System, method, and computer program product for classification based on multiple imbalanced classes of data
The method addresses the challenges of imbalanced datasets in machine learning by dividing data samples into bins based on class proportions and training multiple classifier models on these balanced data blocks, resulting in improved classification accuracy and reduced bias.
Patent Information
- Application Number
- PCT/US2023/081527
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
Existing machine learning models face challenges in classification tasks due to imbalanced datasets, leading to bias, overtraining, and undertraining issues.
A computer-implemented method and system for classification that involves receiving data samples from multiple imbalanced classes, dividing the data samples into bins based on a ratio of class proportions, generating data blocks by combining the majority class samples with selected bins from the minority classes, and training multiple classifier models on these data blocks to improve classification accuracy.
This approach helps to balance the class proportions, reduce bias, and improve the performance of machine learning models by ensuring all data samples are considered and pattern diversity is maintained.
Smart Images

Figure US2023081527_05062025_PF_FP_ABST
Abstract
Description
SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR CLASSIFICATION BASED ON MULTIPLE IMBALANCED CLASSES OF DATABACKGROUND1. Field
[0001] This disclosure relates generally to classification by machine learning models and, in some non-limiting embodiments or aspects, to systems, methods, and computer program products for classification based on multiple imbalanced classes of data.2. Technical Considerations
[0002] Machine learning models are successful in performing many real-world tasks. However, such machine learning models require large datasets for training, and the effectiveness of the model can be impacted by the training dataset. For example, imbalanced data (e.g., class imbalance) is a common issue with datasets for training such machine learning models. A dataset with skewed proportions (e.g., class proportions) may be considered imbalanced. For example, when one or more classes (e.g., the majority class(es)) make up a larger proportion of the dataset than one or more other classes (e.g., the minority class(es)), the dataset may be considered imbalanced. Imbalanced datasets may lead to bias in a machine learning model when such datasets are used for training. Additionally or alternatively, imbalanced datasets may lead to overtraining with respect to some classes (e.g., the larger / majority classes) and / or undertraining with respect to other classes (e.g., the smaller / minority classes).
[0003] Undersampling techniques attempt to balance the classes by using only a portion of the data samples from the majority class(es), but undersampling results in losing useful information from the data samples from the majority class(es) that are not sampled. Oversampling techniques attempt to balance the classes by generating synthetic data samples for the minority class(es), but oversampling may introduce incorrect information and increases the size of the training dataset, which increases the computational burden of training the machine learning model. Partitioning and bagging techniques are limited to binary classification (i.e., only two classes), and partitioning and bagging techniques rely on randomly partitioning the majority class, which may result in the patterns of different bins being correlated (i.e., multiple binsmay share similar information), which results in reduced pattern diversity and negatively impacts performance of the machine learning model.SUMMARY
[0004] Accordingly, it is an object of the present disclosure to provide systems, methods, and computer program products for classification based on multiple imbalanced classes of data (e.g., that overcome some or all of the deficiencies identified above).
[0005] According to non-limiting embodiments or aspects, provided is a computer- implemented method for classification based on multiple imbalanced classes of data. An example method may include receiving a plurality of data samples associated with a plurality of classes including a first class and a plurality of second classes. A first number of data samples of the plurality of data samples may be associated with the first class and a respective second number of data samples may be associated with each second class of the plurality of second classes. For each second class of the plurality of second classes, the respective second number of data samples may be divided into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples. A plurality of data blocks may be generated. Each data block of the plurality of data blocks may include the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes. A plurality of classifier models may be trained based on the plurality of data blocks. Each respective classifier model of the plurality of classifier models may be trained based on a respective data block of the plurality of data blocks. A target data sample may be received. The target data sample may be inputted to the plurality of classifier models to generate a plurality of outputs. A classification of the target data sample may be determined based on the plurality of outputs.
[0006] In some non-limiting embodiments or aspects, dividing the respective second number of data samples into a plurality of bins may include determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples, determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples, and / or determining a respective class proportion ratio for each second class based on rounding a ratiobetween the respective class proportion and the first class proportion. A respective number of bins of the plurality of bins for each second class of the plurality of second classes may be based on the respective class proportion ratio.
[0007] In some non-limiting embodiments or aspects, generating the plurality of data blocks may include cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
[0008] In some non-limiting embodiments or aspects, cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes may include, for a first data block of the plurality of data blocks, selecting a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and, for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
[0009] In some non-limiting embodiments or aspects, at least one unsupervised outlier scoring model may be trained based on the first number of data samples associated with the first class. In some non-limiting embodiments or aspects, for each second class of the plurality of second classes, an outlier score may be generated for each data sample of the respective second number of data samples based on the at least one unsupervised outlier scoring model and / or the respective second number of data samples may be sorted based on the outlier score for each data sample to provide sorted data samples. Dividing the respective second number of data samples into the plurality of bins may include dividing the sorted data samples into the plurality of bins.
[0010] In some non-limiting embodiments or aspects, sorting the respective second number of data samples may include ordering the respective second number of data samples from a lowest outlier score to a highest outlier score. Dividing the sorted data samples into the plurality of bins may include assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
[0011] In some non-limiting embodiments or aspects, the at least one unsupervised outlier scoring model may include a plurality of outlier scoring models. Training the at least one unsupervised outlier scoring model may include training the plurality of unsupervised outlier scoring models based on the first number of data samplesassociated with the first class. Generating the outlier score for each data sample of the respective second number of data samples may include generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models. Sorting the respective second number of data samples may include, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples. Dividing the respective second number of data samples into the plurality of bins may include, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
[0012] In some non-limiting embodiments or aspects, generating the plurality of data blocks may include generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model. Training the plurality of classifier models may include training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
[0013] In some non-limiting embodiments or aspects, inputting the target data sample may include inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model. Determining the classification may include determining a respective classification for each respective unsupervised outlier scoring model.
[0014] In some non-limiting embodiments or aspects, at least one metric may be determined for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model. A selected unsupervised outlier scoring model may be selected from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.
[0015] According to non-limiting embodiments or aspects, provided is a system for classification based on multiple imbalanced classes of data. An example system may include at least one processor configured to receive a plurality of data samples associated with a plurality of classes including a first class and a plurality of secondclasses. A first number of data samples of the plurality of data samples may be associated with the first class, and a respective second number of data samples may be associated with each second class of the plurality of second classes. The at least one processor may be configured to, for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples. The at least one processor may be configured to generate a plurality of data blocks, and each data block of the plurality of data blocks may include the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes. The at least one processor may be configured to train a plurality of classifier models based on the plurality of data blocks, and each respective classifier model of the plurality of classifier models may be trained based on a respective data block of the plurality of data blocks. The at least one processor may be configured to receive a target data sample. The at least one processor may be configured to input the target data sample to the plurality of classifier models to generate a plurality of outputs. The at least one processor may be configured to determine a classification of the target data sample based on the plurality of outputs.
[0016] In some non-limiting embodiments or aspects, dividing the respective second number of data samples into a plurality of bins may include determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples, determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples, and / or determining a respective class proportion ratio for each second class based on rounding a ratio between the respective class proportion and the first class proportion. A respective number of bins of the plurality of bins for each second class of the plurality of second classes may be based on the respective class proportion ratio.
[0017] In some non-limiting embodiments or aspects, generating the plurality of data blocks may include cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
[0018] In some non-limiting embodiments or aspects, cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes may include, for a first data block of the plurality of data blocks, selecting afirst bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and, for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
[0019] In some non-limiting embodiments or aspects, the at least one processor may be further configured to train at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class; and, for each second class of the plurality of second classes, generate an outlier score for each data sample of the respective second number of data samples and sort the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples. Dividing the respective second number of data samples into the plurality of bins may include dividing the sorted data samples into the plurality of bins.
[0020] In some non-limiting embodiments or aspects, sorting the respective second number of data samples may include ordering the respective second number of data samples from a lowest outlier score to a highest outlier score. Dividing the sorted data samples into the plurality of bins may include assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
[0021] In some non-limiting embodiments or aspects, the at least one unsupervised outlier scoring model may include a plurality of outlier scoring models. Training the at least one unsupervised outlier scoring model may include training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class. Generating the outlier score for each data sample of the respective second number of data samples may include generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models. Sorting the respective second number of data samples may include, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples. Dividing the respective second number of data samples into the plurality of bins may include, foreach respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
[0022] In some non-limiting embodiments or aspects, generating the plurality of data blocks may include generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model. Training the plurality of classifier models may include training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
[0023] In some non-limiting embodiments or aspects, inputting the target data sample may include inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model. Determining the classification may include determining a respective classification for each respective unsupervised outlier scoring model. In some non-limiting embodiments or aspects, the at least one processor may be further configured to determine at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model and / or select a selected unsupervised outlier scoring model from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.
[0024] According to non-limiting embodiments or aspects, provided is a computer program product for classification based on multiple imbalanced classes of data. An example computer program product may include at least one non-transitory computer- readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to receive a plurality of data samples associated with a plurality of classes including a first class and a plurality of second classes. A first number of data samples of the plurality of data samples may be associated with the first class and a respective second number of data samples may be associated with each second class of the plurality of second classes. The program instructions, when executed by the at least one processor, may cause the at least one processor to, for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of datasamples. The program instructions, when executed by the at least one processor, may cause the at least one processor to generate a plurality of data blocks. Each data block of the plurality of data blocks may include the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes. The program instructions, when executed by the at least one processor, may cause the at least one processor to train a plurality of classifier models based on the plurality of data blocks. Each respective classifier model of the plurality of classifier models may be trained based on a respective data block of the plurality of data blocks. The program instructions, when executed by the at least one processor, may cause the at least one processor to receive a target data sample. The program instructions, when executed by the at least one processor, may cause the at least one processor to input the target data sample to the plurality of classifier models to generate a plurality of outputs. The program instructions, when executed by the at least one processor, may cause the at least one processor to determine a classification of the target data sample based on the plurality of outputs.
[0025] According to non-limiting embodiments or aspects, provided is a system for classification based on multiple imbalanced classes of data. An example system may include at least one processor configured to perform any of the methods described herein.
[0026] According to non-limiting embodiments or aspects, provided is a computer program product for classification based on multiple imbalanced classes of data. An example computer program product may include at least one non-transitory computer- readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to perform and of the methods described herein.
[0027] Further embodiments or aspects are set forth in the following numbered clauses:
[0028] Clause 1 : A computer-implemented method, comprising: receiving, with at least one processor, a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and a respective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, dividing, with at least one processor, the respective second number of data samplesinto a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generating, with at least one processor, a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; training, with at least one processor, a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks; receiving, with at least one processor, a target data sample; inputting, with at least one processor, the target data sample to the plurality of classifier models to generate a plurality of outputs; and determining, with at least one processor, a classification of the target data sample based on the plurality of outputs.
[0029] Clause 2: The method of clause 1 , wherein dividing the respective second number of data samples into a plurality of bins comprises: determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples; determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples; and determining a respective class proportion ratio for each second class based on rounding a ratio between the respective class proportion and the first class proportion, wherein a respective number of bins of the plurality of bins for each second class of the plurality of second classes is based on the respective class proportion ratio.
[0030] Clause 3: The method of clause 1 or clause 2, wherein generating the plurality of data blocks comprises cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
[0031] Clause 4: The method of any of clauses 1 -3, wherein cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes comprises: for a first data block of the plurality of data blocks, selecting a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
[0032] Clause 5: The method of any of clauses 1 -4, further comprising: training at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class; and for each second class of the plurality of second classes: generating an outlier score for each data sample of the respective second number of data samples based on the at least one unsupervised outlier scoring model; and sorting the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples, wherein dividing the respective second number of data samples into the plurality of bins comprises dividing the sorted data samples into the plurality of bins.
[0033] Clause 6: The method of any of clauses 1 -5, wherein sorting the respective second number of data samples comprises ordering the respective second number of data samples from a lowest outlier score to a highest outlier score, and wherein dividing the sorted data samples into the plurality of bins comprises assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
[0034] Clause 7: The method of any of clauses 1 -6, wherein the at least one unsupervised outlier scoring model comprises a plurality of outlier scoring models, wherein training the at least one unsupervised outlier scoring model comprises training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class, wherein generating the outlier score for each data sample of the respective second number of data samples comprises generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models, wherein sorting the respective second number of data samples comprises, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples, and wherein dividing the respective second number of data samples into the plurality of bins comprises, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
[0035] Clause 8: The method of any of clauses 1 -7, wherein generating the plurality of data blocks comprises generating a respective plurality of data blocks based on therespective plurality of bins for each respective unsupervised outlier scoring model, and wherein training the plurality of classifier models comprises training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
[0036] Clause 9: The method of any of clauses 1 -8, wherein inputting the target data sample comprises inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model, and wherein determining the classification comprises determining a respective classification for each respective unsupervised outlier scoring model.
[0037] Clause 10: The method of any of clauses 1 -9, further comprising: determining, with at least one processor, at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model; and selecting, with at least one processor, a selected unsupervised outlier scoring model from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.
[0038] Clause 1 1 : A system comprising: at least one processor configured to: receive a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and a respective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generate a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; train a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks; receive a target data sample; input the target data sample to the plurality of classifier models to generate a plurality of outputs; and determine a classification of the target data sample based on the plurality of outputs.
[0039] Clause 12: The system of clause 1 1 , wherein dividing the respective second number of data samples into a plurality of bins comprises: determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples; determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples; and determining a respective class proportion ratio for each second class based on rounding a ratio between the respective class proportion and the first class proportion, wherein a respective number of bins of the plurality of bins for each second class of the plurality of second classes is based on the respective class proportion ratio.
[0040] Clause 13: The system of clause 1 1 or clause 12, wherein generating the plurality of data blocks comprises cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
[0041] Clause 14: The system of any of clauses 1 -13, wherein cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes comprises: for a first data block of the plurality of data blocks, selecting a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
[0042] Clause 15: The system of any of clauses 1 -14, wherein the at least one processor is further configured to train at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class; and, for each second class of the plurality of second classes, generate an outlier score for each data sample of the respective second number of data samples; and sort the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples, wherein dividing the respective second number of data samples into the plurality of bins comprises dividing the sorted data samples into the plurality of bins.
[0043] Clause 16: The system of any of clauses 1 -15, wherein sorting the respective second number of data samples comprises ordering the respective second number of data samples from a lowest outlier score to a highest outlier score, andwherein dividing the sorted data samples into the plurality of bins comprises assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
[0044] Clause 17: The system of any of clauses 1 -16, wherein the at least one unsupervised outlier scoring model comprises a plurality of outlier scoring models, wherein training the at least one unsupervised outlier scoring model comprises training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class, wherein generating the outlier score for each data sample of the respective second number of data samples comprises generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models, wherein sorting the respective second number of data samples comprises, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples, and wherein dividing the respective second number of data samples into the plurality of bins comprises, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
[0045] Clause 18: The system of any of clauses 1 -17, wherein generating the plurality of data blocks comprises generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model, and wherein training the plurality of classifier models comprises training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
[0046] Clause 19: The system of any of clauses 1 -18, wherein inputting the target data sample comprises inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model, and wherein determining the classification comprises determining a respective classification for each respective unsupervised outlier scoring model, wherein the at least one processor is further configured to determine at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoringmodel; and select a selected unsupervised outlier scoring model from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.
[0047] Clause 20: A computer program product comprising at least one non- transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to: receive a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and a respective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generate a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; train a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks; receive a target data sample; input the target data sample to the plurality of classifier models to generate a plurality of outputs; and determine a classification of the target data sample based on the plurality of outputs.
[0048] Clause 21 : A system comprising at least one processor configured to perform the method of any of clauses 1 -10.
[0049] Clause 22: A computer program product comprising at least one non- transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any of clauses 1 -10.
[0050] These and other features and characteristics of the present disclosure, as well as the methods of operation and functions of the related elements of structures and the combination of parts and economies of manufacture, will become more apparent upon consideration of the following description and the appended claims with reference to the accompanying drawings, all of which form a part of this specification, wherein like reference numerals designate corresponding parts in the various figures. It is to be expressly understood, however, that the drawings are for the purpose ofillustration and description only and are not intended as a definition of the limits of the disclosed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Additional advantages and details are explained in greater detail below with reference to the non-limiting, exemplary embodiments or aspects that are illustrated in the accompanying schematic figures, in which:
[0052] FIG. 1 is a schematic diagram of a system for classification based on multiple imbalanced classes of data, according to some non-limiting embodiments or aspects;
[0053] FIG. 2 is a flow diagram for a method for classification based on multiple imbalanced classes of data, according to some non-limiting embodiments or aspects;
[0054] FIG. 3 is a diagram of an exemplary environment in which methods, systems, and / or computer program products, described herein, may be implemented, according to some non-limiting embodiments or aspects;
[0055] FIG. 4 is a schematic diagram of example components of one or more devices of FIG. 1 and / or FIG. 3, according to some non-limiting embodiments or aspects; and
[0056] FIGS. 5A-5E are schematic diagrams of implementations of a method for classification based on multiple imbalanced classes of data, according to some nonlimiting embodiments or aspects.DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0057] For purposes of the description hereinafter, the terms “end,” “upper,” “lower,” “right,” “left,” “vertical,” “horizontal,” “top,” “bottom,” “lateral,” “longitudinal,” and derivatives thereof shall relate to the embodiments as they are oriented in the drawing figures. However, it is to be understood that the embodiments may assume various alternative variations and step sequences, except where expressly specified to the contrary. It is also to be understood that the specific devices and processes illustrated in the attached drawings, and described in the following specification, are simply exemplary embodiments or aspects of the disclosed subject matter. Hence, specific dimensions and other physical characteristics related to the embodiments or aspects disclosed herein are not to be considered as limiting.
[0058] Some non-limiting embodiments or aspects are described herein in connection with thresholds. As used herein, satisfying a threshold may refer to a value being greater than the threshold, more than the threshold, higher than the threshold, greater than or equal to the threshold, less than the threshold, fewer than the threshold, lower than the threshold, less than or equal to the threshold, equal to the threshold, etc.
[0059] No aspect, component, element, structure, act, step, function, instruction, and / or the like used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more” and “at least one.” Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, and / or the like) and may be used interchangeably with “one or more” or “at least one.” Where only one item is intended, the term “one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based at least partially on” unless explicitly stated otherwise. In addition, reference to an action being “based on” a condition may refer to the action being “in response to” the condition. For example, the phrases “based on” and “in response to” may, in some non-limiting embodiments or aspects, refer to a condition for automatically triggering an action (e.g., a specific operation of an electronic device, such as a computing device, a processor, and / or the like).
[0060] As used herein, the term “acquirer institution” may refer to an entity licensed and / or approved by a transaction service provider to originate transactions (e.g., payment transactions) using a payment device associated with the transaction service provider. The transactions the acquirer institution may originate may include payment transactions (e.g., purchases, original credit transactions (OCTs), account funding transactions (AFTs), and / or the like). In some non-limiting embodiments or aspects, an acquirer institution may be a financial institution, such as a bank. As used herein, the term “acquirer system” may refer to one or more computing devices operated by or on behalf of an acquirer institution, such as a server computer executing one or more software applications.
[0061] As used herein, the term “account identifier” may include one or more primary account numbers (PANs), tokens, or other identifiers associated with acustomer account. The term “token” may refer to an identifier that is used as a substitute or replacement identifier for an original account identifier, such as a PAN. Account identifiers may be alphanumeric or any combination of characters and / or symbols. Tokens may be associated with a PAN or other original account identifier in one or more data structures (e.g., one or more databases, and / or the like) such that they may be used to conduct a transaction without directly using the original account identifier. In some examples, an original account identifier, such as a PAN, may be associated with a plurality of tokens for different individuals or purposes.
[0062] As used herein, the terms “electronic wallet” and “electronic wallet application” refer to one or more electronic devices and / or software applications configured to initiate and / or conduct payment transactions. For example, an electronic wallet may include a mobile device executing an electronic wallet application, and may further include server-side software and / or databases for maintaining and providing transaction data to the mobile device. An “electronic wallet provider” may include an entity that provides and / or maintains an electronic wallet for a customer, such as Google Pay®, Android Pay®, Apple Pay®, Samsung Pay®, and / or other like electronic payment systems. In some non-limiting examples, an issuer bank may be an electronic wallet provider.
[0063] As used herein, the term “issuer institution” may refer to one or more entities, such as a bank, that provide accounts to customers for conducting transactions (e.g., payment transactions), such as initiating credit and / or debit payments. For example, an issuer institution may provide an account identifier, such as a PAN, to a customer that uniquely identifies one or more accounts associated with that customer. The account identifier may be embodied on a portable financial device, such as a physical financial instrument, e.g., a payment card, and / or may be electronic and used for electronic payments. The term “issuer system” refers to one or more computer devices operated by or on behalf of an issuer institution, such as a server computer executing one or more software applications. For example, an issuer system may include one or more authorization servers for authorizing a transaction.
[0064] As used herein, the term “merchant” may refer to an individual or entity that provides goods and / or services, or access to goods and / or services, to customers based on a transaction, such as a payment transaction. The term “merchant” or “merchant system” may also refer to one or more computer systems operated by oron behalf of a merchant, such as a server computer executing one or more software applications.
[0065] As used herein, a “point-of-sale (POS) device” may refer to one or more devices, which may be used by a merchant to conduct a transaction (e.g., a payment transaction) and / or process a transaction. For example, a POS device may include one or more client devices. Additionally or alternatively, a POS device may include peripheral devices, card readers, scanning devices (e.g., code scanners), Bluetooth® communication receivers, near-field communication (NFC) receivers, radio frequency identification (RFID) receivers, and / or other contactless transceivers or receivers, contact-based receivers, payment terminals, and / or the like. As used herein, a “point- of-sale (POS) system” may refer to one or more client devices and / or peripheral devices used by a merchant to conduct a transaction. For example, a POS system may include one or more POS devices and / or other like devices that may be used to conduct a payment transaction. In some non-limiting embodiments or aspects, a POS system (e.g., a merchant POS system) may include one or more server computers programmed or configured to process online payment transactions through webpages, mobile applications, and / or the like.
[0066] As used herein, the terms “client” and “client device” may refer to one or more client-side devices or systems (e.g., remote from a transaction service provider) used to initiate or facilitate a transaction (e.g., a payment transaction). As an example, a “client device” may refer to one or more POS devices used by a merchant, one or more acquirer host computers used by an acquirer, one or more mobile devices used by a user, and / or the like. In some non-limiting embodiments or aspects, a client device may be an electronic device configured to communicate with one or more networks and initiate or facilitate transactions. For example, a client device may include one or more computers, portable computers, laptop computers, tablet computers, mobile devices, cellular phones, wearable devices (e.g., watches, glasses, lenses, clothing, and / or the like), PDAs, and / or the like. Moreover, a “client” may also refer to an entity (e.g., a merchant, an acquirer, and / or the like) that owns, utilizes, and / or operates a client device for initiating transactions (e.g., for initiating transactions with a transaction service provider).
[0067] As used herein, the term “payment device” may refer to a payment card (e.g., a credit or debit card), a gift card, a smartcard, smart media, a payroll card, a healthcare card, a wristband, a machine-readable medium containing accountinformation, a keychain device or fob, an RFID transponder, a retailer discount or loyalty card, a cellular phone, an electronic wallet mobile application, a personal digital assistant (PDA), a pager, a security card, a computing device, an access card, a wireless terminal, a transponder, and / or the like. In some non-limiting embodiments or aspects, the payment device may include volatile or non-volatile memory to store information (e.g., an account identifier, a name of the account holder, and / or the like).
[0068] As used herein, the term “payment gateway” may refer to an entity and / or a payment processing system operated by or on behalf of such an entity (e.g., a merchant service provider, a payment service provider, a payment facilitator, a payment facilitator that contracts with an acquirer, a payment aggregator, and / or the like), which provides payment services (e.g., transaction service provider payment services, payment processing services, and / or the like) to one or more merchants. The payment services may be associated with the use of portable financial devices managed by a transaction service provider. As used herein, the term “payment gateway system” may refer to one or more computer systems, computer devices, servers, groups of servers, and / or the like, operated by or on behalf of a payment gateway.
[0069] As used herein, the term “transaction service provider” may refer to an entity that receives transaction authorization requests from merchants or other entities and provides guarantees of payment, in some cases through an agreement between the transaction service provider and an issuer institution. For example, a transaction service provider may include a payment network such as Visa® or any other entity that processes transactions. The term “transaction processing system” may refer to one or more computer systems operated by or on behalf of a transaction service provider, such as a transaction processing server executing one or more software applications. A transaction processing server may include one or more processors and, in some non-limiting embodiments or aspects, may be operated by or on behalf of a transaction service provider.
[0070] As used herein, the term “communication” may refer to the reception, receipt, transmission, transfer, provision, and / or the like of data (e.g., information, signals, messages, instructions, commands, and / or the like). For one unit (e.g., a device, a system, a component of a device or system, combinations thereof, and / or the like) to be in communication with another unit means that the one unit is able to directly or indirectly receive information from and / or transmit information to the otherunit. This may refer to a direct or indirect connection (e.g., a direct communication connection, an indirect communication connection, and / or the like) that is wired and / or wireless in nature. Additionally, two units may be in communication with each other even though the information transmitted may be modified, processed, relayed, and / or routed between the first and second unit. For example, a first unit may be in communication with a second unit even though the first unit passively receives information and does not actively transmit information to the second unit. As another example, a first unit may be in communication with a second unit if at least one intermediary unit processes information received from the first unit and communicates the processed information to the second unit. In some non-limiting embodiments or aspects, a message may refer to a network packet (e.g., a data packet and / or the like) that includes data. It will be appreciated that numerous other arrangements are possible.
[0071] As used herein, the term “computing device” may refer to one or more electronic devices configured to process data. A computing device may, in some examples, include the necessary components to receive, process, and output data, such as a processor, a display, a memory, an input device, a network interface, and / or the like. A computing device may be a mobile device. As an example, a mobile device may include a cellular phone (e.g., a smartphone or standard cellular phone), a portable computer (e.g., laptop computer, a tablet computer, and / or the like), a wearable device (e.g., watches, glasses, lenses, clothing, and / or the like), a personal digital assistant (PDA), and / or other like devices. A computing device may also be a desktop computer or other form of non-mobile computer.
[0072] As used herein, the term “server” may refer to or include one or more computing devices that are operated by or facilitate communication and processing for multiple parties in a network environment, such as the Internet, although it will be appreciated that communication may be facilitated over one or more public or private network environments and that various other arrangements are possible. Further, multiple computing devices (e.g., servers, point-of-sale (POS) devices, mobile devices, etc.) directly or indirectly communicating in the network environment may constitute a “system.” Reference to “a server” or “a processor,” as used herein, may refer to a previously-recited server and / or processor that is recited as performing a previous step or function, a different server and / or processor, and / or a combination of servers and / or processors. For example, as used in the specification and the claims,a first server and / or a first processor that is recited as performing a first step or function may refer to the same or different server and / or a processor recited as performing a second step or function.
[0073] As used herein, the term “system” may refer to one or more computing devices or combinations of computing devices (e.g., processors, servers, client devices, software applications, components of such, and / or the like). Reference to “a device,” “a server,” “a processor,” and / or the like, as used herein, may refer to a previously-recited device, server, or processor that is recited as performing a previous step or function, a different server or processor, and / or a combination of servers and / or processors. For example, as used in the specification and the claims, a first server or a first processor that is recited as performing a first step or a first function may refer to the same or different server or the same or different processor recited as performing a second step or a second function.
[0074] Non-limiting embodiments or aspects of the disclosed subject matter are directed to systems, methods, and computer program products for classification based on multiple imbalanced classes of data, including. For example, non-limiting embodiments or aspects of the disclosed subject matter provide receiving a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes. For example, a first number of data samples may be associated with the first class, and a respective second number of data samples may be associated with each second class. For each second class, the respective second number of data samples may be divided into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples. A plurality of data blocks may be generated, and each data block may include the first number of data samples (e.g., a bin of all the data samples associated with the first class) and a selected bin of the plurality of bins associated with each second class of the plurality of second classes. A plurality of classifier models may be trained based on the plurality of data blocks. For example, each respective classifier model may be trained based on a respective data block. A target data sample may be received (e.g., after training). The target data sample may be inputted to the plurality of (trained) classifier models to generate a plurality of outputs. A classification of the target data sample may be determined based on the plurality of outputs. Such embodiments or aspects allow for elimination of bias from imbalanced data (e.g., due to the manner in which the data blocks are generated and used to train multipleclassifier models). Additionally, such embodiments or aspects enable modeling of imbalanced data when any number of classes is present (e.g., rather than being limited to binary classification between two classes). Additionally, such embodiments or aspects enable improved performance compared to oversampling techniques (e.g., because there is no reliance on potentially inaccurate synthetic data samples). Additionally, such embodiments or aspects enable improved performance compared to undersampling techniques (e.g., because all data samples are included in the data blocks to prevent loss of information by disregarding some data samples).
[0075] Additionally, non-limiting embodiments or aspects of the disclosed subject matter provide training at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class. For each second class, an outlier score for each data sample may be generated, and the respective second number of data samples may be sorted based on the outlier scores to provide sorted data samples. The sorted data samples may be divided into the plurality of bins. Such embodiments or aspects allow for improved (e.g., increased) pattern diversity of the data blocks (e.g., compared to randomly sorting the data samples into bins). Additionally, a plurality of outlier scoring models may be used, and the techniques described herein may be repeated for each outlier scoring model to determine which outlier scoring model results in the best performance. As such, the outlier scoring model that enhances performance (e.g., best improved pattern diversity) may be selected, thereby improving performance of the machine learning model trained based on the data blocks.
[0076] For the purpose of illustration, in the following description, while the presently disclosed subject matter is described with respect to systems, methods, and computer program products for classification based on multiple imbalanced classes of data, one skilled in the art will recognize that the disclosed subject matter is not limited to the illustrative embodiments or aspects. For example, the systems, methods, and computer program products described herein may be used with a wide variety of settings, such as imbalanced classes of data for any suitable type of machine learning task, e.g., prediction, recommendation, fraud detection, anomaly detection, risk detection, credit default prediction, credit application approval, transaction approval, and / or the like.
[0077] Referring now to FIG. 1 , shown is an example system 100 for classification based on multiple imbalanced classes of data, according to some non-limitingembodiments or aspects. As shown in FIG. 1 , system 100 may include machine learning model system 102, database 104, user device 106, and / or communication network 1 12.
[0078] Machine learning model system 102 may include one or more devices capable of receiving information from and / or communicating information to database 104 and / or user device 106 (e.g., directly via wired or wireless communication connection, indirectly via communication network 1 12, and / or the like). For example, machine learning model system 102 may include a computing device, such as a server, a group of servers, a desktop computer, a portable computer, a mobile device, and / or other like devices. In some non-limiting embodiments or aspects, machine learning model system 102 may be in communication with a data storage device (e.g., database 104), which may be local or remote to machine learning model system 102. In some non-limiting embodiments or aspects, machine learning model system 102 may be capable of receiving information from, storing information in, communicating information to, or searching information stored in the data storage device (e.g., database 104).
[0079] Database 104 may include one or more devices capable of receiving information from and / or communicating information to machine learning model system 102 and / or user device 106 (e.g., directly via wired or wireless communication connection, indirectly via communication network 1 12, and / or the like). For example, database 104 may include a computing device, such as a server, a group of servers, a desktop computer, a portable computer, a mobile device, and / or other like devices. In some non-limiting embodiments or aspects, database 104 may include a data storage device. In some non-limiting embodiments or aspects, database 104 may be capable of receiving information from, storing information in, communicating information to, or searching information stored in the data storage device. In some non-limiting embodiments or aspects, database 104 may be part of machine learning model system 102 and / or part of the same system as machine learning model system 102.
[0080] User device 106 may include one or more devices capable of receiving information from and / or communicating information to machine learning model system 102 and / or database 104 (e.g., directly via wired or wireless communication connection, indirectly via communication network 1 12, and / or the like). For example, user device 106 may include a computing device, such as a mobile device, a portablecomputer, a desktop computer, and / or other like devices. Additionally or alternatively, each user device 106 may include a device capable of receiving information from and / or communicating information to other user devices 106 (e.g., directly via wired or wireless communication connection, indirectly via communication network 1 12, and / or the like). In some non-limiting embodiments or aspects, user device 106 may be part of machine learning model system 102 and / or part of the same system as machine learning model system 102. For example, machine learning model system 102, database 104, and user device 106 may all be (and / or be part of) a single system and / or a single computing device.
[0081] Communication network 1 12 may include one or more wired and / or wireless networks. For example, communication network 1 12 may include a cellular network (e.g., a long-term evolution (LTE®) network, a third generation (5G) network, a fourth generation (4G) network, a fifth generation (5G) network, a code division multiple access (CDMA) network, and / or the like), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., the public switched telephone network (PSTN)), a private network (e.g., a private network associated with a transaction service provider), an ad hoc network, an intranet, the Internet, a fiber optic-based network, a cloud computing network, and / or the like, and / or a combination of these or other types of networks.
[0082] The number and arrangement of systems and devices shown in FIG. 1 are provided as an example. There may be additional systems and / or devices, fewer systems and / or devices, different systems and / or devices, and / or differently arranged systems and / or devices than those shown in FIG. 1. Furthermore, two or more systems or devices shown in FIG. 1 may be implemented within a single system or device, or a single system or device shown in FIG. 1 may be implemented as multiple, distributed systems or devices. Additionally or alternatively, a set of systems (e.g., one or more systems) or a set of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another set of systems or another set of devices of system 100.
[0083] Referring now to FIG. 2, shown is an example method 200 for classification based on multiple imbalanced classes of data, according to some non-limiting embodiments or aspects. The steps shown in FIG. 2 are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limitingembodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of method 200 may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of method 200 may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106.
[0084] As shown in FIG. 2, at step 202, method 200 may include receiving data samples. For example, machine learning model system 102 may receive (e.g., from database 104, user device 106, and / or the like) a plurality of data samples (e.g., a dataset of imbalanced data).
[0085] In some non-limiting embodiments or aspects, machine learning model system 102 may receive a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes. For example, a first number of data samples of the plurality of data samples may be associated with the first class, and a respective second number of data samples may be associated with each second class of the plurality of second classes. For the purpose of illustration, the first class may be the smallest class (e.g., a minority class). Additionally or alternatively, the second classes (e.g., majority classes) may be larger than the first class.
[0086] In some non-limiting embodiments or aspects, the classes may be sorted (e.g., by machine learning model system 102) based on the respective size of each class. For example, the first class may be the smallest class, and the second classes may be sorted based on the respective size of each class (e.g., from smallest to largest, from largest to smallest, and / or the like).
[0087] As shown in FIG. 2, at step 204, method 200 may include dividing at least some of the classes into bins. For example, machine learning model system 102 may divide (at least some of) the classes into bins.
[0088] In some non-limiting embodiments or aspects, machine learning model system 102 may, for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of datasamples. For example, machine learning model system 102 may determine a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples. Additionally or alternatively, machine learning model system 102 may determine a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples. In some non-limiting embodiments or aspects, machine learning model system 102 may determine a respective class proportion ratio for each second class based on rounding (e.g., to the nearest integer and / or the like) a ratio between the respective class proportion and the first class proportion. For example, a respective number of bins of the plurality of bins for each second class may be based on the respective class proportion ratio.
[0089] In some non-limiting embodiments or aspects, machine learning model system 102 may train at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class. For each second class of the plurality of second classes, machine learning model system 102 may generate an outlier score for each data sample of the respective second number of data samples based on the at least one unsupervised outlier scoring model. Machine learning model system 102 may sort the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples. In some nonlimiting embodiments or aspects, dividing the respective second number of data samples into the plurality of bins may include dividing the sorted data samples into the plurality of bins. As such, the pattern diversity of data blocks generated based on such bins, as described herein, may be improved (e.g., increased), for example, compared to randomly dividing the data samples into bins.
[0090] In some non-limiting embodiments or aspects, the unsupervised outlier scoring model(s) may include at least one of a density-based model (e.g., a kernel density estimation model, a Gaussian mixture model, and / or the like) a tree-based model (e.g., an isolation forest model and / or the like), a neural network model (e.g., an autoencoder and / or the like), any combination thereof, and / or the like.
[0091] In some non-limiting embodiments or aspects, sorting the respective second number of data samples may include ordering the respective second number of data samples from a lowest outlier score to a highest outlier score. For example, dividing the sorted data samples into the plurality of bins may include assigning each data sample of the sorted data samples to a bin of the plurality of bins based on orderingthe respective second number of data samples from a lowest outlier score to a highest outlier score.
[0092] In some non-limiting embodiments or aspects, the at least one unsupervised outlier scoring model may include a plurality of outlier scoring models. For example, training the at least one unsupervised outlier scoring model may include training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class. In some non-limiting embodiments or aspects, generating the outlier score for each data sample of the respective second number of data samples may include generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models. In some non-limiting embodiments or aspects, sorting the respective second number of data samples may include, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples. In some non-limiting embodiments or aspects, dividing the respective second number of data samples into the plurality of bins may include, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
[0093] As shown in FIG. 2, at step 206, method 200 may include generating data blocks. For example, machine learning model system 102 may generate a plurality of data blocks.
[0094] In some non-limiting embodiments or aspects, machine learning model system 102 may generate a plurality of data blocks, and each data block may include the first number of data samples (e.g., all data samples of the first class) and a selected bin of the plurality of bins associated with each second class of the plurality of second classes.
[0095] In some non-limiting embodiments or aspects, machine learning model system 102 may generate the plurality of data blocks by cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes. For example, cyclically iterating through bins may include, for a first data block of the plurality of data blocks, selecting (e.g., by machine learning model system 102) a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with eachsecond class of the plurality of second classes. For each successive data block after the first data block, machine learning model system 102 may select a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected. In some non-limiting embodiments or aspects, after the last bin of a given second class is selected, the first bin of that second class may be selected for the next data block, and the cycle may repeat for that given second class. In some non-limiting embodiments or aspects, the cyclical iteration through bins of each second class may continue until all bins of every second class have been included in at least one data block.
[0096] As such, the disclosed subject matter ensures that, in each data block, the class proportion is balanced. Additionally, the disclosed subject matter ensures that all samples are considered at least once. Moreover, the disclosed subject matter ensures that, within each class, all data samples are used roughly the same amount of times (e.g., at most one more time or one fewer time compared to any other data sample in the same class), thereby avoiding sampling bias.
[0097] In some non-limiting embodiments or aspects, when a plurality of unsupervised outlier scoring models are used, as described herein, generating the plurality of data blocks may include generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model. For example, an entire set of data blocks (e.g., a separate plurality of data blocks) may be generated as described herein for each one of the unsupervised outlier score models.
[0098] As shown in FIG. 2, at step 208, method 200 may include training classifier models. For example, machine learning model system 102 may train a plurality of classifier models based on the plurality of data blocks.
[0099] In some non-limiting embodiments or aspects, each respective classifier model of the plurality of classifier models may be trained based on a respective data block of the plurality of data blocks. As such, the disclosed subject matter allows for elimination of bias from imbalanced data (e.g., due to the manner in which the data blocks are generated and used to train multiple classifier models). Additionally, the disclosed subject matter enables modeling of imbalanced data when any number of classes is present (e.g., rather than being limited to binary classification between two classes). Additionally, the disclosed subject matter enables improved performance compared to oversampling techniques (e.g., because there is no reliance onpotentially inaccurate synthetic data samples) and compared to undersampling techniques (e.g., because all data samples are included in the data blocks to prevent loss of information by disregarding some data samples).
[0100] In some non-limiting embodiments or aspects, training the classifier models may include model analysis system 102 inputting each respective data block to the respective classifier model to generate predicted outputs. The predicted outputs may be compared to known classifications (e.g., labels and / or the like) associated with the data samples of the data block, and / or a loss may be determined based on the predicted outputs and the known classifications (e.g., calculated based on a loss function, the predicted outputs, and the known classifications). In some embodiments, machine learning model system 102 may update (e.g., adjust) the parameters (e.g., weights, connection values, and / or the like) of the respective classifier model based on the loss (e.g., using back propagation, gradient calculations, and / or the like).
[0101] In some non-limiting embodiments or aspects, when a plurality of unsupervised outlier scoring models are used, as described herein, training the plurality of classifier models may include training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model. For example, an entire set of classifier models may be separately trained based on the set of data blocks for each one of the unsupervised outlier score models.
[0102] As shown in FIG. 2, at step 210, method 200 may include receiving at least one target data sample. For example, machine learning model system 102 may receive a target data sample (e.g., from database 104, user device 106, and / or the like).
[0103] As shown in FIG. 2, at step 212, method 200 may include inputting the target data sample(s) to the classifiers to generate outputs. For example, machine learning model system 102 may input a target data sample to (each of) the plurality of classifier models to generate a plurality of outputs (e.g., intermediate outputs).
[0104] In some non-limiting embodiments or aspects, the target data sample may include a new data sample (e.g., for which there is no known classification). As such, machine learning model system 102 may input the new data sample to each of the classifier models to generate a plurality of outputs (e.g., a respective predicted classification from each respective classifier model).
[0105] In some non-limiting embodiments or aspects, the target data sample may include a test data sample (e.g., for which there is a known classification, such as a label and / or the like).
[0106] In some non-limiting embodiments or aspects, when a plurality of unsupervised outlier scoring models are used, as described herein, inputting the target data sample may include inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
[0107] As shown in FIG. 2, at step 214, method 200 may include determining at least one final output (e.g., at least one classification). For example, machine learning model system 102 may determine a classification of the target data sample (e.g., a final output) based on the plurality of outputs (e.g., intermediate outputs) of the plurality of classifier models.
[0108] In some non-limiting embodiments or aspects, machine learning model system 102 may determine the classification based on an average of the plurality of outputs (e.g., intermediate outputs, such as predicted values and / or the like) of the plurality of classifier models. In some non-limiting embodiments or aspects, machine learning model system 102 may determine the classification based on majority voting of the plurality of classifier models (e.g., whichever predicted classification is predicted by the most classifier models is determined to be the final classification).
[0109] In some non-limiting embodiments or aspects, when a plurality of unsupervised outlier scoring models are used, as described herein, machine learning model system 102 may determine a respective classification associated with using each respective unsupervised outlier scoring model (e.g., a plurality of respective classifications, each respective classification associated with using a respective one of the unsupervised outlier scoring models).
[0110] In some non-limiting embodiments or aspects, machine learning model system 102 may determine at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model. Additionally or alternatively, machine learning model system 102 may select an unsupervised outlier scoring model (e.g., a selected unsupervised outlier scoring model) from the plurality of unsupervised outlier scoring models (e.g., based on the metric(s) for each respective unsupervised outlier scoring). For example, the unsupervised outlier scoring model that results in the best performance (e.g., bestdetermination of final classifications based on a plurality of test data samples) may be selected (e.g., machine learning model system 102) as the best unsupervised outlier scoring model. As such, a plurality of outlier scoring models may be used, and the techniques described herein may be repeated for each outlier scoring model to determine which outlier scoring model results in the best performance. In this way, the outlier scoring model that enhances performance may be selected, thereby improving performance of the machine learning models trained based on the data blocks.
[0111] In some non-limiting embodiments or aspects, the metric(s) may include at least one of an accuracy metric, an error metric, a precision metric, a recall metric, an F1 -score, an area under curve (AUG) metric, any combination thereof, and / or the like.
[0112] In some non-limiting embodiments or aspects, the (final) classification may be associated with at least one of a prediction, a recommendation, a detection of fraud, a detection of an anomaly, a detection of risk, a prediction of credit default, an approval of a credit application, an approval of a transaction, and / or the like. For example, a target data sample may be based on a transaction message (e.g., an authorization request and / or the like), and a classification of the target data sample as fraudulent or not fraudulent may be determined, as described herein. For example, a target data sample may be based on a transaction message (e.g., an authorization request and / or the like), and a classification of the target data sample as approved or declined may be determined, as described herein. For example, a target data sample may be based on a credit application, and a classification of credit risk (e.g., high, medium, low, and / or the like) may be determined, as described herein. For example, a target data sample may be based on a credit application, and a classification of the target data sample as approved or rejected may be determined, as described herein. For example, a target data sample may be based on a credit application, and a classification (e.g., prediction) of credit default risk (e.g., short term default risk, long term default risk, no default risk, and / or the like) may be determined, as described herein. For example, a target data sample may be based on a user profile, and a classification of at least one recommendation (e.g., recommended product, service, and / or the like) may be determined, as described herein.
[0113] Referring now to FIG. 3, depicted is a diagram of a non-limiting embodiment or aspect of an exemplary environment 300 in which systems, products, and / or methods, as described herein, may be implemented. As shown in FIG. 3, environment300 may include transaction service provider system 302, issuer system 304, customer device 306, merchant system 308, acquirer system 310, and communication network 312. In some non-limiting embodiments or aspects, each of machine learning model system 102, database 104, and / or user device 106 of FIG. 1 may be implemented by (e.g., part of) transaction service provider system 302. In some nonlimiting embodiments or aspects, at least one of machine learning model system 102, database 104, and / or user device 106 of FIG. 1 may be implemented by (e.g., part of) another system, another device, another group of systems, or another group of devices, separate from or including transaction service provider system 302, such as issuer system 304, customer device 306, merchant system 308, acquirer system 310, and / or the like. For example, user device 106 may be implemented by (e.g., part of) at least one of issuer system 304, customer device 306, merchant system 308, and / or acquirer system 310.
[0114] Transaction service provider system 302 may include one or more devices capable of receiving information from and / or communicating information to issuer system 304, customer device 306, merchant system 308, and / or acquirer system 310 via communication network 312. For example, transaction service provider system 302 may include a computing device, such as a server (e.g., a transaction processing server), a group of servers, and / or other like devices. In some non-limiting embodiments or aspects, transaction service provider system 302 may be associated with a transaction service provider as described herein. In some non-limiting embodiments or aspects, transaction service provider system 302 may be in communication with a data storage device, which may be local or remote to transaction service provider system 302. In some non-limiting embodiments or aspects, transaction service provider system 302 may be capable of receiving information from, storing information in, communicating information to, or searching information stored in the data storage device.
[0115] Issuer system 304 may include one or more devices capable of receiving information and / or communicating information to transaction service provider system 302, customer device 306, merchant system 308, and / or acquirer system 310 via communication network 312. For example, issuer system 304 may include a computing device, such as a server, a group of servers, and / or other like devices. In some non-limiting embodiments or aspects, issuer system 304 may be associated with an issuer institution as described herein. For example, issuer system 304 may beassociated with an issuer institution that issued a credit account, debit account, credit card, debit card, and / or the like to a user associated with customer device 306.
[0116] Customer device 306 may include one or more devices capable of receiving information from and / or communicating information to transaction service provider system 302, issuer system 304, merchant system 308, and / or acquirer system 310 via communication network 312. Additionally or alternatively, each customer device 306 may include a device capable of receiving information from and / or communicating information to other customer devices 306 via communication network 312, another network (e.g., an ad hoc network, a local network, a private network, a virtual private network, and / or the like), and / or any other suitable communication technique. For example, customer device 306 may include a client device and / or the like. In some non-limiting embodiments or aspects, customer device 306 may or may not be capable of receiving information (e.g., from merchant system 308 or from another customer device 306) via a short-range wireless communication connection (e.g., an NFC communication connection, an RFID communication connection, a Bluetooth® communication connection, a Zigbee® communication connection, and / or the like), and / or communicating information (e.g., to merchant system 308) via a short-range wireless communication connection.
[0117] Merchant system 308 may include one or more devices capable of receiving information from and / or communicating information to transaction service provider system 302, issuer system 304, customer device 306, and / or acquirer system 310 via communication network 312. Merchant system 308 may also include a device capable of receiving information from customer device 306 via communication network 312, a communication connection (e.g., an NFC communication connection, an RFID communication connection, a Bluetooth® communication connection, a Zigbee® communication connection, and / or the like) with customer device 306, and / or the like, and / or communicating information to customer device 306 via communication network 312, the communication connection, and / or the like. In some non-limiting embodiments or aspects, merchant system 308 may include a computing device, such as a server, a group of servers, a client device, a group of client devices, and / or other like devices. In some non-limiting embodiments or aspects, merchant system 308 may be associated with a merchant as described herein. In some non-limiting embodiments or aspects, merchant system 308 may include one or more client devices. For example, merchant system 308 may include a client device that allowsa merchant to communicate information to transaction service provider system 302. In some non-limiting embodiments or aspects, merchant system 308 may include one or more devices, such as computers, computer systems, and / or peripheral devices capable of being used by a merchant to conduct a transaction with a user. For example, merchant system 308 may include a POS device and / or a POS system.
[0118] Acquirer system 310 may include one or more devices capable of receiving information from and / or communicating information to transaction service provider system 302, issuer system 304, customer device 306, and / or merchant system 308 via communication network 312. For example, acquirer system 310 may include a computing device, a server, a group of servers, and / or the like. In some non-limiting embodiments or aspects, acquirer system 310 may be associated with an acquirer as described herein.
[0119] Communication network 312 may include one or more wired and / or wireless networks. For example, communication network 312 may include a cellular network (e.g., a long-term evolution (LTE®) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (3G) network, a code division multiple access (CDMA) network, and / or the like), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., the public switched telephone network (PSTN)), a private network (e.g., a private network associated with a transaction service provider), an ad hoc network, an intranet, the Internet, a fiber optic-based network, a cloud computing network, and / or the like, and / or a combination of these or other types of networks.
[0120] The number and arrangement of systems, devices, and / or networks shown in FIG. 3 are provided as an example. There may be additional systems, devices, and / or networks; fewer systems, devices, and / or networks; different systems, devices, and / or networks; and / or differently arranged systems, devices, and / or networks than those shown in FIG. 3. Furthermore, two or more systems or devices shown in FIG. 3 may be implemented within a single system or device, or a single system or device shown in FIG. 3 may be implemented as multiple, distributed systems or devices. Additionally or alternatively, a set of systems (e.g., one or more systems) or a set of devices (e.g., one or more devices) of environment 300 may perform one or more functions described as being performed by another set of systems or another set of devices of environment 300.
[0121] Referring now to FIG. 4, shown is a diagram of example components of a device 400 according to non-limiting embodiments. Device 400 may correspond to machine learning model system 102, database 104, and / or user device 106 of FIG. 1 and / or transaction service provider system 302, issuer system 304, customer device 306, merchant system 308, acquirer system 310, and communication network 312 of FIG. 3, as an example. In some non-limiting embodiments, such systems or devices may include at least one device 400 and / or at least one component of device 400. The number and arrangement of components shown are provided as an example. In some non-limiting embodiments, device 400 may include additional components, fewer components, different components, or differently arranged components than those shown. Additionally or alternatively, a set of components (e.g., one or more components) of device 400 may perform one or more functions described as being performed by another set of components of device 400.
[0122] As shown in FIG. 4, device 400 may include a bus 402, a processor 404, memory 406, a storage component 408, an input component 410, an output component 412, and a communication interface 414. Bus 402 may include a component that permits communication among the components of device 400. In some non-limiting embodiments, processor 404 may be implemented in hardware, a combination of hardware and firmware, and / or a combination of hardware and software. For example, processor 404 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component (e.g., a field-programmable gate array (FPGA), an applicationspecific integrated circuit (ASIC), etc.) that can be programmed to perform a function. Memory 406 may include random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, optical memory, etc.) that stores information and / or instructions for use by processor 404.
[0123] With continued reference to FIG. 4, storage component 408 may store information and / or software related to the operation and use of device 400. For example, storage component 408 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, a solid-state disk, etc.) and / or another type of computer-readable medium. Input component 410 may include a component that permits device 400 to receive information, such as via user input (e.g., a touch screendisplay, a keyboard, a keypad, a mouse, a button, a switch, a microphone, etc.). Additionally or alternatively, input component 410 may include a sensor for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). Output component 412 may include a component that provides output information from device 400 (e.g., a display, a speaker, one or more light-emitting diodes (LEDs), etc.). Communication interface 414 may include a transceiver-like component (e.g., a transceiver, a separate receiver and transmitter, etc.) that enables device 400 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 414 may permit device 400 to receive information from another device and / or provide information to another device. For example, communication interface 414 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a cellular network interface, and / or the like.
[0124] Device 400 may perform one or more processes described herein. Device 400 may perform these processes based on processor 404 executing software instructions stored by a computer-readable medium, such as memory 406 and / or storage component 408. A computer-readable medium may include any non- transitory memory device. A memory device includes memory space located inside of a single physical storage device or memory space spread across multiple physical storage devices. Software instructions may be read into memory 406 and / or storage component 408 from another computer-readable medium or from another device via communication interface 414. When executed, software instructions stored in memory 406 and / or storage component 408 may cause processor 404 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, embodiments described herein are not limited to any specific combination of hardware circuitry and software. The term “programmed or configured,” as used herein, refers to an arrangement of software, hardware circuitry, or any combination thereof on one or more devices.
[0125] Referring now to FIG. 5A, shown is an example implementation 500a of a method for classification based on multiple imbalanced classes of data (e.g., an implementation of method 200 of FIG. 2), according to some non-limiting embodimentsor aspects. The steps shown in FIG. 5A are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of implementation 500a may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of implementation 500a may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106.
[0126] In some non-limiting embodiments or aspects, implementation 500a may include a plurality of data samples 502 (e.g., machine learning model system 102 may receive a plurality of data samples 502). For example, the data samples 502 may be associated with a plurality of classes (e.g., M classes, denoted ci, C2, ... , CM). A first class ci may be the smallest class (e.g., the class with the fewest data samples 502). A plurality of other classes (e.g., C2 through CM) each may have a respective number of data samples. For example, each class may have a class proportion (e.g., the proportion of the total number of data samples 502, which may be determined by dividing the number of data samples 502 in the class by the total number of data samples 502). For example, the class proportions of each class ci through CM may be denoted as qi through qM. Assuming the classes are in order (or if the classes are sorted to be in order, e.g., by machine learning model system 102) of increasing class proportion (e.g., increasing number of data samples per class), then qi< q2 ^ ... ^ qM. A respective class proportion ratio R may be determined for each class C2 through CM based on rounding (e.g., to the nearest integer) a ratio between the respective class proportion and the first class proportion qi. For example, the class proportion ratio R may be determined for each class based on the following equation: Equation 1R = [1 , round(q2 / qi), round(q3 / qi),..., round(qM / qi)]
[0127] In some non-limiting embodiments or aspects, all data samples 502 of the first class ci may be included (e.g., by machine learning model system 102) in a single bin 504 (e.g., bin 1 of class ci). In some non-limiting embodiments or aspects, thedata samples 502 of each other class may be divided (e.g., by machine learning model system 102) into a plurality of bins 504 of roughly equal size. For example, the number of bins nmof the m-th class (e.g., m = 1 , 2, ..., M) may be based on the respective class proportion ratio (e.g., nm= Rm = round(qm / qi)). As such, a second class C2 may have n2 bins 504, a third class C3 may have ns bins 504, and so on through an M-th class CM having OM bins.
[0128] Referring now to FIG. 5B, shown is an example implementation 500b of a method for classification based on multiple imbalanced classes of data (e.g., an implementation of method 200 of FIG. 2), according to some non-limiting embodiments or aspects. The steps shown in FIG. 5B are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of implementation 500b may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of implementation 500b may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106. In some non-limiting embodiments or aspects, implementation 500b may be combined with implementation 500a.
[0129] In some non-limiting embodiments or aspects, a plurality of data blocks 506 may be generated (e.g., by machine learning model system 102). For example, each data block 506 may include bin 504 from the first class ci and a selected bin 504 from the plurality of bins 504 associated with each other class (e.g., C2 through CM). For example, generating the plurality of data blocks may include cyclically iterating through bins 504 of each other class (e.g., C2 through CM).
[0130] In some non-limiting embodiments or aspects, cyclically iterating through bins 504 of each other class (e.g., C2 through CM) may include, for a first data block 506 (e.g., data block 1 ), selecting a first bin (e.g., bin 1 ) of the bins 504 associated with each other class (e.g., C2 through CM). For each successive data block 506 after the first data block 506, a successive bin 504 of the bins 504 associated with eachother class (e.g., 02 through CM) may be selected until a last bin associated with each second class is selected. For the purpose of illustration, a second data block 506 (e.g., data block 2) may include bin 2 from each class C2 through CM. Data block 02 may include bin n2 from each class C2 through CM, and since bin n2 is the last bin for class C2, the next data block n2+1 may include bin 1 from class C2 (e.g., and bin n2+1 from every other class C3 through CM). In some non-limiting embodiments or aspects, this cyclical iteration may repeat for each class until each bin 504 from each class is included in at least one data block 506 (e.g., until bin HM of class CM is included in the last data block 506, which may be denoted as data block OM).
[0131] Referring now to FIG. 5C, shown is an example implementation 500c of a method for classification based on multiple imbalanced classes of data (e.g., an implementation of method 200 of FIG. 2), according to some non-limiting embodiments or aspects. The steps shown in FIG. 5C are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of implementation 500c may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of implementation 500c may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106. In some non-limiting embodiments or aspects, implementation 500c may be combined with implementation 500a, implementation 500b, or any combination thereof.
[0132] In some non-limiting embodiments or aspects, a plurality of classifier models 508 may be trained (e.g., by machine learning model system 102) based on the data blocks 506. For example, each respective classifier model 508 may be trained based on a respective data block 506 (e.g., classifier 1 may be trained based on data block 1 , classifier 2 may be trained based on data block 2, and so on through classifier OM being trained based on data block OM).
[0133] In some non-limiting embodiments or aspects, a target data sample 510 may be received (e.g., by machine learning model system 102). The target datasample 510 may be inputted (e.g., by machine learning model system 102) to the classifier models 508 to generate a plurality of outputs 512 (e.g., intermediate outputs, which may include classifier output 1 from classifier 1 , classifier output 2 from classifier 2, and so on through classifier output OM from classifier OM).
[0134] In some non-limiting embodiments or aspects, a classification 514 (e.g., a final classification) of target data sample 510 may be determined (e.g., by machine learning model system 102) based on the outputs 512. For example, the classification 514 may be determined based on an average of outputs 512, based on majority voting of the outputs 512, and / or the like.
[0135] Referring now to FIG. 5D, shown is an example implementation 500d of a method for classification based on multiple imbalanced classes of data (e.g., an implementation of method 200 of FIG. 2), according to some non-limiting embodiments or aspects. The steps shown in FIG. 5D are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of implementation 500d may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of implementation 500d may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106. In some non-limiting embodiments or aspects, implementation 500d may be combined with implementation 500a, implementation 500b, implementation 500c, or any combination thereof.
[0136] In some non-limiting embodiments or aspects, at least one unsupervised outlier scoring model 516 may be trained (e.g., by machine learning system 102) based on bin 1 of class m. For each other class (C2 through CM), an outlier score for each data sample of unsorted data samples 502a of the respective class may be generated (e.g., by machine learning system 102) based on the unsupervised outlier scoring model(s) 516. The respective data samples of each class may be sorted (e.g., by machine learning system 102) based on the outlier score for each data sample to provide sorted data samples 502b.
[0137] In some non-limiting embodiments or aspects, dividing the data samples into bins 504 may include dividing the sorted data samples 502b into the plurality of bins 504. For example, sorting the data samples may include ordering the data samples from a lowest outlier score to a highest outlier score (e.g., sorted data samples 502b may be ordered from lowest to highest outlier score). As such, dividing sorted data samples 502b into the plurality of bins 504 may include assigning each data sample of sorted data samples 502b to a bin 504 based on the order of the data samples (e.g., bin 1 includes the data samples with the lowest outlier scores, bin 2 includes the data samples with the second lowest outlier scores, and so on through bin nmincluding the data samples with the highest outlier scores).
[0138] Referring now to FIG. 5E, shown is an example implementation 500e of a method for classification based on multiple imbalanced classes of data (e.g., an implementation of method 200 of FIG. 2), according to some non-limiting embodiments or aspects. The steps shown in FIG. 5E are for example purposes only. It will be appreciated that additional, fewer, different, and / or a different order of steps may be used in non-limiting embodiments or aspects. In some non-limiting embodiments or aspects, a step may be automatically performed in response to performance and / or completion of a prior step. In some non-limiting embodiments or aspects, one or more of the steps of implementation 500e may be performed (e.g., completely, partially, and / or the like) by machine learning model system 102 (e.g., one or more devices of machine learning model system 102). In some non-limiting embodiments or aspects, one or more of the steps of implementation 500e may be performed (e.g., completely, partially, and / or the like) by another system, another device, another group of systems, or another group of devices, separate from or including machine learning model system 102, such as database 104 and / or user device 106. In some non-limiting embodiments or aspects, implementation 500e may be combined with implementation 500a, implementation 500b, implementation 500c, implementation 500d, or any combination thereof.
[0139] In some non-limiting embodiments or aspects, a plurality of unsupervised outlier scoring models 516 (e.g., T outlier scoring models, denoted outlier score model 1 , outlier score model 2, ..., outlier scoring model T) may be used. For example, the scoring, sorting, dividing, and data block generation, as described herein, may be repeated (e.g., by machine learning model system 102) for each unsupervised outlier scoring model 516. Additionally or alternatively, the training of the classifier modelsand generating of outputs (e.g., classifier outputs 512 and / or final output 514), as described herein, may be repeated (e.g., by machine learning model system 102) for each unsupervised outlier scoring model 516.
[0140] In some non-limiting embodiments or aspects, at least one metric 518 may be determined (e.g., by machine learning model system 102) for each unsupervised outlier scoring model 516. For example, the metric(s) may be determined for each respective unsupervised outlier scoring model 516 based on the respective classification (e.g., final classification 514) for each respective unsupervised outlier scoring model 516. In some non-limiting embodiments or aspects, a selected (e.g., best) unsupervised outlier scoring model 516 may be selected (e.g., by machine learning model system 102) based on the metric(s) 518 for each respective unsupervised outlier scoring model 516.
[0141] Although embodiments have been described in detail for the purpose of illustration, it is to be understood that such detail is solely for that purpose and that the disclosure is not limited to the disclosed embodiments or aspects, but, on the contrary, is intended to cover modifications and equivalent arrangements that are within the spirit and scope of the appended claims. For example, it is to be understood that the present disclosure contemplates that, to the extent possible, one or more features of any embodiment or aspect can be combined with one or more features of any other embodiment or aspect.
Claims
WHAT IS CLAIMED IS:1 . A computer-implemented method, comprising: receiving, with at least one processor, a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and a respective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, dividing, with at least one processor, the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generating, with at least one processor, a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; training, with at least one processor, a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks; receiving, with at least one processor, a target data sample; inputting, with at least one processor, the target data sample to the plurality of classifier models to generate a plurality of outputs; and determining, with at least one processor, a classification of the target data sample based on the plurality of outputs.
2. The method of claim 1 , wherein dividing the respective second number of data samples into a plurality of bins comprises: determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples; determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples; anddetermining a respective class proportion ratio for each second class based on rounding a ratio between the respective class proportion and the first class proportion, wherein a respective number of bins of the plurality of bins for each second class of the plurality of second classes is based on the respective class proportion ratio.
3. The method of claim 1 , wherein generating the plurality of data blocks comprises cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
4. The method of claim 3, wherein cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes comprises: for a first data block of the plurality of data blocks, selecting a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
5. The method of claim 1 , further comprising: training at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class; and for each second class of the plurality of second classes: generating an outlier score for each data sample of the respective second number of data samples based on the at least one unsupervised outlier scoring model; and sorting the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples, wherein dividing the respective second number of data samples into the plurality of bins comprises dividing the sorted data samples into the plurality of bins.
6. The method of claim 5, wherein sorting the respective second number of data samples comprises ordering the respective second number of data samples from a lowest outlier score to a highest outlier score, and wherein dividing the sorted data samples into the plurality of bins comprises assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
7. The method of claim 5, wherein the at least one unsupervised outlier scoring model comprises a plurality of outlier scoring models, wherein training the at least one unsupervised outlier scoring model comprises training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class, wherein generating the outlier score for each data sample of the respective second number of data samples comprises generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models, wherein sorting the respective second number of data samples comprises, for each respective unsupervised outlier scoring model, sorting the respective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples, and wherein dividing the respective second number of data samples into the plurality of bins comprises, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
8. The method of claim 7, wherein generating the plurality of data blocks comprises generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model, and wherein training the plurality of classifier models comprises training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
9. The method of claim 8, wherein inputting the target data sample comprises inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model, and wherein determining the classification comprises determining a respective classification for each respective unsupervised outlier scoring model.
10. The method of claim 9, further comprising: determining, with at least one processor, at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model; and selecting, with at least one processor, a selected unsupervised outlier scoring model from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.1 1. A system comprising: at least one processor configured to: receive a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and a respective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generate a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; train a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks;receive a target data sample; input the target data sample to the plurality of classifier models to generate a plurality of outputs; and determine a classification of the target data sample based on the plurality of outputs.
12. The system of claim 11 , wherein dividing the respective second number of data samples into a plurality of bins comprises: determining a first class proportion based on the first number of data samples and a total number of data samples of the plurality of data samples; determining a respective class proportion based on the respective second number of data samples of each second class and the total number of data samples of the plurality of data samples; and determining a respective class proportion ratio for each second class based on rounding a ratio between the respective class proportion and the first class proportion, wherein a respective number of bins of the plurality of bins for each second class of the plurality of second classes is based on the respective class proportion ratio.
13. The system of claim 1 1 , wherein generating the plurality of data blocks comprises cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes.
14. The system of claim 13, wherein cyclically iterating through bins of the plurality of bins associated with each second class of the plurality of second classes comprises: for a first data block of the plurality of data blocks, selecting a first bin of the plurality of bins associated with each second class of the plurality of second classes as the selected bin of the plurality of bins associated with each second class of the plurality of second classes; and for each successive data block after the first data block, selecting a successive bin of the plurality of bins associated with each second class until a last bin of the plurality of bins associated with each second class is selected.
15. The system of claim 1 1 , wherein the at least one processor is further configured to: train at least one unsupervised outlier scoring model based on the first number of data samples associated with the first class; and for each second class of the plurality of second classes: generate an outlier score for each data sample of the respective second number of data samples; and sort the respective second number of data samples based on the outlier score for each data sample to provide sorted data samples, wherein dividing the respective second number of data samples into the plurality of bins comprises dividing the sorted data samples into the plurality of bins.
16. The system of claim 15, wherein sorting the respective second number of data samples comprises ordering the respective second number of data samples from a lowest outlier score to a highest outlier score, and wherein dividing the sorted data samples into the plurality of bins comprises assigning each data sample of the sorted data samples to a bin of the plurality of bins based on ordering the respective second number of data samples from a lowest outlier score to a highest outlier score.
17. The system of claim 15, wherein the at least one unsupervised outlier scoring model comprises a plurality of outlier scoring models, wherein training the at least one unsupervised outlier scoring model comprises training the plurality of unsupervised outlier scoring models based on the first number of data samples associated with the first class, wherein generating the outlier score for each data sample of the respective second number of data samples comprises generating respective outlier scores for each data sample of the respective second number of data samples based on each respective unsupervised outlier scoring model of the plurality of unsupervised outlier scoring models, wherein sorting the respective second number of data samples comprises, for each respective unsupervised outlier scoring model, sorting therespective second number of data samples based on the respective outlier score for each data sample to provide respective sorted data samples, and wherein dividing the respective second number of data samples into the plurality of bins comprises, for each respective unsupervised outlier scoring model, dividing the respective second number of data samples into a respective plurality of bins based on the respective sorted data samples.
18. The system of claim 17, wherein generating the plurality of data blocks comprises generating a respective plurality of data blocks based on the respective plurality of bins for each respective unsupervised outlier scoring model, and wherein training the plurality of classifier models comprises training the plurality of classifier models based on the respective plurality of data blocks for each respective unsupervised outlier scoring model.
19. The method of claim 18, wherein inputting the target data sample comprises inputting the target data sample to the plurality of classifier models as trained based on the respective plurality of data blocks for each respective unsupervised outlier scoring model, and wherein determining the classification comprises determining a respective classification for each respective unsupervised outlier scoring model, wherein the at least one processor is further configured to: determine at least one metric for each respective unsupervised outlier scoring model based on the respective classification for each respective unsupervised outlier scoring model; and select a selected unsupervised outlier scoring model from the plurality of unsupervised outlier scoring models based on the at least one metric for each respective unsupervised outlier scoring model.
20. A computer program product comprising at least one non- transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to: receive a plurality of data samples associated with a plurality of classes comprising a first class and a plurality of second classes, wherein a first number of data samples of the plurality of data samples are associated with the first class and arespective second number of data samples is associated with each second class of the plurality of second classes; for each second class of the plurality of second classes, divide the respective second number of data samples into a plurality of bins based on a ratio based on the respective second number of data samples and the first number of data samples; generate a plurality of data blocks, each data block of the plurality of data blocks comprising the first number of data samples and a selected bin of the plurality of bins associated with each second class of the plurality of second classes; train a plurality of classifier models based on the plurality of data blocks, wherein each respective classifier model of the plurality of classifier models is trained based on a respective data block of the plurality of data blocks; receive a target data sample; input the target data sample to the plurality of classifier models to generate a plurality of outputs; and determine a classification of the target data sample based on the plurality of outputs.
Citation Information
Patent Citations
Method, System, and Computer Program Product for Detecting Fraudulent Interactions
US20210049418A1