Bayesian estimators and expectation maximization for classifier testing from noisy labels
The method evaluates classifier performance using noisy-label models to determine detection and false alarm probabilities, enhancing classifier deployment and improving performance in classification tasks.
Patent Information
- Application Number
- PCT/US2024/046762
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-09-13
- Publication Date
- 2025-12-26
AI Technical Summary
Classifier models trained with noisy labels perform less accurately than those trained with correct labels due to labeling errors, particularly in environments where gold-standard truths are absent, leading to improper deployment and suboptimal selection.
A method and system for evaluating classifier performance using a noisy-label model to determine the probability of detection and false alarm, allowing for improved evaluation and deployment decision-making based on these probabilities.
Enhances the deployment of classifiers by selecting models better suited for specific environments, improving performance in classification tasks by 10 times compared to conventional methods, and reducing testing errors significantly.
Smart Images

Figure US2024046762_26122025_PF_FP_ABST
Abstract
Description
[0001] BAYESIAN ESTIMATORS AND EXPECTATION MAXIMIZATION FOR CLASSIFIER TESTING FROM NOISY LABELS
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application Serial No. 63 / 600,022, filed November 16, 2023, under attorney docket number M0437.70164US00, and entitled “BAYESIAN ESTIMATORS FOR CLASSIFIER TESTING FROM NOISY LABELS”, which is hereby incorporated herein by reference in its entirety.
[0004] FEDERALLY SPONSORED RESEARCH
[0005] This invention was made with government support under FA8702-15-D-0001 awarded by the U.S. Air Force. The government has certain rights in the invention.
[0006] BACKGROUND
[0007] Classifier models may be trained to apply labels to input data. Classifier models may be trained with noisy labels.
[0008] SUMMARY
[0009] According to aspects of the disclosure, there is provided a method of evaluating performance of a trained classifier model, the method comprising providing a set of predicted labels output by the trained classifier model, providing a set of noisy labels, providing a noisy- label model having a set of parameters, based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model, and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0010] According to aspects of the disclosure, there is provided at least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by at least one processor, cause the at least one processor to perform a method of a method of evaluating performance of a trained classifier model, the method comprising providing a set of predicted labels output by the trained classifier model, providing a set of noisy labels, providing a noisy-label model having a set of parameters, based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model, and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0011] According to aspects of the disclosure, there is provided a system for evaluating performance of a trained classifier model, the system comprising at least one processor and at least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by the at least one processor, cause the at least one processor to perform a method comprising providing a set of predicted labels output by the trained classifier model, providing a set of noisy labels, providing a noisy-label model having a set of parameters, based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model, and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0012] In some embodiments, the method further comprises providing an additional set of predicted labels output by the trained classifier model and an additional set of correct labels and, based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises, based on the set of predicted labels of the trained classifier model, the set of noisy labels, the set of parameters of the noisy-label model, and the additional set of predicted labels of the trained classifier model and the additional set of correct labels, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model. In some embodiments, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises determining the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model and outputting the performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises outputting the performance evaluation of the trained classifier model based on the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model.
[0013] In some embodiments, the method further comprises providing a set of scores of the trained classifier model, each score of the set of score providing an indication of probability of truth for a predicted of the set of predicted labels of the trained classifier model and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises based on the set of scores of the trained classifier model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0014] In some embodiments, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises determining a first mean of a representation of an average number of true positives for the trained classifier model, determining a first variance of the representation of the average number of true positives for the trained classifier model, determining a second mean of a representation of an average number of false negatives for the trained classifier model, and determining a second variance of the representation of the average number of false negatives for the trained classifier model.
[0015] In some embodiments, the method further comprises based on the first mean, first variance, second mean, and second variance, determining the representation of the probability of detection of the trained classifier model, determining the representation of the probability of false alarm of the trained classifier model, redetermining the first mean, first variance, second mean, and second variance, and based on the redetermined first mean, first variance, second mean, and second variance, iterating the representation of the probability of detection of the trained classifier model, and iterating the representation of the probability of false alarm of the trained classifier model.
[0016] In some embodiments, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises simulating M possible combinations of correct labels based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model and, based on the simulated M possible combinations of the correct labels iterating the representation of the probability of detection of the trained classifier model and iterating the representation of the probability of false alarm of the trained classifier model.
[0017] In some embodiments, the trained classifier model comprises a multi-class trained classifier model and the method further comprises determining, for each possible value of a correct label, a representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label and outputting a performance evaluation of the trained classifier model based on, for each possible value of a correct label, the representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label.
[0018] In some embodiments, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization.
[0019] In some embodiments, estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization comprises calculating an expectation of a likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model and calculating a maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model. In some embodiments, estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization further comprises repeatedly performing calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model and subsequent to calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model, calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
[0020] In some embodiments, calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises calculating posterior probabilities of correct labels being equal to a first label and calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises determining a representation of a sum of posterior probabilities of events where predicted labels are equal to a second label and correct labels are equal to the first label divided by a sum of posterior probabilities of the events where correct labels are equal to the first label.
[0021] In some embodiments, the method further comprises determining whether to deploy or undeploy the trained classifier model based on the performance evaluation of the trained classifier model.
[0022] In some embodiments, the method further comprises deploying the trained classifier model to a classification environment.
[0023] In some embodiments, the method further comprises receiving input data for the trained classifier model and using the trained classifier model, applying one or more labels to the input data.
[0024] BRIEF DESCRIPTION OF DRAWINGS
[0025] Various aspects and embodiments of the application will be described with reference to the following figures. It should be appreciated that the figures are not necessarily drawn to scale. Items appearing in multiple figures may be indicated by the same reference number in all the figures in which they appear. FIGs. 1-1 and 1-2 show a comparison of supervised classification and estimation theory.
[0026] FIG. 2 shows a graphical model for a testing approach.
[0027] FIG. 3 shows a graphical model of iterative estimation for testing.
[0028] FIG. 4 shows an exemplary embodiment of multi-class classification.
[0029] FIG. 5 shows an exemplary embodiment of a binary symmetric broadcast channel.
[0030] FIGs. 6 and 7 shows exemplary graphs of mutual information.
[0031] FIG. 8 shows an exemplary graph of number of labelers and mutual information.
[0032] FIG. 9 shows an exemplary graph of a main testing example.
[0033] FIG. 9A shows an exemplary graph of a main testing example for iterative estimation methods.
[0034] FIGs. 10 shows an exemplary graph of a main testing example.
[0035] FIGs. 11-1, 11-2, 11-3, and 11-4 show exemplary graphs of a main testing example.
[0036] FIGs. 12-1 and 12-2 show exemplary graphs of a main testing example.
[0037] FIGs. 13-1 and 13-2 show exemplary graphs of estimates of metric random variables.
[0038] FIGs. 14-1 and 14-2 show exemplary graphs of scalar metric estimation errors.
[0039] FIGs. 14A-1 and 14A-2 show exemplary graphs of scalar metric estimation errors.
[0040] FIGs. 15-1, 15-2, 15-3, 15-4, 15-5, and 15-6 show exemplary graphs of P-R analysis estimation errors.
[0041] FIGs. 15A-1, 15A-2, 15A-3, 15A-4, 15A-5, and 15A-6 show exemplary graphs of P-R analysis estimation errors.
[0042] FIGs. 16-1, 16-2, 16-3, 16-4, 16-5, and 16-6 show exemplary graphs of ROC analysis estimation errors.
[0043] FIGs. 16A-1, 16A-2, 16A-3, 16A-4, 16A-5, and 16A-6 show exemplary graphs of ROC analysis estimation errors.
[0044] FIGs. 17-1 and 17-2 show exemplary graphs of a testing example for a single labeler.
[0045] FIGs. 18-1 and 18-2 show exemplary graphs of a testing example for small sample size.
[0046] FIGs. 19-1, 19-2, 19-3, and 19-4 show exemplary graphs of a training example with different numbers of labelers.
[0047] FIGs. 20-1, 20-2, 20-3, and 20-4 show exemplary graphs of a training example with different numbers of labelers.
[0048] FIGs. 21-1, 21-2, 21-3, and 21-4 show exemplary graphs of a training example with different numbers of labelers. FIGs. 22-1, 22-2, 22-3, and 22-4 show exemplary graphs of a training example with different numbers of labelers for ML-trained classifiers.
[0049] FIGs. 23-1, 23-2, 23-3, and 23-4 show exemplary graphs of a training example with different numbers of labelers for MMSE-trained classifiers.
[0050] FIGs. 24-1, 24-2, 24-3, 24-4, 24-5, and 24-6 show exemplary graphs of a ROC analysis for a single good labeler and multiple mediocre labelers.
[0051] FIGs. 25-1, 25-2, 25-3, and 25-4 show exemplary graphs of estimated P-R curve posteriors for a single expert labeler and many poor labelers.
[0052] FIG. 26 is a block diagram of a computer system on which various functions can be implemented, according to one exemplary embodiment.
[0053] FIG. 27 shows a process flow for a method of evaluating performance of a trained classifier model.
[0054] FIG. 28 shows a first algorithm related to evaluating performance of a trained classifier model.
[0055] FIG. 29 shows a second algorithm related to evaluating performance of a trained classifier model.
[0056] FIG. 30 shows a third algorithm related to evaluating performance of a trained classifier model.
[0057] FIG. 31 shows a fourth algorithm related to evaluating performance of a trained classifier model.
[0058] FIG. 32 shows a fifth algorithm related to evaluating performance of a trained classifier model.
[0059] FIG. 33 shows a sixth algorithm related to evaluating performance of a trained classifier model.
[0060] FIGs. 34-1 and 34-2 illustrate exemplary cases of ideal and noisy labels used in training classifier models.
[0061] FIG. 34-3 shows a training classifier model labeling an input.
[0062] DETAILED DESCRIPTION
[0063] The present disclosure provides methods and systems for evaluating performance of classifier models trained on noisy labels. The algorithms described herein provide for improved evaluation of performance of classifiers compared to conventional techniques. By providing improved evaluation of classifiers, performance of deployed classifiers in classification environments may be improved, because classifiers better suited for particular environments may be selected for deployment. In contrast, conventional techniques for evaluating classifiers trained with noisy labels may not take the noisy labels into account, and therefore, may result in improper or suboptimal selection of classifiers for deployment.
[0064] I. Overview
[0065] Based on a set of predicted labels for a trained classifier model, a set of noisy labels, and a set of parameters of a noisy-label model, a representation of a probability of detection and / or a representation of a probability of false alarm are determined. An additional set of predicted labels and / or an additional set of correct labels may also be used to determine the representations of the probabilities. A performance evaluation of the trained classifier model is output based on the at least one of the representation of the probability of detection and / or the representation of the probability of false alarm. Then, it may be determined whether to deploy or undeploy the trained classifier model based on the performance evaluation. Further, the trained classifier model may be deployed to a classification environment, where input data for the trained classifier model is received and one or more labels are applied to the input data using the trained classifier model.
[0066] Methods provided herein may be used to evaluate the performance of classifier models, such as classifier models trained on noisy labels. While ideal supervised classification may assume known correct labels, the inventor has recognized and appreciated that various truthing issues arise in practice, particularly with respect to noisy labels. When a classifier or other model is trained, labels may be applied with training data by a labeler, such as a human labeler. When labels applied to training data are not the correct labels, they may be referred to as noisy labels. Models trained on said noisy labels may perform less accurately than models trained on correct labels.
[0067] For example, FIGs. 34-1 and 34-2 illustrate cases of ideal and noisy labels, while FIG. 34-3 shows a training classifier model labeling an input. FIG. 34-1 shows an ideal case with ideal labels. In FIG. 34-1, a human labeler has correctly provided labels of 1 or 0 responsive to the feature vector. Such data may be used to generate a trained classifier model which may be used to label feature vectors, as shown in FIG. 34-3. However, training data for classifier models is often not ideal, and instead exhibits many real world problems. Humans may make labeling errors, in other words, they may label a feature as present when it is absent or vice versa. The noise of labels may differ by case, as some labeling cases may be easy and some may be hard. Issues that may arise with labeling for training classifier models include: noise errors in labeling, the presence of multiple labels per sample (e.g., when multiple human labelers provide different labels for the same feature), adversarial labelers (e.g., humans who purposely choose the wrong label, or insert incorrect training data), as well as issues that may arise when training data is itself generated by a trained classifier model. Illustrative examples of environments in which noisy labels arise include, among other things: medical environments such as radiologist assessments of images, where there is no gold-standard test; and security or defense settings where analysts label operational data but there is no definitive truth. As such, the inventor has recognized that there are many problems that involve noisy labels, and that these noisy labels degrade ability to train and test classifiers due to the presence of multiple labels per sample (e.g., due to crowd sourcing).
[0068] Accordingly, the inventor has recognized and appreciated that a method of evaluating the performance of classifier models trained on noisy labels would enhance the deployment of these models. For example, when the performance of one or more classifier models trained on noisy labels is evaluated, a decision can be made as to whether to deploy a model to a particular environment, or whether to select a model best suited for a particular environment from a group of multiple models.
[0069] Trained classifier models may be deployed in a variety of environments for classifying input data. For example, such models have applicability to image recognition, target detection, sentiment analysis, digital marketing and spam detection, and may include a variety of algorithms, such as logistic regression, support vector machines, decision trees, random forests, convolutional neural networks, long short-term memory networks and transformers.
[0070] In some embodiments, there may be provided a noisy label model, which treats noisy labels as random variables conditioned on unobserved correct labels. Such a model may estimate the conditional distribution of the noisy labels and the class prior, and may estimate the correct labels or training with noisy labels.
[0071] Based on the evaluations of classifier models described herein, it may be determined whether to deploy or undeploy a trained classifier model, or to select for deployment a particular trained classifier model from a group of trained classifier models. Once a deployment decision has been made for a trained classifier model, the model may be deployed to a classification environment. In the classification environment, the model may receive input data and the trained classifier model may be used to apply one or more labels to the input data. After labels have been applied to the input data, further processing may be performed using the labeled data.
[0072] Accordingly, the disclosure provides systems and methods for improved training and testing of classifiers using only noisy labels. The improved training and testing of classifiers using only noisy labels is applicable to many real-world problems that involve noisy labels, which degrade the ability to train and test classifiers due to crowdsourced data having multiple labels per sample. Using estimation theory as described herein, the disclosure enables improved training and testing with noisy labels. The systems and methods described herein apply to virtually all classifiers and noisy-label models and may not require estimation of correct labels. Compared with conventional methods, the systems and methods described herein provide improved testing errors that may be about 10 times smaller than other, conventional methods. Indeed, the improved training and testing described herein is more comparable to using correct labels. Further, the disclosure provides systems and methods for performing multi-class classification, and provides for the effects of having multiple labels per sample. Using the training and testing systems and methods described herein, multiple mediocre labelers may be used to generate a trained classifier model that is as informative a model generated using a single expert.
[0073] Furthermore, according to aspects of the disclosure, trained classifier models may be improved after testing. For example, after determining the performance of a model using the systems and methods described herein, the weights or other operating parameters of a trained classifier model may be adjusted, after which the model may be tested again to determine an improvement. Several such improvement stages may be performed to determine an optimal or improved weighting or parameterized implementation of the model. Alternatively or additionally, a set of models with different weights or parameters may be provided as a set, and the testing systems and methods described herein may evaluate each model in the set to determine an improved or optimal model. In this manner, the testing systems and methods described herein may be used to provide improved or optimal trained classifier models.
[0074] Supervised classification is a broad category within artificial intelligence and machine learning that uses labeled data to train and test a predictive model, e.g., a classifier. In various embodiments, these trained classifiers may provide high performance in various applications, such as image classification, detecting diseases and tumors, forecasting weather, detecting street signs and obstacles during autonomous navigation, and performing sentiment analysis on tweets, which makes them applicable to almost every industry.
[0075] In various embodiments, classification may rely on having correctly-labeled data, which may require companies spend a vast amount of resources labeling data sets, for example, using the Amazon Mechanical Turk is a crowdsourcing service where one can hire people to label data sets.
[0076] Errors in labeling frequently occur, which results in noisy labels. In various embodiments, there are provided noisy-label models and methods of determining these models. Further, there are methods for training classifiers using noisy labels and such noisy-label models.
[0077] In various embodiments, trained classifier models may be tested to evaluate their performance. The inventor has recognized and appreciated that in testing trained classifiers, practitioners have continued to test the models as if the noisy labels were correct, resulting in inaccurate evaluation of classifier performance.
[0078] Aspects of the disclosure provide various algorithms for classifier testing using noisy labels. Methods described herein may be used with any classifier and virtually any noisy label model. Provided are two algorithms for binary, e.g., two-class, classification and a third algorithm for multi-class classification. The algorithms apply Bayesian estimation techniques to estimate testing metrics.
[0079] The algorithms provided herein provide enhanced evaluation of classifier models compared to conventional techniques. For example, the techniques described herein may produce estimates with errors about 10 times smaller than other methods, with results within 1- 5% of actual metrics, while conventional methods may be only within 10-40% of actual values. According to aspects of the disclosure, these algorithms may be used to test with noisy labels nearly or just as reliably as testing with correct labels, and these algorithms may employed in any classification application where noisy labels are encountered.
[0080] According to the disclosure, novel and generally applicable testing algorithms are provided. For typical testing methods, inputs may comprise pairs (e.g., predicted label and reference or correct label), where the predicted label is the output of the classifier and the reference label is calculated from the noisy label(s). The pairs may then be plugged into standard formulas for accuracy, probability of detection, precision, recall, or F-score. Described in further detail below, Algorithms 1 and 2 provide for binary classification, while Algorithm 4 provides for multi-class classification. As inputs, these algorithms take pairs (e.g., predicted label and noisy labels) for each sample, any noisy-label model with the form p(noisy labels | correct label) p(correct label) and accompanying parameters (e.g., Ψ and π herein), and an optional additional set of pairs (e.g., predicted label and correct label). For Algorithms 2 and 4, an integer M with the number of sample realizations to draw is another input. Rather than plugging into a formula, these Algorithms use an iterative procedure known as empirical Bayes to estimate parameters of the testing model.
[0081] After an algorithm completes, subsequent calculations (described in Subsection 2.5 of Section II) provide different kinds of estimates of the testing metrics. For Algorithms 1 and 2, the posterior probability distribution of one-dimensional metrics (accuracy, probability of detection, probability of false alarm, precision, recall, and / or F-score) or two-dimensional metrics ((probability of detection, probability of false alarm), (precision, recall)). For all algorithms, the conditional mean of the preceding metrics, which is an estimated value of the metric that has the smallest mean-square error, the most-probable value of the preceding metrics, and credible regions of the metrics, for example, for 1-D metrics, this may be an interval [a, b] such that there is a 95% probability that the metric lies between a and b. For 2-D metrics, this may be an arbitrarily shaped region.
[0082] Algorithms 1, 2, and 4 may be implemented according Section II below. For example, Section II includes a pseudocode implementation of these algorithms and exemplary uses of the algorithms. Furthermore, the person of skill in the art may readily envision modifications of the algorithms within the scope of the disclosure and by using the package or examining the pseudocode to create a modified implementation.
[0083] According to various embodiments, the algorithms described herein may be applied to and produce accurate results with any classifier (because the algorithms operate on the classifier outputs, it may not matter what kind of classifier was used) and any noisy-label model having the form p(noisy labels | correct label) p(correct label). Furthermore, implementations of the algorithms may provide a standard form for the predicted and noisy labels, and a manner for users to specify a noisy-label model and its parameters. Algorithms 2 and 4 use random sampling, which, according to some embodiments, may be parallelized. In various embodiments, portions of subsequent calculations may also be parallelized. Algorithms 1, 2, and 4 may be applied to testing environments in various ways. A software package, such as scikit-learn’s “classification metrics functions” may be used to provide the algorithms available to various users or deployers of classifier models. Furthermore, the algorithms may also be integrated into various machine-learning frameworks, such as TensorBoard or Microsoft Azure. In some embodiments, the algorithms may be available as a service: a user may could upload the predicted labels, noisy labels, noisy-label model parameters, and, in some embodiments, an additional set of predicted labels and correct labels, and the service may apply the algorithms and provide testing results.
[0084] Described in further detail below, Algorithm 5 provide for binary classification using expectation maximation, while Algorithm 6 provides for multi-class classification for the same. As inputs, these algorithms may take similar input as the MMSE algorithms: pairs (e.g., predicted label and noisy labels) for each sample, any noisy-label model with the form p(noisy labels | correct label) p(correct label) and accompanying parameters (e.g., Ψ and π herein), and an optional additional set of pairs (e.g., predicted label and correct label).
[0085] Algorithms 5 and 6 may be implemented according Section III below. For example, Section III includes a pseudocode implementation of these algorithms and exemplary uses of the algorithms. Furthermore, the person of skill in the art may readily envision modifications of the algorithms within the scope of the disclosure and by using the package or examining the pseudocode to create a modified implementation.
[0086] In some embodiments, the algorithms take as input a suitable noisy-label model. This model may be learned in addition to the main classifier. For example, a published noisy-label model and methods for learning them may be provided in a package alongside the algorithms. Accordingly, a user may use the package to learn the noisy-label model before providing it to the algorithms to evaluate the trained classifier.
[0087] In some embodiments, a package comprising the algorithms may provide functions for the algorithms and subsequent calculations, one or more noisy-label models and references on learning them, and documentation and examples of how to apply the algorithms to a set of predicted labels, noisy labels, a learned noisy-label model, and, in some embodiments, an additional set of predicted labels and correct labels. A user may then learn the noisy-label model outside of the package. The Python package described in the disclosure provides one exemplary embodiment of this capability. FIGs. 14-1 and 14-2, 15-1, 15-2, 15-3, 15-4, 15-5, and 15-6, and 16-1, 16-2, 16-3, 16-4, 16-5, and 16-6 illustrate comparisons of estimation errors of the techniques described herein (e.g., Ratios and Sampling) and conventional techniques. The axes for the techniques described herein are lOx smaller (-5% to +5%) than for the conventional techniques (-50% to +50%). For multi-class classification, Tables 13, 14, and 15 of Section II below show that the techniques described herein (e.g., MMSE testing) produces better estimates of a classifier’s accuracy and confusion matrix compared with conventional techniques.
[0088] In various embodiments, trained classifiers may provide high performance in various applications. Some exemplary embodiments of classification environments for classifiers, include image classification, medical diagnosis (such as detecting diseases and tumors), forecasting weather, autonomous driving (such as detecting street signs and obstacles during autonomous navigation), sentiment analysis, (such as performing sentiment analysis on tweets). The techniques described herein provide improved performance evaluation of classifier models in each of these environments. Because the algorithms described herein provide more accurate performance evaluation for classifiers tested with noisy labels, it may be easier to determine what is a better classifier to deploy in a particular classification environment. By selecting a better classifier for a particular classification environment, performance in that classification environment will be improved. Accordingly, the techniques described herein provide for improved image classification, improved medical diagnosis, improved forecasting of weather, improved autonomous driving, improved sentiment analysis, as well as in every other classification environment, even without directly improving the performance of a particular classifier.
[0089] In some embodiments, such as the exemplary Algorithm 1 and Algorithm 2 of Section II, methods may include steps of providing a set of predicted labels output by the trained classifier model, providing a set of noisy labels, providing a noisy-label model having a set of parameters, (and, in some embodiments, providing an additional set of predicted labels and correct labels), and based on the set of predicted labels of the trained classifier model, the set of noisy labels, the set of parameters of the noisy-label model, (and, in some embodiments, based further on the additional set of predicted labels and additional set of correct labels), determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model, and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0090] In some embodiments, systems described herein may evaluate performance of a classifier using predicted scores from a trained classifier model, either alternatively or in addition to using predicted labels of the model. For example, for binary classification, systems may provide a set of predicted scores from the trained classifier model instead of predicted labels, and may then sweep a decision threshold over the scores to generate precision-recall and receiver operating characteristic (ROC) curves and their probability density functions (PDFs). In some embodiments, users of such a system may use this process to generate performance evaluations of classifier models as in FIGs. 20-1, 20-2, 20-3, and 20-4, 22-1, 22-2, 22-3, and 22-4, 23-1, 23- 2, 23-3, and 23-4, 24-1, 24-2, 24-3, 24-4, 24-5, and 24-6, and 25-1, 25-2, 25-3, and 25-4.
[0091] In some embodiments, such as the exemplary Algorithm 1, methods may include steps of determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises determining a first mean of a representation of an average number of true positives for the trained classifier model, determining a first variance of the representation of the average number of true positives for the trained classifier model, determining a second mean of a representation of an average number of false negatives for the trained classifier model, and determining a second variance of the representation of the average number of false negatives for the trained classifier model. In some embodiments, methods may further include steps of, based on the first mean, first variance, second mean, and second variance, determining the representation of the probability of detection of the trained classifier model, determining the representation of the probability of false alarm of the trained classifier model, redetermining the first mean, first variance, second mean, and second variance, and based on the redetermined first mean, first variance, second mean, and second variance, iterating the representation of the probability of detection of the trained classifier model, and iterating the representation of the probability of false alarm of the trained classifier model.
[0092] In some embodiments, such as the exemplary Algorithm 2, methods may include steps of determining the at least one of the operating parameter representing the probability of detection of the trained classifier model or the operating parameter representing the probability of false alarm of the trained classifier model comprises simulating the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
[0093] In some embodiments, such as the exemplary Algorithms 5 and 6, methods may include steps of determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprising estimating the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model using expectation maximization.
[0094] In some embodiments, such as the exemplary Algorithms 5 and 6, methods may further include steps of estimating the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model using expectation maximization comprising calculating an expectation of the logarithm of the likelihood function from the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model and calculating a maximization of the expectation of the logarithm of the likelihood function from the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model.
[0095] In some embodiments, such as the exemplary Algorithms 5 and 6, methods may include steps of estimating the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model using expectation maximization further comprising repeatedly performing calculating the expectation of the logarithm of the likelihood function from the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model and subsequent to calculating the expectation of the logarithm of the likelihood function from the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model, calculating the maximization of the logarithm of the expectation of the likelihood function from the probability of detection of the trained classifier model and the probability of false alarm of the trained classifier model.
[0096] In some embodiments, such as the exemplary Algorithms 5 and 6, methods may include steps of calculating the expectation of the logarithm of the likelihood function of the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprising calculating posterior probabilities of correct labels being equal to a first label and calculating the maximization of the logarithm of the expectation of the likelihood function from the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprising determining a representation of a sum of posterior probabilities of events where predicted labels are equal to a second label and correct labels are equal to the first label divided by a sum of posterior probabilities of the events where correct labels are equal to the first label.
[0097] An illustrative implementation of a computer system 100 that may be used in connection with any of the embodiments of the disclosure provided herein is shown in FIG. 26. The computer system 100 may include one or more processors 110 and one or more articles of manufacture that comprise non-transitory computer-readable storage media, for example, memory 120 and one or more non-volatile storage media 130. The processor 110 may control writing data to and reading data from the memory 120 and the non-volatile storage device 130 in any suitable manner. To perform any of the functionality described herein, the processor 110 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media, for example the memory 120, which may serve as non- transitory computer-readable storage media storing processor-executable instructions for execution by the processor 110.
[0098] FIG. 27 shows a process flow 200 for a method of evaluating performance of a trained classifier model. Process flow 200 includes step 202, step 204, step 206, step 208, and step 210. At step 202, at least one processor may provide a set of predicted labels output by the trained classifier model. At step 204, at least one processor may provide a set of noisy labels. At step 206, at least one processor may provide a noisy-label model having a set of parameters. At step 208, at least one processor may based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model. At step 210, at least one processor may output a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model. In various embodiments, the process flow 200 may comprise additional steps, such as the steps of the algorithms shown and described with respect to FIGs. 28, 29, 30, 31, 32, and 33.
[0099] Section II below provides exemplary embodiments of truthing for supervised classification with minimum mean-square error (MMSE) testing using empirical Bayes algorithms. As described in Section II, ideal supervised classification may assume known correct labels. However, various truthing issues can arise in practice: noisy labels; multiple, conflicting labels for a sample; missing labels; and different labeler combinations for different samples.
[0100] There may be provided a noisy-label model, which views the observed noisy labels as random variables conditioned on the unobserved correct labels. Noisy-label models may be learned by estimating the conditional distribution of the noisy labels and a class prior, and they may be used to estimate the correct labels or to train a classifier with noisy labels. As described in Section II given a conditional distribution and class prior, systems and methods described herein may apply estimation theory to classifier testing, training, and comparison of different combinations of labelers. Systems and methods described herein provide, for binary classification, a testing model, and Section II provides approximate marginal posteriors for accuracy, precision, recall, probability of false alarm, and F-score, and joint posteriors for ROC and precision-recall analysis. The algorithms described in Section II may provide minimum mean-square error (MMSE) testing, which employs empirical Bayes algorithms to estimate the testing-model parameters and then computes optimal point estimates and credible regions for the metrics. Section II also provides an extended approach for multi-class classification to obtain optimal estimates of accuracy and individual confusion-matrix elements.
[0101] Additionally, Section II provides a unified view of training that covers probabilistic (i.e., discriminative or generative) and non-probabilistic models. For the former, maximum-likelihood or maximum a posteriori training may be adjusted for truthing issues. For the latter, MMSE training may be used which minimizes the MMSE estimate of the empirical risk. Section II also provides compatible suboptimal training.
[0102] Section II also provides that mutual information allows expression of any labeler combination as an equivalent single labeler, such that multiple mediocre labelers can be as informative as, or more informative than, a single expert labeler. Section II demonstrates the effectiveness of the methods and confirm the implication.
[0103] Section III below provides exemplary embodiments of truthing for supervised classification with expectation-maximization (EM) algorithms. Section III provides systems and methods for parameter estimation for classifier testing with noisy labels. Section III provides a specialized expectation-maximization (EM) algorithm as an alternative to the empirical Bayes estimation methods that use iterative minimum mean-square error (MMSE) estimation described in Section II. As described in Section III, systems and methods may use the EM algorithm for both binary and multi-class classifier testing with noisy labels. To compare the performance of the EM algorithm with the systems and methods of Section II, Section II also provides comparison experiments. According to various embodiments, the EM algorithm may perform similarly to the empirical Bayes methods of Section II.
[0104] As described above, Section II provides systems and method for classifier testing using noisy labels rather than correct labels. Section II provides testing models that relate the observed predicted and noisy labels and the unobserved correct-label random variable (RV). To estimate the parameters of the testing models from the predicted and noisy labels, systems and methods according to section II use empirical Bayes algorithms (Algorithms 1, 2, 4) and use minimum mean-square error (MMSE) estimation, and these methods may be referred to as MMSE testing. After the parameters are estimated, the posteriors of the metric RVs and optimal estimates of the metric RVs can be calculated.
[0105] Section III provides systems and methods using an expectation-maximization (EM) algorithm to estimate these parameters. The assumptions and notation in Section III may be the same as those in Section II. Equation (n) in Section II may be referenced in Section III as Equation (O-n). Similar exemplary experiments as in Section II are revisited in Section III. For easy cross-referencing with the similar figures or tables in Section II (e.g., “Figure 14” or “FIGs. 14-1 and 14-2”), any revisited figures of Section III may append ‘A’ to the same number (e.g., “Figure 14A” or “FIGs. 14A-1 and 14A-2”). Other tables, algorithms, appendices may continue the numbering sequence from Section II.
[0106] Subsection 2 of Section III describes latent-variable problems and reviews the EM algorithm. Subsection 3 shows how classifier testing with noisy labels can be viewed as a latent- variable problem. For binary classifier testing with noisy labels, Subsection 4 summarizes the EM algorithm, and Algorithm 5 provides pseudocode. Subsection 5 and Algorithm 6 address a multi-class case. Subsection 6 describes MMSE testing and the EM algorithm, and Subsection 7 revisits some of the relevant experiments from Section III to compare the EM algorithm with the existing methods. Section 8 concludes Section III. Derivations of the EM algorithm for binary and multi-class classifier testing with noisy labels appear in Appendices G and H. II. Truthing for Supervised Classification with Minimum Mean-Square Error (MMSE) Testing using Empirical Bayes Algorithms FIGs. 1-1 and 1-2 show a comparison of supervised classification and estimation theory. In supervised classification as illustrated by FIG. 1-1, a system or method may use a set of correctly-labeled samples from an unknown actual process to learn a predictive model that generalizes to future, out-of-sample realizations from the process. In estimation theory as illustrated by FIG. 1-2, a system or method may estimate the in-sample value of an unobserved variable from noisy measurements produced by a known measurement process. In each of FIGs. 1-1 and 1-2 there is an unknown element such as the unknown actual process of FIG. 1-1 or the observed actual of FIG. 1-2, a known element such as the correctly labeled samples of FIG. 1-1 or the known measurement process of FIG. 1-2, and a desired element, such as the learned predictive model of FIG. 1-1 or the estimated current value of FIG. 1-2.
[0107]
[0108]
[0109]
[0110] FIG. 2 shows a graphical model for a testing approach. Small rectangles indicate nonrandom variables and circles indicate random variables. Shading indicates a variable that is fully observed. Large rectangles indicate N independent instances of the enclosed variables indexed by i.
[0111] FIG. 3 shows a graphical model of iterative estimation for testing. Similar to FIG. 2, Small rectangles indicate non-random variables and circles indicate random variables, shading indicates a variable that is fully observed, and large rectangles indicate N independent instances of the enclosed variables indexed by i. The common random variables U and V may depend on separate partitions of F and
[0112] FIG. 28 shows a first algorithm related to evaluating performance of a trained classifier model. FIG. 29 shows a second algorithm related to evaluating performance of a trained classifier model.
[0113] FIG. 30 shows a third algorithm related to evaluating performance of a trained classifier model.
[0114]
[0115] FIG. 4 shows an exemplary embodiment of multi-class classification. FIG. 4 shows a graphical model of iterative estimation for testing. FIG. 31 shows a fourth algorithm related to evaluating performance of a trained classifier model.
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124] FIG. 5 shows an exemplary embodiment of a binary symmetric broadcast channel.
[0125]
[0126]
[0127] FIG. 8 shows an exemplary graph of the number of labelers used to achieve same mutual information as a single labeler for π(1) = 0.4.
[0128]
[0129]
[0130]
[0131] FIG. 9 shows an exemplary graph of a main testing example. FIG. 9 shows progression of in iterative estimation methods. FIG. 10 shows an exemplary graph of a main testing example. FIG. 10 shows accuracy estimates by all methods, with markers spaced vertically for easier readability. FIGs. 11-1, 11-2, 11-3, and 11-4 show exemplary graphs of a main testing example.
[0132] FIGs. 11-1, 11-2, 11-3, and 11-4 show estimates of scalar metric random variables.
[0133] FIGs. 12-1 and 12-2 show exemplary graphs of a main testing example. FIGs. 12-1 and 12-2 show estimates of joint metric random variables for ROC and P-R analysis.
[0134] FIGs. 13-1 and 13-2 show exemplary graphs of estimates of metric random variables if the predicted labels and OP parameters are not fully exploited. As may be appreciated by reviewing FIGs. 10, 12-1, and 12-2, the plots of FIGs. 13-1 and 13-2 do not show the labelers’ results because they are unchanged from the earlier figures.
[0135] FIGs. 14-1 and 14-2 show exemplary graphs of scalar metric estimation errors for different testing approaches over 100 different ideal operating points. The axis limits differ for MMSE testing with empirical Bayes (FIG. 14-1 : -0.05 to +0.05) and the other approaches (FIG. 14-2: -0.5 to +0.5). Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and as text above each circle. Multiples of ±1 and ±2 times the standard deviation of the errors appear as crosses, and text below the first cross gives the standard deviation.
[0136] FIGs. 15-1, 15-2, 15-3, 15-4, 15-5, and 15-6 show exemplary graphs of P-R analysis estimation errors for different testing approaches over 100 different ideal operating points. The axis limits differ for MMSE testing with empirical Bayes (FIGs. 15-1 and 15-2: -0.05 to +0.05) and the other approaches (FIGs. 15-3, 15-4, 15-5, and 15-6: -0.5 to +0.5). Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and listed as an ordered pair. Ellipses denote areas that account for 0.6827 and 0.9545 of the density of a bivariate normal distribution fitted to the errors, and text adjacent to the semi-major and semiminor axes gives the square roots of the eigenvalues of the error covariance matrix.
[0137] FIGs. 16-1, 16-2, 16-3, 16-4, 16-5, and 16-6 show exemplary graphs of ROC analysis estimation errors for different testing approaches over 100 different ideal operating points. The axis limits differ for MMSE testing with empirical Bayes (FIGs. 16-1 and 16-2: -0.05 to +0.05) and the other approaches (FIGs. 16-3, 16-4, 16-5, and 16-6: -0.5 to +0.5). Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and listed as an ordered pair. Ellipses denote areas that account for 0.6827 and 0.9545 of the density of a bivariate normal distribution fitted to the errors, and text adjacent to the semi-major and semiminor axes gives the square roots of the eigenvalues of the error covariance matrix. FIGs. 17-1 and 17-2 show exemplary graphs of a testing example for a single labeler (T = 1) when a constant labeling-error probability ε0= 0.1 is assumed for all samples. The mean and median of the labelers’ metrics are not shown because T= 1.
[0138] FIGs. 18-1 and 18-2 show exemplary graphs of a testing example for small sample size (N= 70),
[0139]
[0140] FIGs. 19-1, 19-2, 19-3, and 19-4 show exemplary graphs of a training example with different numbers of labelers T. FIGs. 19-1, 19-2, 19-3, and 19-4 show estimated testing results from MMSE testing with Algorithm 1 on the held-out testing set. A chance line appears as a black dotted line.
[0141] FIGs. 20-1, 20-2, 20-3, and 20-4 show exemplary graphs of a training example with different numbers of labelers T. FIGs. 20-1, 20-2, 20-3, and 20-4 show estimated ROC curves from MMSE testing with Algorithm 1 on the held-out testing set. A chance line appears as a black dotted line.
[0142] FIGs. 21-1, 21-2, 21-3, and 21-4 show exemplary graphs of a training example with different numbers of labelers T. FIGs. 21-1, 21-2, 21-3, and 21-4 show actual ROC curves on the held-out testing set. A chance line appears as a black dotted line.
[0143] FIGs. 22-1, 22-2, 22-3, and 22-4 show exemplary graphs of a training example with different numbers of labelers T for ML-trained classifiers. FIGs. 22-1, 22-2, 22-3, and 22-4 show estimated ROC curve posteriors from Algorithm 1 on the held-out testing set. The posteriors appear as heat maps with a base- 10 logarithmic color scale; the color bar ticks correspond to the exponents of the scale. A chance line appears as a black dotted line.
[0144] FIGs. 23-1, 23-2, 23-3, and 23-4 show exemplary graphs of a training example with different numbers of labelers T for MMSE-trained classifiers. FIGs. 23-1, 23-2, 23-3, and 23-4 show estimated P-R curve posteriors from Algorithm 1 on the held-out testing set. The posteriors appear as heat maps with a base- 10 logarithmic color scale; the color bar ticks correspond to the exponents of the scale. A chance line appears as a black dotted line.
[0145]
[0146] FIGs. 24-1, 24-2, 24-3, 24-4, 24-5, and 24-6 show exemplary graphs of a ROC analysis for a single good labeler (FIGs. 24-1, 24-3, and 24-5) and multiple mediocre labelers (FIGs. 24- 2, 24-4, and 24-6). FIGs. 24-1 and 24-2 show performance for the single threshold FIGs. 24-3 and 24-4 show the estimated ROC curve posteriors for ML-trained classifiers. FIGs. 24-5 and 24-6 show the estimated ROC curve posteriors for MMSE-trained classifiers. The posteriors appear as heat maps with a base- 10 logarithmic color scale; the color bar ticks correspond to the exponents of the scale. A chance line appears as a black dotted line.
[0147] FIGs. 25-1, 25-2, 25-3, and 25-4 show exemplary graphs of estimated P-R curve posteriors for a single expert labeler (FIGs. 25-1 and 25-3) and many poor labelers (FIGs. 25-2 and 25-4). FIGs. 25-1 and 25-2 show results for ML-trained classifiers; FIGs. 25-3 and 25-4 lower graphs show results for MMSE-trained classifiers. The posteriors appear as heat maps with a base- 10 logarithmic color scale; the color bar ticks correspond to the exponents of the scale. A chance line appears as a black dotted line.
[0148]
[0149] Appendix A, Metrics in Terms of Common RVs
[0150] Appendix B. Ratios of Jointly Normal Random Variables
[0151] Appendix C. Review of MMSE Estimation
[0152]
[0153]
[0154] Appendix E. Training with Truthing Issues
[0155]
[0156] Appendix F. Regularized Logistic Regression Training Equations
[0157]
[0158] III. Truthing for Supervised Classification with Expectation-Maximization (EM) Algorithms
[0159] FIG. 32 shows a fifth algorithm related to evaluating performance of a trained classifier model. FIG. 33 shows a sixth algorithm related to evaluating performance of a trained classifier model. FIG. 9A shows, similar to FIG. 9 described above, an exemplary graph of a main testing example. FIG. 9 A shows progression of in iterative estimation methods.
[0160] FIGs. 14A-1 and 14A-2 show, similar to FIGs. 14-1 and 14-2 described above, exemplary graphs of scalar metric estimation errors for different testing approaches over 100 different ideal operating points and 10 simulations per operating point. Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and as text above each circle. Multiples of ±1 and ±2 times the standard deviation of the errors appear as crosses, and text below the first cross gives the standard deviation.
[0161] FIGs. 15A-1, 15A-2, 15A-3, 15A-4, 15A-5, and 15A-6 show, similar to FIGs. 15-1, 15- 2, 15-3, 15-4, 15-5, and 15-6 described above, exemplary graphs of P-R analysis estimation errors for different testing approaches over 100 different ideal operating points and 10 simulations per operating point. Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and listed as an ordered pair. Ellipses denote areas that account for 0.6827 and 0.9545 of the density of a bivariate normal distribution fitted to the errors, and text adjacent to the semi-major and semi-minor axes gives the square roots of the eigenvalues of the error covariance matrix.
[0162] FIGs. 16A-1, 16A-2, 16A-3, 16A-4, 16A-5, and 16A-6 show, similar to FIGs. 16-1, 16- 2, 16-3, 16-4, 16-5, and 16-6 described above, exemplary graphs of ROC analysis estimation errors for different testing approaches over 100 different ideal operating points and 10 simulations per operating point. Miniature scatterplots of the estimation errors appear as dots. The average error is marked with a circle and listed as an ordered pair. Ellipses denote areas that account for 0.6827 and 0.9545 of the density of a bivariate normal distribution fitted to the errors, and text adjacent to the semi-major and semi-minor axes gives the square roots of the eigenvalues of the error covariance matrix.
[0163] Appendix G. EM Algorithm Derivation for Binary Classifier Testing
[0164]
[0165] Appendix H, EM Algorithm Derivation for Multi-class Classifier Testing IV. Definitions
[0166] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of processor-executable instructions that can be employed to program a computer or other processor to implement various aspects of embodiments as discussed above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the disclosure provided herein need not reside on a single computer or processor, but may be distributed in a modular fashion among different computers or processors to implement various aspects of the disclosure provided herein.
[0167] Processor-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0168] Also, data structures may be stored in one or more non-transitory computer-readable storage media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a non-transitory computer-readable medium that convey relationship between the fields. However, any suitable mechanism may be used to establish relationships among information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationships among data elements.
[0169] Also, various concepts of the disclosure may be embodied as one or more methods, of which examples have been provided. The acts performed as part of each process may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0170] All definitions, as defined and used herein, should be understood to control over dictionary definitions, and / or ordinary meanings of the defined terms. As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0171] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0172] Use of ordinal terms such as “first,” “second,” “third,” etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed. Such terms are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term).
[0173] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” “having,” “containing”, “involving”, and variations thereof, is meant to encompass the items listed thereafter and additional items.
[0174] The terms “approximately,” “substantially,” and “about” may be used to mean within ±20% of a target value in some embodiments, within ±10% of a target value in some embodiments, within ±5% of a target value in some embodiments, and yet within ±2% of a target value in some embodiments. The terms “approximately” and “about” may include the target value.
[0175] Having described several embodiments of the techniques described herein in detail, various modifications, and improvements will readily occur to those skilled in the art. Such modifications and improvements are intended to be within the spirit and scope of the disclosure. Accordingly, the foregoing description is by way of example only, and is not intended as limiting. The techniques are limited only as defined by the following claims and the equivalents thereto.
Claims
CLAIMSWhat is claimed is:
1. A method of evaluating performance of a trained classifier model, the method comprising: providing a set of predicted labels output by the trained classifier model; providing a set of noisy labels; providing a noisy-label model having a set of parameters; based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model: determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model; and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
2. The method of claim 1, wherein: the method further comprises providing an additional set of predicted labels output by the trained classifier model and an additional set of correct labels; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of predicted labels of the trained classifier model, the set of noisy labels, the set of parameters of the noisy-label model, and the additional set of predicted labels of the trained classifier model and the additional set of correct labels, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
3. The method of claim 1, wherein:determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model; and outputting the performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: outputting the performance evaluation of the trained classifier model based on the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model.
4. The method of claim 1, further comprising: providing a set of scores of the trained classifier model, each score of the set of score providing an indication of probability of truth for a predicted of the set of predicted labels of the trained classifier model; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of scores of the trained classifier model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
5. The method of claim 1, wherein determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining a first mean of a representation of an average number of true positives for the trained classifier model;determining a first variance of the representation of the average number of true positives for the trained classifier model; determining a second mean of a representation of an average number of false negatives for the trained classifier model; and determining a second variance of the representation of the average number of false negatives for the trained classifier model.
6. The method of claim 5, wherein the method further comprises: based on the first mean, first variance, second mean, and second variance: determining the representation of the probability of detection of the trained classifier model; determining the representation of the probability of false alarm of the trained classifier model; redetermining the first mean, first variance, second mean, and second variance; and based on the redetermined first mean, first variance, second mean, and second variance: iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
7. The method of claim 1, wherein determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: simulating M possible combinations of correct labels based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model; and based on the simulated M possible combinations of the correct labels: iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
8. The method of claim 1, wherein the trained classifier model comprises a multi-class trained classifier model and wherein the method further comprises: determining, for each possible value of a correct label, a representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label; and outputting a performance evaluation of the trained classifier model based on, for each possible value of a correct label, the representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label.
9. The method of claim 1, wherein: determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization.
10. The method of claim 9, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization comprises: calculating an expectation of a likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; and calculating a maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
11. The method of claim 10, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization further comprises repeatedly performing:calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; and subsequent to calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model, calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
12. The method of claim 10, wherein: calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises: calculating posterior probabilities of correct labels being equal to a first label; and calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises: determining a representation of a sum of posterior probabilities of events where predicted labels are equal to a second label and correct labels are equal to the first label divided by a sum of posterior probabilities of the events where correct labels are equal to the first label.
13. The method of any of claims 1-12, further comprising: determining whether to deploy or undeploy the trained classifier model based on the performance evaluation of the trained classifier model.
14. The method of claim 13, further comprising: deploying the trained classifier model to a classification environment.
15. The method of claim 14, further comprising:receiving input data for the trained classifier model; and using the trained classifier model, applying one or more labels to the input data.
16. At least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by at least one processor, cause the at least one processor to perform a method of a method of evaluating performance of a trained classifier model, the method comprising: providing a set of predicted labels output by the trained classifier model; providing a set of noisy labels; providing a noisy-label model having a set of parameters; based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model: determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model; and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
17. The at least one non-transitory computer-readable storage medium of claim 16, wherein: the method further comprises providing an additional set of predicted labels output by the trained classifier model and an additional set of correct labels; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of predicted labels of the trained classifier model, the set of noisy labels, the set of parameters of the noisy-label model, and the additional set of predicted labels of the trained classifier model and the additional set of correct labels, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
18. The at least one non-transitory computer-readable storage medium of claim 16, wherein: determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model; and outputting the performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: outputting the performance evaluation of the trained classifier model based on the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model.
19. The at least one non-transitory computer-readable storage medium of claim 16, wherein the method further comprises: providing a set of scores of the trained classifier model, each score of the set of score providing an indication of probability of truth for a predicted of the set of predicted labels of the trained classifier model; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of scores of the trained classifier model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
20. The at least one non-transitory computer-readable storage medium of claim 16, wherein determining the at least one of the representation of the probability of detection of the trainedclassifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining a first mean of a representation of an average number of true positives for the trained classifier model; determining a first variance of the representation of the average number of true positives for the trained classifier model; determining a second mean of a representation of an average number of false negatives for the trained classifier model; and determining a second variance of the representation of the average number of false negatives for the trained classifier model.
21. The at least one non-transitory computer-readable storage medium of claim 20, wherein the method further comprises: based on the first mean, first variance, second mean, and second variance: determining the representation of the probability of detection of the trained classifier model; determining the representation of the probability of false alarm of the trained classifier model; redetermining the first mean, first variance, second mean, and second variance; and based on the redetermined first mean, first variance, second mean, and second variance: iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
22. The at least one non-transitory computer-readable storage medium of claim 16, wherein determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: simulating M possible combinations of correct labels based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model; andbased on the simulated M possible combinations of the correct labels: iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
23. The at least one non-transitory computer-readable storage medium of claim 16, wherein the trained classifier model comprises a multi-class trained classifier model and wherein the method further comprises: determining, for each possible value of a correct label, a representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label; and outputting a performance evaluation of the trained classifier model based on, for each possible value of a correct label, the representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label.
24. The at least one non-transitory computer-readable storage medium of claim 16, wherein: determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization.
25. The at least one non-transitory computer-readable storage medium of claim 24, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization comprises: calculating an expectation of a likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; andcalculating a maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
26. The at least one non-transitory computer-readable storage medium of claim 25, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization further comprises repeatedly performing: calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; and subsequent to calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model, calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
27. The at least one non-transitory computer-readable storage medium of claim 25, wherein: calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises: calculating posterior probabilities of correct labels being equal to a first label; and calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises: determining a representation of a sum of posterior probabilities of events where predicted labels are equal to a second label and correct labels are equal to the first label divided by a sum of posterior probabilities of the events where correct labels are equal to the first label.
28. The at least one non-transitory computer-readable storage medium of any of claims 16- 27, wherein the method further comprises: determining whether to deploy or undeploy the trained classifier model based on the performance evaluation of the trained classifier model.
29. The at least one non-transitory computer-readable storage medium of claim 28, wherein the method further comprises: deploying the trained classifier model to a classification environment.
30. The at least one non-transitory computer-readable storage medium of claim 29, wherein the method further comprises: receiving input data for the trained classifier model; and using the trained classifier model, applying one or more labels to the input data.
31. A system for evaluating performance of a trained classifier model, the system comprising: at least one processor; and at least one non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by the at least one processor, cause the at least one processor to perform a method comprising: providing a set of predicted labels output by the trained classifier model; providing a set of noisy labels; providing a noisy-label model having a set of parameters; based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model: determining at least one of a representation of a probability of detection of the trained classifier model or a representation of a probability of false alarm of the trained classifier model; and outputting a performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
32. The system of claim 31, wherein: the method further comprises providing an additional set of predicted labels output by the trained classifier model and an additional set of correct labels; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of predicted labels of the trained classifier model, the set of noisy labels, the set of parameters of the noisy-label model, and the additional set of predicted labels of the trained classifier model and the additional set of correct labels, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
33. The system of claim 31, wherein: determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model; and outputting the performance evaluation of the trained classifier model based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: outputting the performance evaluation of the trained classifier model based on the representation of the probability of detection of the trained classifier model and the representation of the probability of false alarm of the trained classifier model.
34. The system of claim 31, wherein the method further comprises:providing a set of scores of the trained classifier model, each score of the set of score providing an indication of probability of truth for a predicted of the set of predicted labels of the trained classifier model; and based on the set of predicted labels of the trained classifier model, the set of noisy labels, and the set of parameters of the noisy-label model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: based on the set of scores of the trained classifier model, determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model.
35. The system of claim 31, wherein determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: determining a first mean of a representation of an average number of true positives for the trained classifier model; determining a first variance of the representation of the average number of true positives for the trained classifier model; determining a second mean of a representation of an average number of false negatives for the trained classifier model; and determining a second variance of the representation of the average number of false negatives for the trained classifier model.
36. The system of claim 35, wherein the method further comprises: based on the first mean, first variance, second mean, and second variance: determining the representation of the probability of detection of the trained classifier model; determining the representation of the probability of false alarm of the trained classifier model; redetermining the first mean, first variance, second mean, and second variance; and based on the redetermined first mean, first variance, second mean, and second variance:iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
37. The system of claim 31, wherein determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises: simulating M possible combinations of correct labels based on the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model; and based on the simulated M possible combinations of the correct labels: iterating the representation of the probability of detection of the trained classifier model; and iterating the representation of the probability of false alarm of the trained classifier model.
38. The system of claim 31, wherein the trained classifier model comprises a multi-class trained classifier model and wherein the method further comprises: determining, for each possible value of a correct label, a representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label; and outputting a performance evaluation of the trained classifier model based on, for each possible value of a correct label, the representation of a probability of the predicted label of the trained classifier model and the correct label divided by the probability of the correct label.
39. The system of claim 31, wherein: determining the at least one of the representation of the probability of detection of the trained classifier model or the representation of the probability of false alarm of the trained classifier model comprises:estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization.
40. The system of claim 39, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization comprises: calculating an expectation of a likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; and calculating a maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
41. The system of claim 40, wherein: estimating the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model using expectation maximization further comprises repeatedly performing: calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model; and subsequent to calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model, calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model.
42. The system of claim 40, wherein: calculating the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises:calculating posterior probabilities of correct labels being equal to a first label; and calculating the maximization of the expectation of the likelihood function based on the probability of detection of the trained classifier model or the probability of false alarm of the trained classifier model comprises: determining a representation of a sum of posterior probabilities of events where predicted labels are equal to a second label and correct labels are equal to the first label divided by a sum of posterior probabilities of the events where correct labels are equal to the first label.
43. The system of any of claims 31-42, wherein the method further comprises: determining whether to deploy or undeploy the trained classifier model based on the performance evaluation of the trained classifier model.
44. The system of claim 43, wherein the method further comprises: deploying the trained classifier model to a classification environment.
45. The system of claim 44, wherein the method further comprises: receiving input data for the trained classifier model; and using the trained classifier model, applying one or more labels to the input data.