Method for validating a trained model
A method for detecting and correcting misclassifications in supervised learning models by comparing their outputs to a test dataset, ensuring model integrity and security in critical applications.
Patent Information
- Application Number
- JP2025501373
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-11
- Filing Date
- 2023-07-07
- Publication Date
- 2025-07-17
AI Technical Summary
Existing models trained using supervised learning methods, such as neural networks, lack a theoretical proof of correctness and are susceptible to misclassifications due to biased design or attacks, posing security and efficiency threats in critical applications.
A method involving a computer system that compares the answers of multiple trained models to a test dataset, performing a homogeneity test to identify models with abnormal behavior by analyzing probability distribution laws or direct comparisons, and taking corrective actions when deviations are detected.
Effectively detects and corrects misclassifications in models, preventing attacks and ensuring model integrity by identifying and addressing biased training or poisoning, thus enhancing security and efficiency.
Smart Images

Figure 2025523025000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the verification of models trained using supervised learning methods such as neural networks, and more specifically, to a method for detecting a model that serves as a control against other models among a plurality of such models.
Background Art
[0002] Models trained using supervised learning methods such as neural networks are becoming increasingly valuable tools for solving classification problems such as image recognition, pattern recognition, or speech recognition, or even for making predictions, for example, by regression. Such neural networks can be used for biometric verification. Accurate verification authenticates a person correctly; inaccurate verification can result in false positives, i.e., interpreting a fraudster as an authenticated person, and false negatives, i.e., misinterpreting a person as a fraudster.
[0003] Furthermore, as the use of automation, such as in autonomous vehicles and factory automation, increases, it is necessary to interpret operating environments such as traffic conditions, road conditions, and factory floor conditions. Such models have become central tools used in these applications of image processing.
[0004] The efficiency and security of such models are of utmost importance for the security of systems that rely on those models for mission-critical purposes. As an example, consider a neural network used to verify whether a given person should be allowed to pass through a certain filter, such as a gate, a secure door, or border control. It may be based on facial recognition, i.e., whether that person matches an image digitally recorded of a specific person or identification object in a database. The operation of that neural network is one mechanism by which its security system can be attacked. For example, the neural network may be manipulated to recognize a forger as an authorized person.
[0005] Similarly, such models used in automation must be efficient and secure. For example, one can envision a terrorist or threat attack based on an intrusion into a neural network used in an autonomous vehicle. For example, imagine the derivative problems of an autonomous vehicle after an attack on an image processing neural network that fails to recognize a pedestrian at a crosswalk. Similar threats can occur in the field of an automated factory, a delivery system, aviation, etc.
[0006] The problem with models trained using supervised classification is that their correctness cannot yet be theoretically proven. Their correctness can only be inferred by using performance metrics such as the percentage of correct answers.
[0007] However, such statistical metrics do not prevent the possibility of having misclassified data that results in labels different from the expected labels for this particular data or type of data. Such misclassification can be the result of an attacker causing a defect in the model to spontaneously acquire such misclassification and use it for their own benefit.
[0008] Alternatively, such misclassifications can be an undesirable effect of model drawbacks resulting from, for example, a biased design or overfitting during the training phase.
[0009] In both cases, such misclassifications are a threat to the efficiency and security of such models and should be detected and properly processed. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION
[0010] From the foregoing, it is clear that there is a need for an improved method for detecting models that have a biased design or are erroneously trained using supervised learning methods and lead to such misclassifications. MEANS FOR SOLVING THE PROBLEM
[0011] For this purpose, according to a first aspect, the present invention is thus a method for detecting a deviating model among a plurality of different models trained using a supervised learning method, the method being performed by a computer system programmed by the trained model: - obtaining a test dataset, - presenting the test dataset to each of the trained models and generating an answer of each trained model for the test dataset, - performing at least one homogeneity test based on the answers generated by at least two of the plurality of trained models, - when the homogeneity test fails, performing a predetermined action indicating that one of the at least two models has been detected as deviating from other trained models and relates to a method.
[0012] Such a method makes it possible to determine that one of the tested models provided an answer for a test dataset that is significantly different from what was provided by other models, and to take appropriate action.
[0013] The model may be among a neural network model, a k-nearest neighbors (KNN) model, a support vector machine (SVM) model, a decision tree model, and a quadratic discriminant analysis (QDA) model.
[0014] The model may be based on several different learning methods or on the same learning method using different parameters and / or hyperparameters.
[0015] In one embodiment, the step of performing a homogeneity test based on the answers for the test dataset generated by at least two of the plurality of trained models includes determining, for each model, the probability distribution law that its answer follows based on the generated answers for the test dataset, and comparing the probability distribution laws determined for the at least two models.
[0016] This makes it possible to check that the answers of the models for the test dataset follow the same distribution, even if they are not identical, and that the behavior of the tested models is equivalent.
[0017] The step of determining, for each model, the probability distribution law that its answer follows may include establishing the law based on the answers for the test dataset generated by the model using regression techniques.
[0018] In one embodiment, the step of performing a homogeneity test based on the answer for a test data set generated by at least two of a plurality of trained models includes performing a direct comparison of the answer for the test data set.
[0019] The step of performing a direct comparison of the answer for the test data set may include performing a pairwise homogeneity test, a Cramer-von-Mises test, or a Kolmogorov-Smirnov test.
[0020] Such a test makes it possible to compare the distributions of the answers of the models being tested, even if the probability distribution law for the answers of the models cannot be determined.
[0021] The step of performing a direct comparison of the answer for a test data set generated by two models may include determining a test statistic value related to the answers of the two models for the test data set and comparing the test statistic value with a predetermined threshold.
[0022] The step of performing a direct comparison of the answer for a test data set generated by two models may include determining a p-value for the test statistic and comparing the value of the p-value with a predetermined threshold.
[0023] Calculating the test statistic or the p-value makes it possible to quantify the distance between the distributions of the answers of the test data sets of the two models and to determine whether the answers of these two models can be considered to follow the same distribution.
[0024] In one embodiment, the step of performing a predetermined action when the homogeneity test based on the answers generated by two models fails includes declaring one of the two models to be defective and discarding it.
[0025] This makes it possible to prevent any use of the model that has been caused to have a defect by an attacker during training.
[0026] In one embodiment, the step of performing a predetermined action when the homogeneity test based on the answers generated by the two models fails includes performing new training on one of the two models using a new training dataset.
[0027] This makes it possible to correct any defect in the training applied to the model, for example, any bias or any overfitting in the training database used to train the model.
[0028] According to a second aspect, the present invention is thus a computer system programmed using a plurality of different models trained using a supervised learning method, comprising: - a processor configured to obtain a test dataset, - at least one memory storing the trained models and connected to a processor including instructions executable by the processor, the instructions being: · presenting the test dataset to each of the trained models and generating an answer of each trained model for the test dataset, · performing at least one homogeneity test based on the answers generated by at least two of the plurality of trained models, · when the homogeneity test fails, performing a predetermined action indicating that one of the at least two models has been detected as deviating from other trained models and relates to a computer system.
[0029] According to a third aspect, the present invention relates to a computer program product that can be directly loaded into the memory of at least one computer, the product including software code instructions for performing the steps of the method according to the first aspect of the present invention when the product is executed on a computer.
[0030] To achieve the foregoing and related objects, one or more embodiments are fully described below and include features particularly pointed out in the claims.
Brief Description of the Drawings
[0031]
Figure 1
Figure 2
Figure 3
Modes for Carrying Out the Invention
[0032] The technology described in this specification provides a method for detecting abnormal behavior of a model trained using a supervised learning method such as a statistical learning method or a machine learning method.
[0033] Such a model can be, for example, a neural network (such as a Reduced Boltzmann Machine (RBM)), a k-nearest neighbor method (KNN), a support vector machine (SVM), a decision tree model, or a quadratic discriminant analysis (QDA) model.
[0034] Such abnormal behavior is, for example, a misclassification of a given input to an unexpected class to which the input does not belong, or a prediction of a value with a large error.
[0035] Such incorrect output can be the result of an attack during the training of the model. An attacker can, for example, add one or more incorrect samples to the training dataset used to train the model in order to spontaneously induce the incorrect classifications observed when the model is being used. Such incorrect output can also be the result of an inadequate selection of the model's parameters or hyperparameters, such as dropout values or the number of layers of an MLP model.
[0036] Such incorrect classifications can also result from flaws in the model, for example, caused by an inadequate or biased training dataset, or by overfitting the model to the training dataset such that it is impossible to correctly handle inputs different from this training dataset.
[0037] To detect such abnormal behavior and take appropriate actions to prevent it, the main idea of the present invention is to obtain several trained models that should provide similar answers for a given input and compare their answers for a given test dataset.
[0038] As an example, such models can be models trained by the same training data but based on different learning methods such as neural network and decision tree models. Alternatively, they can be models based on the same learning method but having different parameters and / or hyperparameters, because they are trained and provided by different entities such as models provided for free on the Internet and models provided by subcontractors.
[0039] The models to be compared can even be different versions of a single model at multiple different stages of its training. In such cases, comparing these different versions can make it possible to check that the output of the model does not diverge when it is further trained.
[0040] When such models have similar performance, their outputs provided as answers for a given test dataset should be almost identical and should follow the same probability distribution law. Therefore, the second main idea of the present invention is to determine which model among the compared models has abnormal behavior by comparing the probability distribution laws of the models to be compared, and to identify a model whose output does not follow the same probability distribution law as other outputs as a "deviating model".
[0041] Figure 1 is a high-level architecture diagram showing a possible architecture for a computer system 100 that performs the steps of the method according to the present invention. The computer system 100 includes a processor 101, and the processor 101 can be a microprocessor, a graphics processing unit, or any processing unit suitable for executing the steps of the method described below.
[0042] The processor 101 is connected to the memory 102 via a bus, for example. The memory 102 may be a non-volatile memory (NVM) such as a flash memory or an erasable programmable read-only memory (EPROM). The memory 102 stores the trained models to be compared, including the parameters and hyperparameters of the trained models, such as the weights associated with the neurons that create the neural network and are adjusted during the training of the neural network.
[0043] The computer system 100 further includes an input / output interface 103 for receiving inputs, such as data regarding the models to be compared and settings for the comparisons to be made between the models to be compared.
[0044] Figure 2 is a flowchart showing the steps of a method for detecting a deviating model among a plurality of different models.
[0045] During the first step S1, the computer system, which is used to compare the models to be compared as described above, is programmed with the trained models to be compared. As described above, the trained models may, for example, be pre-trained outside the computer system 100 by a supplier and imported into the computer system through the input / output interface 103. Alternatively, the trained models may be trained within the computer system 100 itself. During this step, the trained models are set in the computer system, particularly its memory 102, whereby the computer system can interrogate each model with an input and obtain the answer of each model to this model as an output. For example, in the case of a classification application, it is a vector of scores.
[0046] In the second step S2, the computer system obtains a test data set containing a plurality of data used as inputs to the model to test the behavior of the model. Such a test data set may, for example, contain thousands or millions of test samples. Such a test data set may be imported into the computer system through the input / output interface 103 or may be generated by the computer system, for example, from a larger data set.
[0047] In the third step S3, the computer system presents the test data set to each of the trained models to be compared, which generates the answers of each trained model to the test data set. When the data set is properly selected, the answers of the model to the data set should characterize the behavior of the model and reveal any abnormal behavior of the model, such as classifying a picture of a cat as a dog when the model is used to classify pictures of animals.
[0048] Unfortunately, especially when the test dataset contains thousands of samples, it is very tedious and time-consuming to have a human check the model's answers for each sample in the dataset in order to detect any abnormal behavior as described above. Therefore, in the fourth step S4, the computer system performs at least one homogeneity test based on the answers generated by at least two of the trained models tested in the third step S3. Instead of identifying abnormal behavior of the model solely from those answers, such a homogeneity test rather identifies abnormal behavior of the model by comparing those answers with the answers of at least another model for the same test dataset.
[0049] In a first embodiment, such a homogeneity test includes determining, for each model undergoing the homogeneity test, the probability distribution law that its answer follows, based on the answers generated for the test dataset. Such a law can be, for example, the Gaussian law or the gamma distribution. Such a law can be determined using regression techniques in the answers for the test dataset generated by each model. The computer system then compares the probability distribution laws determined for the models. The computer system can first check whether the answers of all the models follow the same type of distribution law, for example, whether they all follow the gamma distribution. The computer system can then compare the parameters (mean value, variance, etc.) of the laws followed by the answers of all the models. If all the tested models are equivalent, their answers for the same test dataset should all follow the same probability distribution law. If one model has abnormal behavior leading to inaccurate answers for some samples of the test dataset, the answers of this particular model for the test dataset will follow a probability distribution law different from that of the other models. An example is provided in FIG. 3 which shows the distributions of the answers of three models for a test dataset. It can be seen that the distribution of the answers of the left model is significantly different from the distributions of the answers of the other two models which appear identical.
[0050] In some cases, such as when the model's answer does not follow a parametric law, it can be difficult to characterize the probability distribution law that the model's answer follows. Or, for example, in order to avoid errors by the regression method, it can be considered preferable not to do so. Therefore, in the second embodiment, the homogeneity test includes a direct comparison of the answers of the two models for the test data set. Such a direct comparison may include, for example, performing a pairwise homogeneity test, a Cramer von Mises test, or a Kolmogorov-Smirnov test.
[0051] The performance index can be defined to measure the distance between the distributions of the answers of the two models for the test data set.
[0052] In the first embodiment, such a direct comparison of the answers for the test data set generated by the two models includes determining a test statistic value related to the answers. For example, it is an effect size such as Pearson's correlation, coefficient of determination, or Eta-squared value.
[0053] In the second embodiment, such a direct comparison of the answers for the test data set generated by the two models further includes determining a p-value for such a test statistic.
[0054] In the third embodiment, such a distance between the distributions of the answers of the two models for the test data set can be determined using a non-statistical method such as the Wasserstein distance method.
[0055] Next, this test statistic value, or the value of the p-value, or any other distance metric, can be compared by a computer system to a predetermined threshold value. When the test statistic value or other distance is higher than such a predetermined threshold value, or when the p-value is lower than such a predetermined threshold value, the distance between the answers of the two models being compared is considered too large, and the answers of the two models are considered not to follow the same distribution. In such a case, the homogeneity test is considered to have failed.
[0056] In the fifth step S5, when the homogeneity test fails, the computer system performs a predetermined action indicating that one of the at least two models that received the failed test has been detected as deviating from the other trained models.
[0057] In the first embodiment, the method according to the present invention is used to test the correctness of a model that may have been poisoned by an attacker during training. In such a case, failing the test indicates that one of the models gives unreasonable answers for some of the samples of the test dataset, which may be the result of poisoning of the training dataset used to train the model. In such a case, such an action may be, for example, declaring one of the two models as defective and discarding it. The computer system may issue an alarm.
[0058] In the second embodiment, the method according to the present invention is used to determine whether model parameters such as dropout values, the number of layers, and the number of neurons per layer are appropriate for the problem to be solved. In such a case, failing the test may indicate that one or more parameters are inappropriate. In such a case, the predetermined action may include updating at least one parameter of one of the two models and, optionally, performing new training of the updated model using the same training dataset as before.
[0059] In a third embodiment, the method according to the invention is used to determine whether the training of the tested model was sufficient. In such a case, failing the test may indicate that the model is under-trained or over-learned, which can lead to inaccurate results for at least a part of the test dataset. In such a case, the predetermined action may include retraining one of the two models using a new training dataset.
[0060] The homogeneity test described above can only determine whether the distributions of the answers of the two models are the same and conclude that one of the two models being compared exhibits abnormal behavior, without indicating which of the two models has this abnormal behavior. To identify the model with the abnormal behavior, a number of tests may be performed on various pairs of models, and combining the results of these multiple tests makes it possible to determine which model has a distribution of answers that deviates from the distribution shared by the other models. The multiple performance indices calculated for these pairs of models can also be used to rank the models and define the preferred model, which can be either one of the tested models or a weighted combination of the tested models. In the case of a model that is periodically updated by continuous learning, the weights used in such a combination can be periodically updated by performing the method according to the invention again on the updated model.
[0061] According to a second aspect, the invention is a computer system programmed using a plurality of different models trained using a supervised learning method as previously described herein, - a processor 101 configured to obtain a test dataset, - at least one memory 102 storing the trained models and connected to a processor containing instructions executable by the processor comprising, the instructions being · Presenting the test dataset to each of the trained models and generating an answer of each trained model for the test dataset; · Performing at least one homogeneity test based on the answers generated by at least two of the plurality of trained models; · When the homogeneity test fails, performing a predetermined action indicating that one of the at least two models has been detected as deviating from other trained models. relates to a computer system including.
[0062] According to a third aspect, the present invention is a computer program product that can be directly loaded into the memory of at least one computer, and when the product is executed on a computer, it includes software code instructions for performing the steps of the method described previously herein. It relates to a computer program product.
[0063] In addition to these features, the computer program according to the second aspect of the present invention can be configured to perform any other features described previously herein or can include any other features thereof.
[0064] Such a method, computer system, and computer program product can detect abnormal behavior of a model that provides inaccurate answers for a limited number of inputs, even if such inaccurate answers have only a limited impact on the overall performance of the model. Therefore, this makes it possible to detect at low cost any poisoning of the model training by an attacker or any insufficient training of the model resulting from, for example, a biased training dataset.
Claims
**Claim 1** A method for detecting an outlier model among a plurality of different models trained using a supervised learning method, the method being performed by a computer system (100) programmed by a trained model, - obtaining a test data set (S2), - presenting the test data set to each of the trained models and generating an answer of each trained model for the test data set (S3), - performing at least one homogeneity test based on the answers generated by at least two of the plurality of trained models (S4), - when the homogeneity test fails, performing a predetermined action indicating that one of the at least two models has been detected as deviating from other trained models (S5) and including. **Claim 2** The method according to claim 1, wherein the model is one of a neural network model, a k-nearest neighbor (KNN) model, a support vector machine (SVM) model, a decision tree model, and a quadratic discriminant analysis (QDA) model. **Claim 3** The method according to claim 1 or 2, wherein the model is based on several different learning methods. **Claim 4** The method according to claim 1 or 2, wherein the model is based on the same learning method using different parameters and / or hyperparameters. **Claim 5** Performing a homogeneity test based on the answers for the test data set generated by at least two of the plurality of trained models includes determining a probability distribution law that the answer of each model follows based on the generated answer for the test data set, and comparing the probability distribution laws determined for the at least two models. The method according to any one of claims 1 to 4. **Claim 6** Determining the probability distribution law that the answer of each model follows includes establishing the law based on the answer for the test data set generated by the model using a regression technique. The method according to claim 5. **Claim 7** Performing a homogeneity test based on the answer for a test data set generated by at least two of a plurality of trained models, including performing a direct comparison of the answer for the test data set, the method according to any one of claims 1 to 4.
8. Performing a direct comparison of the answer for the test data set includes performing a pairwise homogeneity test, a Cramer von Mises test, or a Kolmogorov-Smirnov test, the method according to claim 7.
9. Performing a direct comparison of the answer for a test data set generated by two models includes determining a test statistic value related to the answer of the two models for the test data set and comparing the test statistic value with a predetermined threshold value, the method according to claim 7.
10. Performing a direct comparison of the answer for a test data set generated by two models includes determining a p-value for the test statistic and comparing the value of the p-value with a predetermined threshold value, the method according to claim 7.
11. Performing a predetermined action when the homogeneity test based on the answer generated by two models fails includes declaring one of the two models as defective and discarding it, the method according to any one of claims 1 to 10.
12. Performing a predetermined action when the homogeneity test based on the answer generated by two models fails includes performing new training on one of the two models using a new training data set, the method according to any one of claims 1 to 10.
13. A computer system (100) programmed using a plurality of different models trained using a supervised learning method, - a processor (101) configured to obtain a test data set, - at least one memory (102) storing the trained models and connected to a processor including instructions executable by the processor comprising, the instructions being · presenting the test data set to each of the trained models and generating an answer of each trained model for the test data set, - Performing at least one homogeneity test based on the answers generated by at least two of the plurality of trained models; - When the homogeneity test fails, performing a predetermined action indicating that one of the at least two models has been detected as deviating from other trained models comprise a computer system (100).
14. A computer program product that can be directly loaded into the memory of at least one computer, the product comprising software code instructions for performing the steps according to any one of claims 1 to 12 when the product is executed on a computer.
Citation Information
Patent Citations
Method and an apparatus for evaluating generative machine learning model
US20190012581A1
Data Analytics Model Selection through Champion Challenger Mechanism
US20200184398A1
Automated data and label creation for supervised machine learning regression testing
US20210142222A1