Model performance evaluation method and apparatus

By employing a merge sort adversarial evaluation method, only model pairs generated during the merge sort process of the model set are labeled, which solves the problems of high labeling resource consumption and low accuracy in existing technologies, and achieves efficient and accurate model performance evaluation.

CN118860821BActive Publication Date: 2025-10-31ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410853179.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-10-31
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

Existing model performance evaluation methods require full-match adversarial evaluation between all models, which results in high consumption of human resources for annotation, long time consumption, and low annotation quality, making it difficult to accurately evaluate model performance.

Method used

The merge sort adversarial evaluation method is adopted, which only labels the model pairs generated during the merge sort process of the model set, reducing the amount of labeling and improving the evaluation efficiency and accuracy.

Benefits of technology

By reducing the amount of annotation, the probability of annotation errors is lowered, thereby improving the efficiency and accuracy of model performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118860821B_ABST
    Figure CN118860821B_ABST
Patent Text Reader

Abstract

This application provides one or more embodiments of a model performance evaluation method and apparatus. The method includes: acquiring a model set containing multiple models and a sample set containing at least one evaluation sample; sequentially determining the evaluation samples in the sample set as target evaluation samples, and acquiring a first type of model pair generated during the merging and sorting process of the model set; publishing the first type of model pair to annotators, so that the annotators can annotate the first type of model pair according to the model outputs obtained by inputting the target evaluation sample into the two models in the first type of model pair, and obtain annotation results indicating the comparison results of the model performance of the two models in the first type of model pair on the target evaluation sample; acquiring the annotation results of the first type of model pair on the target evaluation sample, and continuing to complete the merging and sorting of the model set according to the annotation results, to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this application relate to the field of artificial intelligence technology, and in particular to a model performance evaluation method and apparatus. Background Technology

[0002] With the rapid development of artificial intelligence and machine learning technologies, various complex models, such as deep neural networks and ensemble learning models, are widely used in fields such as image recognition, natural language processing, and recommendation systems. The performance of these models directly determines the accuracy and efficiency of the final application; therefore, model performance evaluation has become an indispensable part of ensuring algorithm reliability and optimizing model design. Summary of the Invention

[0003] One or more embodiments of this application provide the following technical solutions:

[0004] This application provides a model performance evaluation method, the method comprising:

[0005] Obtain a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample;

[0006] The evaluation samples in the sample set are sequentially determined as target evaluation samples, and the first type of model pairs to be compared are obtained during the merging and sorting process of the model set.

[0007] The first type of model pair is published to the annotation party, which then annotates the first type of model pair based on the model outputs obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, thereby obtaining the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample;

[0008] Obtain the annotation results of the first type of model on the target evaluation sample, and continue to complete the merge sort of the model set according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.

[0009] This application also provides a model performance evaluation device, the device comprising:

[0010] The acquisition module acquires a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample.

[0011] The first sorting module sequentially determines the evaluation samples in the sample set as target evaluation samples, and obtains the first type of model pairs to be compared during the merging and sorting process of the model set.

[0012] The publishing module publishes the first type of model pair to the annotator, so that the annotator can annotate the first type of model pair according to the model output obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, and obtain the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample;

[0013] The second sorting module obtains the annotation results of the first type of model on the target evaluation sample, and continues to perform merge sorting on the model set based on the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.

[0014] This application also provides an electronic device, including:

[0015] processor;

[0016] Memory used to store processor-executable instructions;

[0017] The processor executes the executable instructions to implement the steps of the method as described in any of the preceding descriptions.

[0018] This application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method as described in any of the preceding claims.

[0019] In the above technical solution, for a model set containing multiple models to be evaluated, the model performance of the models in the model set can be evaluated on any evaluation sample in the sample set. Specifically, model pairs generated during the merge sorting process of the model set can be obtained first. The model performance of the two models in these model pairs on the evaluation sample needs to be compared. Then, the obtained model pairs can be published to the annotation party, which can annotate each model pair according to the model output obtained by inputting the evaluation sample into each model pair, and obtain the corresponding annotation results. The annotation results can be used to indicate the comparison results of the model performance of the two models in each model pair on the evaluation sample. Finally, the merge sorting of the model set can be completed according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the evaluation sample.

[0020] By employing the aforementioned model performance evaluation method, only the model pairs requiring comparison generated during the merge sorting process of the model set need to be labeled, eliminating the need to label every pair of models in the set. This reduces the amount of labeling required during model performance evaluation. Consequently, it reduces the consumption of manpower and time spent on labeling, thereby improving the efficiency of model performance evaluation. Furthermore, the reduced amount of labeling lowers the probability of errors, thus improving the accuracy of model performance evaluation. Attached Figure Description

[0021] The accompanying drawings used in the description of the exemplary embodiments will now be explained, wherein:

[0022] Figure 1 This is a schematic diagram illustrating a process of merging and sorting a data set according to an exemplary embodiment of this application.

[0023] Figure 2 This is a flowchart illustrating a model performance evaluation method in an exemplary embodiment of this application.

[0024] Figure 3 This is a schematic diagram illustrating a process of merging and sorting a set of models, as shown in an exemplary embodiment of this application.

[0025] Figure 4 This is a schematic diagram of the structure of a device shown in an exemplary embodiment of this application.

[0026] Figure 5 This is a block diagram illustrating a model performance evaluation device according to an exemplary embodiment of this application. Detailed Implementation

[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this application. Rather, they are merely examples consistent with some aspects of one or more embodiments of this application.

[0028] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this application in other embodiments. In some other embodiments, the methods may include more or fewer steps than those described in this application. Furthermore, a single step described in this application may be broken down into multiple steps in other embodiments; and multiple steps described in this application may be combined into a single step in other embodiments.

[0029] When evaluating the performance of multiple models, adversarial evaluation can be used. Adversarial evaluation is a method for assessing and comparing model performance. It evaluates models by simulating adversarial scenarios, revealing the relative performance of models and providing guidance for model improvement and optimization.

[0030] Adversarial evaluation not only focuses on a model's performance on standard test sets but also delves deeper into its capabilities by designing challenging or adversarial test conditions. This approach reveals a more comprehensive picture of a model's strengths and weaknesses, particularly its response to carefully crafted inputs or complex environments. By comparing the performance of different models under the same adversarial tests, it becomes clear which model is more robust in handling specific types of tasks or resisting specific attacks. This comparison helps identify which model designs or training strategies are more effective.

[0031] The results of adversarial evaluations can serve as feedback, guiding researchers or developers to identify the shortcomings of a model and then make targeted improvements to the model structure, training data, or learning algorithms to enhance the model's generalization ability and robustness.

[0032] In adversarial evaluation, multiple models can be selected and configured with evaluation samples. Manual annotation is then used to label the outputs of two models in response to these samples, thereby evaluating the model performance in a specific task or scenario. For example, assuming the model is a binary classification model, the evaluation sample can be the data to be classified, and the model's output for the evaluation sample can be the classification result and its confidence probability corresponding to the evaluation sample. Alternatively, assuming the model is a Large Language Model (LLM), the evaluation sample can be a prompt text, and the model's output for the evaluation sample can be the text generated under the guidance of the evaluation sample.

[0033] Specifically, two models can be used as a comparison instance. For the same input or task, experts or trained annotators evaluate the quality of the two models' outputs. They might label the output results of the two models in the comparison instance as "left better," "right better," or "same" ("equally good" or "equally bad") for the evaluation samples. The distinction between "equally good" and "equally bad" is to facilitate reasoning and analysis of the differences in model performance between different models.

[0034] For example, suppose a comparison instance is "Model 1 vs. Model 2". If the accuracy of Model 1's output is higher than that of Model 2, then Model 1's performance is considered better than Model 2's, and this comparison instance is labeled "Left Better". If the accuracy of Model 2's output is higher than that of Model 1, then Model 2's performance is considered better than Model 1's, and this comparison instance is labeled "Right Better". If the accuracy of Model 1 and Model 2 is not significantly different and both are relatively high, then Model 1 and Model 2's performance is considered to be basically the same and both are good, and this comparison instance is labeled "Equally Good". If the accuracy of Model 1 and Model 2 is not significantly different and both are relatively low, then Model 1 and Model 2's performance is considered to be basically the same and both are relatively poor, and this comparison instance is labeled "Equally Poor". Here, "Left Better" means Model 1 is better than Model 2, "Right Better" means Model 2 is better than Model 1, and "Equally Good" or "Equally Poor" means Model 1 and Model 2 are tied.

[0035] Manually annotating and comparing the model's output results for the evaluation samples is an intuitive and practical way to evaluate model performance. It can analyze the performance differences of models at a more detailed level, especially in scenarios where automatic evaluation metrics may be difficult to distinguish accurately, thereby gaining in-depth insights into the performance differences of models in specific tasks.

[0036] Current model performance evaluation typically employs a fully matched adversarial evaluation method. Fully matched adversarial evaluation means that every two models in the evaluation pool are treated as comparison instances, and experts or trained annotators annotate the output of each comparison instance for each evaluation sample. That is, for N models and M evaluation samples, a full matching adversarial evaluation is required. Secondary annotation; among which, Therefore, full-match adversarial evaluation not only consumes a large amount of manpower for annotation but also takes a long time, resulting in low efficiency in model performance evaluation. Furthermore, full-match adversarial evaluation is usually done by a single person, leading to low annotation quality and frequent circular annotations. This means that the win-loss relationships between multiple models form a circular structure (e.g., Model 1 beats Model 2, Model 2 beats Model 3, and Model 3 beats Model 1), making it impossible to accurately evaluate the performance of multiple models.

[0037] This application provides one or more embodiments of a technical solution for model performance evaluation, which employs a merge sort adversarial evaluation method. In this solution, for a model set containing multiple models to be evaluated, the model performance of each model in the model set can be evaluated on any evaluation sample in the sample set. Specifically, model pairs generated during the merge sort process of the model set can be obtained first. The performance of the two models in these model pairs on the evaluation sample needs to be compared. Then, the obtained model pairs can be published to annotators, who can annotate each model pair based on the model outputs obtained by inputting the evaluation sample into each pair, thus obtaining corresponding annotation results. These annotation results can be used to indicate the comparison results of the model performance of the two models in each pair on the evaluation sample. Finally, based on the annotation results, the merge sort of the model set can be completed to obtain the ranking result of the model performance of the models in the model set on the evaluation sample.

[0038] By employing the aforementioned model performance evaluation method, only the model pairs requiring comparison generated during the merge sorting process of the model set need to be labeled, eliminating the need to label every pair of models in the set. This reduces the amount of labeling required during model performance evaluation. Consequently, it reduces the consumption of manpower and time spent on labeling, thereby improving the efficiency of model performance evaluation. Furthermore, the reduced amount of labeling lowers the probability of errors, thus improving the accuracy of model performance evaluation.

[0039] In this application, for a model set containing multiple models to be evaluated, a merge sort method can be used. The model set is merged and sorted according to the performance of the models to obtain the ranking result of the model performance in the model set. That is, all models in the model set are sorted in order of their performance from best to worst (or from worst to best). This reduces the amount of annotation required in the model performance evaluation process and improves the efficiency and accuracy of model performance evaluation.

[0040] For example, suppose a model set contains model 1, model 2, and model 3 to be evaluated, with M1 representing model 1, M2 representing model 2, and M3 representing model 3. Then the model set can be represented as {M1, M2, M3}. Further suppose that the performance of model 1 is worse than that of model 2, but better than that of model 3; the performance of model 2 is better than that of both model 1 and model 3; the performance of model 3 is better than that of both model 1 and model 2. Then, after using the merge sort algorithm to merge sort the model set according to the performance of the models, the resulting ranking of model performance could be M2 > M1 > M3.

[0041] The merge sort algorithm is explained below.

[0042] Merge sort is an efficient sorting algorithm based on the merge operation. It is a typical application of the divide-and-conquer method. Its core idea is to merge already sorted subsequences to obtain a completely sorted sequence; that is, first make each subsequence sorted, and then make the subsequences sorted relative to each other.

[0043] Merge sort divides the dataset into two subsets, sorts each subset, and finally merges the subsets into a single ordered dataset. Specifically, merge sort recursively divides the dataset into two subsets repeatedly until each subset contains only one element (if the dataset has an odd number of elements, a subset may have one more element). Then, it merges adjacent subsets into a single ordered dataset. This process of merging is repeated until the original dataset is fully sorted.

[0044] The specific steps of merge sort can be summarized into the following stages:

[0045] Decomposition Phase: Divide the dataset into two subsets starting from the middle position. If the dataset length is odd, the middle element can be placed in the left subset. This process is repeated recursively until each subset contains only one element, at which point the subsets are considered sorted.

[0046] Recursive sorting phase: Merge sort is performed on the two data subsets obtained from the decomposition. This means that each data subset will be further decomposed until they cannot be decomposed any further.

[0047] Merge Phase: Begin merging the already sorted subsets of data. The merge operation is bottom-up, meaning it first merges the smallest sorted subset, then progressively merges larger sorted subsets. The specific steps for merging two subsets are as follows: Create two pointers, each pointing to the starting position of one subset; compare the current elements of the two subsets, copy the smaller (or larger) element to a new result set, and move the corresponding pointer forward one position; repeat the above process until all elements of one subset have been copied to the result set; copy the remaining elements from the other subset to the remaining positions in the result set.

[0048] Recursive return phase: After the merge is complete, the recursive call will return a sorted subset of data. As the recursion returns step by step, each merge produces a larger ordered set of data.

[0049] Complete sorting: When the recursion reaches the initial data set, the entire data set will be sorted.

[0050] The merge sort algorithm will be illustrated with an example below.

[0051] Please refer to Figure 1 , Figure 1 This is a schematic diagram illustrating a process of merging and sorting a data set according to an exemplary embodiment of this application.

[0052] like Figure 1 As shown, taking the dataset {7,3,5,9} as an example, this dataset can be decomposed into two subsets: {7,3} and {5,9}. These two subsets can then be further decomposed into four subsets: {7}, {3}, {5}, and {9}. {7} and {3} can be merged by comparing 7 and 3. Since 7 > 3, {7} and {3} are merged into {7,3}. {5} and {9} can also be merged by comparing 5 and 9. Since 5 < 9, {5} and {9} are merged into {9,5}. Finally, {7,3} and {9,5} can be merged by first comparing 3 and 5. Since 3 < 5, then comparing 7 and 5. Since 7 > 5, then comparing 7 and 9. Since 7 < 9, {7,3} and {9,5} are merged into {9,7,5,3}. That is, the sorting result of the elements in the data set {7,3,5,9} according to their numerical values ​​from largest to smallest is 9>7>5>3.

[0053] The technical solution for model performance evaluation provided in this application is described below.

[0054] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating a model performance evaluation method in an exemplary embodiment of this application.

[0055] In this embodiment, the above-described model performance evaluation method can be applied to a server. This server can be a server containing a single independent physical host, or a server cluster consisting of multiple independent physical hosts; alternatively, the server can be a virtual server, cloud server, or similar service hosted by a host cluster. Alternatively, the above-described model performance evaluation method can be applied to electronic devices with a certain computing capability, such as tablet computers, laptops, desktop computers, personal computers (PCs), and handheld digital assistants (PDAs).

[0056] In practical applications, an evaluation platform capable of performing model performance evaluation tasks can be deployed on a server or an electronic device with a certain computing power. In this case, the above-mentioned model performance evaluation method can be applied to the evaluation platform.

[0057] like Figure 2 As shown, the above model performance evaluation method may include the following steps:

[0058] Step 202: Obtain a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample.

[0059] In this embodiment, a model set containing multiple models to be evaluated and a sample set containing at least one evaluation sample can be obtained first.

[0060] The aforementioned evaluation samples can be used to evaluate model performance. Specifically, the evaluation samples can be input into the model, which will then perform calculations based on the evaluation samples and output the corresponding results. Thus, the model performance can be evaluated based on the model's output results for the evaluation samples.

[0061] For example, assuming the model is a binary classification model, the evaluation sample can be the data to be classified, and the model's output for the evaluation sample can be the classification result corresponding to the evaluation sample and its confidence probability. In this case, the accuracy of the model's output can be calculated based on the true classification result of the evaluation sample and the classification result and its confidence probability predicted by the model for the evaluation sample, so as to reflect the model's performance through the accuracy of the output result.

[0062] For example, assuming the model is a large language model, the evaluation samples can be prompt text, and the model's output for the evaluation samples can be the text generated by the model under the guidance of the evaluation samples. In this case, the accuracy of the model's output can be calculated based on the semantics of the evaluation samples, as well as the semantics, structure, and word choice of the text predicted by the model for the evaluation samples, so as to reflect the model's performance through the accuracy of the output.

[0063] It should be noted that the accuracy of the output results of the same model may differ for different evaluation samples; that is, the model performance of the same model may vary on different evaluation samples. Therefore, the output result of a model for a single evaluation sample is usually used to evaluate the model's performance on that evaluation sample.

[0064] Step 204: Sequentially determine the evaluation samples in the sample set as target evaluation samples, and obtain the first type of model pairs to be compared during the merging and sorting process of the model set.

[0065] In this embodiment, the evaluation samples in the aforementioned sample set can be sequentially determined as target evaluation samples, thereby enabling the model performance of the models in the aforementioned model set on the target evaluation samples. For example, the sample set can be traversed, and the traversed evaluation samples can be determined as target evaluation samples. In this way, the model performance of the models in the model set can be evaluated on at least one evaluation sample, allowing for a more comprehensive analysis of the model performance differences in the model set and providing a more accurate evaluation of the model performance of the models in the model set.

[0066] For reference Figure 1 The merge sorting process shown involves sorting a dataset based on the merge sort algorithm, prioritizing model performance. First, it involves identifying the model pairs (referred to as first-class model pairs) generated during the merge sorting process. Specifically, the comparison focuses on the performance of each model within the first-class pair on a single evaluation sample (the target evaluation sample) within the dataset. For example, assuming a first-class model pair includes Model 1 and Model 2, the comparison is between the performance of Model 1 and Model 2 on the target evaluation sample. Therefore, the merge sorting process for this model set is essentially a process of merging and sorting the model performance on the target evaluation sample.

[0067] Step 206: Publish the first type of model pair to the annotator, so that the annotator can annotate the first type of model pair according to the model output obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, and obtain the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample.

[0068] In this embodiment, upon obtaining the aforementioned first type of model pair, the obtained first type of model pair can be published to the annotation party (e.g., experts capable of manual annotation or trained annotators, automatic annotation tools capable of automatic annotation, etc.). For a first type of model pair, the annotation party can annotate the first type of model pair based on the model outputs obtained by inputting the aforementioned target evaluation sample into the two models in the first type of model pair, thereby obtaining the annotation result of the first type of model pair on the target evaluation sample.

[0069] As mentioned earlier, the model output obtained by inputting the target evaluation sample into one of the models in a first-class model pair can be used to evaluate the model performance of that model in the first-class model pair on the target evaluation sample. The annotation results of the two models in the first-class model pair on the target evaluation sample can be used to indicate the comparison results of the model performance of the two models in the first-class model pair on the target evaluation sample.

[0070] For example, suppose a first-class model pair includes model 1 and model 2, where M1 represents model 1 and M2 represents model 2. This first-class model pair can be represented as (M1, M2). If model 1 performs better on the target evaluation sample than model 2, then the performance comparison of the two models in the first-class model pair on the target evaluation sample is M1 > M2, and the labeler can label this first-class model pair as "better on the left". If model 2 performs better on the target evaluation sample than model 1, then the performance comparison of the two models in the first-class model pair on the target evaluation sample is M2 > M1, and the labeler can label this first-class model pair as "better on the left". The model pair is labeled "Good on the right"; if the model performance of Model 1 on the target evaluation sample is basically the same as that of Model 2 on the target evaluation sample, then the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample is M1=M2, and the labeler can label the first type of model pair as "Same"; specifically, if the model performance of Model 1 on the target evaluation sample is basically the same as that of Model 2 on the target evaluation sample and both are good, then the labeler can label the first type of model pair as "Equally Good", and if the model performance of Model 1 on the target evaluation sample is basically the same as that of Model 2 on the target evaluation sample and both are poor, then the labeler can label the first type of model pair as "Equally Poor".

[0071] In some embodiments, to avoid situations where first-type model pairs are missed and not labeled, a labeling pool can be set up; wherein, the labeling pool can be centralized or distributed. Upon obtaining the aforementioned first-type model pairs, the identified first-type model pairs can be published to the labeling pool, thereby enabling the labelers to obtain these first-type model pairs from the labeling pool.

[0072] Furthermore, after the annotator completes the annotation of the aforementioned first type of model pairs, the annotation results can be fed back to the evaluation platform in real time through a message queue. Specifically, the annotator can act as a message producer in this message queue, and the evaluation platform can act as a message consumer in this message queue, with the annotator sending the annotation results to the evaluation platform through this message queue.

[0073] In some embodiments, to minimize the amount of annotation required during model performance evaluation, when the aforementioned first-type model pairs are obtained during the current model performance evaluation process, it can be first determined whether there are historical annotation results for each first-type model pair on the aforementioned target evaluation sample, i.e., whether each first-type model pair has been previously annotated on the target evaluation sample. For a first-type model pair, if there are historical annotation results for that first-type model pair on the target evaluation sample, these historical annotation results can be directly used as the annotation results for that first-type model pair on the target evaluation sample in the current model performance evaluation process, without re-annotating the first-type model pair on the target evaluation sample; however, if there are no historical annotation results for that first-type model pair on the target evaluation sample, it indicates that the first-type model pair needs to be annotated on the target evaluation sample, and therefore the first-type model pair can be published to the annotation party.

[0074] Step 208: Obtain the annotation results of the first type of model on the target evaluation sample, and continue to complete the merge sort of the model set according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.

[0075] In this embodiment, the annotation results of the first type of model pair on the target evaluation sample can be obtained. As mentioned earlier, the annotation results of a first type of model pair on the target evaluation sample can be used to indicate the comparison results of the model performance of the two models in the first type of model pair on the target evaluation sample. Therefore, after obtaining the annotation results, the model set can be merged and sorted according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample. In this way, the model performance differences of the models in the model set on the target evaluation sample can be analyzed through the ranking results, thereby providing a more accurate evaluation of the model performance of the models in the model set.

[0076] It should be noted that, since when performing merge sort on a set of models, it is also necessary to treat two models in the set as a comparison instance (i.e., a model pair), and the labeler labels the output results of the comparison instance for the evaluation sample, the technical solution provided in this application can be called merge sort adversarial evaluation.

[0077] In some embodiments, in order to enable the merge sorting process for the above-mentioned model set to be completed automatically and to ensure the execution efficiency of the merge sorting for the model set, a scheduled task can be used to advance the rounds of the merge sorting for the model set.

[0078] Specifically, since the models in the aforementioned model set are sorted according to their performance from best to worst (or worst to best), and the comparison results are indicated by the annotation results, in practical applications, the entire merge sorting process for this model set can be divided into multiple rounds based on annotation requirements. In each round, the model pairs to be compared generated in that round can be published to the annotation team, and the corresponding annotation results can be obtained through a scheduled task. The merge sorting of the model set can then continue based on the annotation results, completing the round. Once all the annotation results obtained in that round have been utilized, the round is considered complete. After completing the round, the process can proceed to the next round, generating new model pairs to be compared. These new model pairs will then be the model pairs to be compared in the next round. Alternatively, each time a new model pair to be compared is generated, the process can be considered a new round, and the generated new model pairs will be the model pairs to be compared in the new round. The division of rounds may depend on the specific steps of the above merge sort algorithm in its implementation (e.g., whether the decomposition, sorting and merging of multiple model subsets are performed in parallel), and this application does not impose any restrictions on this.

[0079] In this scenario, when obtaining the first type of model pairs generated during the merge sorting process for the aforementioned model set, specifically, the first type of model pairs generated in the current round of the merge sorting process can be obtained. Subsequently, the first type of model pairs generated in the current round of the merge sorting process for the model set can be published to the annotation team. The annotation team can then annotate each first type of model pair generated in this current round based on the model outputs obtained by inputting the target evaluation sample into each of the two models in the first type of model pair, thus obtaining the annotation results of each first type of model pair on the target evaluation sample.

[0080] Accordingly, when obtaining the annotation results of the first type of model pairs on the target evaluation samples, and continuing to complete the merge sort of the model set based on the annotation results to obtain the model performance ranking results of the models in the model set on the target evaluation samples, specifically, based on a preset timing strategy, the annotation results of the first type of model pairs generated in the current round on the target evaluation samples can be obtained periodically, and the current round of merge sort of the model set can be completed based on the annotation results. After completing the current round, the process can proceed to the next round of merge sort of the model set.

[0081] In other words, after completing the current round of merge sorting for the aforementioned model set, we can continue to obtain the first-class model pairs generated in the next round of merge sorting for this model set. Subsequently, the first-class model pairs generated in the next round of merge sorting for this model set can be published to the annotation team. The annotation team can then annotate each first-class model pair generated in the next round based on the model outputs obtained by inputting the target evaluation sample into each of the two models in the first-class model pair, thus obtaining the annotation results of each first-class model pair on the target evaluation sample.

[0082] Similarly, based on the above timing strategy, the annotation results of the first type of model pair generated in the next round on the target evaluation sample can still be obtained at regular intervals, and the next round of merge sorting for the model set can be completed according to the annotation results. After completing the next round, the round of merge sorting for the model set can be continued.

[0083] This process continues until the final round of merging and sorting the aforementioned model set is completed. After completing this final round, the ranking results of the model performance in this model set on the aforementioned target evaluation samples can be obtained.

[0084] The following example illustrates this. Figure 2 The illustrated embodiments will be used for explanation.

[0085] Please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating a process of merging and sorting a set of models, as shown in an exemplary embodiment of this application.

[0086] like Figure 3 As shown, suppose a model set contains model 1, model 2, model 3, model 4, model 5, model 6, model 7 and model 8, where M1 represents model 1, M2 represents model 2, M3 represents model 3, M4 represents model 4, M5 represents model 5, M6 represents model 6, M7 represents model 7 and M8 represents model 8, then the model set can be represented as {M1,M2,M3,M4,M5,M6,M7,M8}.

[0087] Taking the above model set as an example, the model set can be decomposed into two model subsets: {M1,M2,M3,M4} and {M5,M6,M7,M8}. These two model subsets can be further decomposed into four model subsets: {M1,M2}, {M3,M4}, {M5,M6}, and {M7,M8}. These four model subsets can be further decomposed into eight model subsets: {M1}, {M2}, {M3}, {M4}, {M5}, {M6}, {M7}, and {M8}.

[0088] Furthermore, {M1} and {M2}, {M3} and {M4}, {M5} and {M6}, and {M7} and {M8} can be merged. Specifically, the model performance of Model 1 and Model 2 on the aforementioned target evaluation samples can be compared, as can the model performance of Model 3 and Model 4 on the target evaluation samples, the model performance of Model 5 and Model 6 on the target evaluation samples, and the model performance of Model 7 and Model 8 on the target evaluation samples. Therefore, the aforementioned first-class model pairs generated in the first round can be (M1,M2), (M3,M4), (M5,M6), and (M7,M8). Assuming the first type of model labels (M1, M2) as "right better" on the target evaluation sample, (M3, M4) as "equally good" on the target evaluation sample, (M5, M6) as "left better" on the target evaluation sample, and (M7, M8) as "equally bad" on the target evaluation sample, then we can merge {M1} and {M2} into {M2, M1}, {M3} and {M4} into {M3, M4}, {M5} and {M6} into {M5, M6}, and {M7} and {M8} into {M7, M8}.

[0089] Furthermore, we can merge {M2,M1} and {M3,M4}, and merge {M5,M6} and {M7,M8}. Specifically, we can first compare the model performance of Model 2 and Model 3 on the aforementioned target evaluation samples, and then compare the model performance of Model 5 and Model 7 on the target evaluation samples. Therefore, the first type of model pair generated in the second round can be (M2,M3) and (M5,M7). Assuming that the labeling result of the first type of model pair (M2,M3) on the target evaluation samples is "left is better", and the labeling result of the first type of model pair (M5,M7) on the target evaluation samples is also "left is better", we can then compare the model performance of Model 1 and Model 3 on the aforementioned target evaluation samples, and then compare the model performance of Model 6 and Model 7 on the target evaluation samples. Therefore, the first type of model pair generated in the third round can be (M1,M3) and (M6,M7). Suppose that the labeling result of the first type of model pair (M1, M3) on the target evaluation sample is "right best", and the labeling result of the first type of model pair (M6, M7) on the target evaluation sample is "equally good". Then we can compare the model performance of Model 1 and Model 4 on the target evaluation sample, and merge {M5, M6} and {M7, M8} into {M5, M6, M7, M8}. Since M3 = M4, there is no need to construct the first type of model pair (M1, M4) and label it. Instead, we can automatically infer that M4 > M1, and merge {M2, M1} and {M3, M4} into {M2, M3, M4, M1}.

[0090] Furthermore, we can merge {M2, M3, M4, M1} and {M5, M6, M7, M8}. Specifically, we can first compare the model performance of Model 2 and Model 5 on the aforementioned target evaluation samples. Therefore, the first type of model pair generated in the fourth round can be (M2, M5). Assuming that the labeling result of the first type of model pair (M2, M5) on the target evaluation samples is "left is better", we can then compare the model performance of Model 3 and Model 5 on the target evaluation samples. Therefore, the first type of model pair generated in the fifth round can be (M3, M5). Assuming that the labeling result of the first type of model pair (M3, M5) on the target evaluation samples is "left is better", we can then compare the model performance of Model 4 and Model 5 on the target evaluation samples. Since M3 = M4, there is no need to construct the first type of model pair (M4, M5) and label it. Instead, we can automatically infer that M4 > M5, and thus we can compare the model performance of Model 1 and Model 5 on the target evaluation samples. Therefore, the first type of model pair generated in the 6th round can be (M1, M5). Assuming that the labeling result of the first type of model pair (M1, M5) on the target evaluation sample is "left is better", then {M2, M3, M4, M1} and {M5, M6, M7, M8} can be merged into {M2, M3, M4, M1, M5, M6, M7, M8}.

[0091] In other words, the sixth round mentioned above is the final round of merge sorting of the model set {M1, M2, M3, M4, M5, M6, M7, M8} on the aforementioned target evaluation samples. The ranking of the model performance of the models in this model set on the aforementioned target evaluation samples is M2>M3=M4>M1>M5>M6>M7=M8.

[0092] In the process of merging and sorting the model set {M1, M2, M3, M4, M5, M6, M7, M8} on the aforementioned target evaluation samples, the first type of model pairs to be compared include: (M1, M2), (M3, M4), (M5, M6), (M7, M8), (M2, M3), (M5, M7), (M1, M3), (M6, M7), (M2, M5), (M3, M5), and (M1, M5). That is, only 11 annotations are needed in total, and no further annotations are required. Secondary annotation.

[0093] It should be noted that the time complexity of the merge sort algorithm is O(n log n). For a model set containing 8 models, the merge sort process for that model set requires at most [number] steps. Secondary annotation, and no longer needed Secondary annotation.

[0094] Furthermore, since the number of annotations in merge sort adversarial evaluation is significantly reduced when the number of models or evaluation samples is large compared to full matching adversarial evaluation, in merge sort adversarial evaluation, for a first-class model pair, multiple annotators can vote to annotate the first-class model pair, thereby improving the annotation quality.

[0095] In some embodiments, in order to facilitate the analysis of the model performance differences of models in the model set on the target evaluation samples, the models in the model set can be combined in pairs to construct model pairs. The model pairs other than the first type of model pairs mentioned above are identified as the second type of model pairs. Subsequently, the labeling results of the second type of model pairs on the target evaluation samples can be inferred based on the labeling results of the first type of model pairs on the target evaluation samples.

[0096] Continue as Figure 3 Taking the model set shown as an example, as mentioned earlier, in the process of merging and ranking the model set {M1,M2,M3,M4,M5,M6,M7,M8} on the above target evaluation samples, the first type of model pairs that need to be compared include: (M1,M2), (M3,M4), (M5,M6), (M7,M8), (M2,M3), (M5,M7), (M1,M3), (M6,M7), (M2,M5), (M3,M5) and (M1,M5).

[0097] The second type of model pairs, which are constructed by combining the models in the above model set in pairs, are as follows, in addition to the first type of model pairs mentioned above: (M1,M4), (M1,M6), (M1,M7), (M1,M8), (M2,M4), (M2,M6), (M2,M7), (M2,M8), (M3,M6), (M3,M7), (M3,M8), (M4,M5), (M4,M6), (M4,M7), (M4,M8), (M5,M8), (M6,M8).

[0098] The sum of the number of the first type of model pairs and the number of the second type of model pairs is . .

[0099] Based on the annotation results of the first type of model on the target evaluation samples, the annotation results of the second type of model on the target evaluation samples can be determined recursively by size relationship. This allows the annotation results from the merge sorting process of the model set on the target evaluation samples to be completed as full-match adversarial evaluation annotation results for the model set on the target evaluation samples. The completed full-match annotation results are shown in Table 1 below.

[0100] Table 1

[0101]

[0102] In some embodiments, as described above, for the target evaluation sample, the annotation results of the full-match adversarial evaluation of the model set on the target evaluation sample can be obtained by completion. Similarly, for any evaluation sample in the sample set, the annotation results of the full-match adversarial evaluation of the model set on that evaluation sample can be obtained by completion. In this case, the scores corresponding to each model in the model set can be calculated based on the annotation results of the first type of model on each evaluation sample in the sample set and the annotation results of the second type of model on each evaluation sample in the sample set; wherein, the score corresponding to a model can be used to characterize the model performance of that model on the sample set. For example, based on a preset cyclic integral rule, the integrals corresponding to each model in the model set can be calculated based on the annotation results of the first type of model on each evaluation sample and the annotation results of the second type of model on each evaluation sample, and used as the scores corresponding to each model in the model set.

[0103] Round-robin points are typically used to determine the ranking of participating teams or individuals in a round-robin tournament. These rules allocate points based on the results of the matches (win, loss, draw), and the rankings are determined by the total points after all matches are played.

[0104] As mentioned earlier, for a model pair (M1, M2) consisting of Model 1 and Model 2, "left better" means Model 1 is better than Model 2, "right better" means Model 2 is better than Model 1, and "equally good" or "equally bad" means Model 1 and Model 2 are tied. The round-robin scoring rules adopted in this application can specifically be a win-draw-loss points system, that is, a win earns 3 points, a draw earns 1 point, and a loss (or forfeit) earns 0 points.

[0105] In some embodiments, the win rate corresponding to each model in the model set can be calculated based on the annotation results of the first type of model on each evaluation sample in the sample set and the annotation results of the second type of model on each evaluation sample in the sample set; wherein, the win rate corresponding to a model is proportional to the model performance of that model on the sample set.

[0106] Win-draw percentage is used to assess the ratio of wins to draws in competitive matches. It can be calculated using the following formula: Win-draw percentage = (Number of wins + Number of draws) / Total number of matches.

[0107] In some embodiments, to reduce the impact of mislabeled samples, a rough initial ranking of the models can be given based on their strength. Therefore, when obtaining the aforementioned model set containing multiple models to be evaluated, the historical win rate corresponding to each model among the multiple models to be evaluated can be obtained. The historical win rate is determined based on the historical labeling results corresponding to the model pairs containing each model. Thus, these models can be sorted according to the order of their historical win rates, and the model set can be constructed based on the sorting result. For example, assuming that the historical win rate of model 1 is 50%, that of model 2 is 60%, that of model 3 is 40%, and that of model 4 is 80%, and let M1 represent model 1, M2 represent model 2, M3 represent model 3, and M4 represent model 4, then the sorting result of these four models according to their historical win rates from largest to smallest is M4>M2>M1>M3, and the constructed model set can be represented as {M4,M2,M1,M3}.

[0108] In practical applications, offline calculations can be used to calculate the scores corresponding to each model in the above model set, as well as the winning odds corresponding to each model in the above model set.

[0109] In the above technical solution, for a model set containing multiple models to be evaluated, the model performance of the models in the model set can be evaluated on any evaluation sample in the sample set. Specifically, model pairs generated during the merge sorting process of the model set can be obtained first. The model performance of the two models in these model pairs on the evaluation sample needs to be compared. Then, the obtained model pairs can be published to the annotation party, which can annotate each model pair according to the model output obtained by inputting the evaluation sample into each model pair, and obtain the corresponding annotation results. The annotation results can be used to indicate the comparison results of the model performance of the two models in each model pair on the evaluation sample. Finally, the merge sorting of the model set can be completed according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the evaluation sample.

[0110] By employing the aforementioned model performance evaluation method, only the model pairs requiring comparison generated during the merge sorting process of the model set need to be labeled, eliminating the need to label every pair of models in the set. This reduces the amount of labeling required during model performance evaluation. Consequently, it reduces the consumption of manpower and time spent on labeling, thereby improving the efficiency of model performance evaluation. Furthermore, the reduced amount of labeling lowers the probability of errors, thus improving the accuracy of model performance evaluation.

[0111] Corresponding to the embodiments of the methods described above, this application also provides embodiments of the apparatus.

[0112] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the structure of a device according to an exemplary embodiment of this application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, memory 408, and non-volatile memory 410, and may also include other necessary hardware. One or more embodiments of this application can be implemented in software, for example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into memory 408 and then runs it. Of course, besides software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic modules, but can also be hardware or logic devices.

[0113] Please refer to Figure 5 , Figure 5 This is a block diagram illustrating a model performance evaluation device according to an exemplary embodiment of this application.

[0114] The above-mentioned model performance evaluation device can be applied to Figure 4 The apparatus shown is used to implement the technical solution of this application. The apparatus includes:

[0115] Module 502 obtains a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample.

[0116] The determination module 504 sequentially determines the evaluation samples in the sample set as target evaluation samples, and obtains the first type of model pairs to be compared generated during the merging and sorting process of the model set.

[0117] The publishing module 506 publishes the first type of model pair to the annotator, so that the annotator can annotate the first type of model pair according to the model output obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, and obtain the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample;

[0118] The sorting module 508 obtains the annotation results of the first type of model on the target evaluation sample, and continues to perform merge sorting on the model set according to the annotation results to obtain the model performance ranking results of the models in the model set on the target evaluation sample.

[0119] In some embodiments, the apparatus further includes:

[0120] The construction module combines the models in the model set in pairs to construct model pairs, and determines the other model pairs in the constructed model pairs other than the first type of model pairs as the second type of model pairs;

[0121] The inference module infers the annotation results of the second type of model on the target evaluation sample based on the annotation results of the first type of model on the target evaluation sample.

[0122] In some embodiments, the apparatus further includes:

[0123] The first calculation module calculates the score corresponding to each model in the model set based on the annotation results of the first type of model on each evaluation sample and the annotation results of the second type of model on each evaluation sample; wherein the score is used to characterize the model performance of each model in the model set on the sample set.

[0124] In some embodiments, the apparatus further includes:

[0125] The second calculation module calculates the win rate corresponding to each model in the model set based on the annotation results of the first type of model on each evaluation sample and the annotation results of the second type of model on each evaluation sample; wherein the win rate is proportional to the model performance of each model in the model set on the sample set.

[0126] In some embodiments, obtaining a model set containing multiple models to be evaluated includes:

[0127] Obtain the historical win rate for each model among multiple models to be evaluated; wherein the historical win rate is determined based on the historical annotation results corresponding to the model pair containing the model;

[0128] The multiple models are sorted according to their historical win and draw rates, and a model set is constructed based on the sorting results.

[0129] In some embodiments, publishing the first type of model pair to the annotation party includes:

[0130] Determine whether there are historical annotation results for each type I model pair on the target evaluation sample; if so, determine the historical annotation results as the annotation results of the type I model pair on the target evaluation sample; otherwise, publish the type I model pair to the annotator.

[0131] In some embodiments, obtaining the first type of model pairs to be compared, generated during the merge sorting process for the model set, includes:

[0132] Obtain the first type of model pair to be compared in the current round of merge sorting during the merge sorting process for the model set;

[0133] The step of obtaining the annotation results of the first type of model on the target evaluation sample, and further performing merge sorting on the model set based on the annotation results to obtain the model performance ranking results of the models in the model set on the target evaluation sample, includes:

[0134] Based on a preset timing strategy, the annotation results of the first type of model pairs on the target evaluation sample are obtained periodically, and the current round of merge sorting of the model set is completed according to the annotation results. After the current round is completed, the first type of model pairs to be compared generated in the next round of merge sorting of the model set are triggered, until the last round of merge sorting of the model set is completed, and the model performance ranking results of the models in the model set on the target evaluation sample are obtained.

[0135] For the device embodiments, they basically correspond to the method embodiments; therefore, relevant details can be found in the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the technical solution of this application according to actual needs.

[0136] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0137] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0138] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0139] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0140] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0141] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of this application. In some cases, the actions or steps described in this application may be performed in a different order than those shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0142] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. The term “and / or” refers to and includes any or all possible combinations of one or more associated listed items.

[0143] The terms "an embodiment," "some embodiments," "example," "specific example," or "one implementation," as used in one or more embodiments of this application, refer to specific features or characteristics described in connection with that embodiment, which are included in at least one embodiment of this application. Illustrative descriptions of these terms do not necessarily refer to the same embodiment. Furthermore, the described specific features or characteristics may be combined in a suitable manner in one or more embodiments of this application. In addition, different embodiments and specific features or characteristics from different embodiments may be combined without contradiction.

[0144] It should be understood that although the terms first, second, third, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of one or more embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0145] The above description is merely a preferred embodiment of one or more embodiments of this application and is not intended to limit the scope of one or more embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this application should be included within the protection scope of one or more embodiments of this application.

[0146] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

Claims

1. A model performance evaluation method, the method comprising: Obtain a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample; The evaluation samples in the sample set are sequentially determined as target evaluation samples, and the first type of model pairs to be compared are obtained during the merging and sorting process of the model set. The first type of model pair is published to the annotation party, which then annotates the first type of model pair based on the model outputs obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, thereby obtaining the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample; Obtain the annotation results of the first type of model on the target evaluation sample, and continue to complete the merge sort of the model set according to the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.

2. The method according to claim 1, further comprising: The models in the model set are combined in pairs to construct model pairs, and the model pairs other than the first type of model pairs in the constructed model pairs are determined as the second type of model pairs; Based on the annotation results of the first type of model on the target evaluation sample, the annotation results of the second type of model on the target evaluation sample are inferred.

3. The method according to claim 2, further comprising: Based on the annotation results of the first type of model on each evaluation sample and the annotation results of the second type of model on each evaluation sample, a score corresponding to each model in the model set is calculated; wherein, the score is used to characterize the model performance of each model in the model set on the sample set.

4. The method according to claim 2, further comprising: Based on the annotation results of the first type of model on each evaluation sample and the annotation results of the second type of model on each evaluation sample, calculate the win rate corresponding to each model in the model set; wherein the win rate is proportional to the model performance of each model in the model set on the sample set.

5. The method according to claim 1, wherein obtaining the model set containing multiple models to be evaluated comprises: Obtain the historical win rate for each model among multiple models to be evaluated; wherein the historical win rate is determined based on the historical annotation results corresponding to the model pair containing the model; The multiple models are sorted according to their historical win and draw rates, and a model set is constructed based on the sorting results.

6. The method according to claim 1, wherein publishing the first type of model pair to the annotation party comprises: Determine whether each of the first-class models has historical annotation results on the target evaluation samples; If so, the historical annotation results are determined as the annotation results of the first type of model pair on the target evaluation sample; otherwise, the first type of model pair is published to the annotator.

7. The method according to claim 1, wherein obtaining the first type of model pairs to be compared generated during the merge sorting process for the model set includes: Obtain the first type of model pair to be compared generated in the current round of merge sorting during the merge sorting process for the model set; The step of obtaining the annotation results of the first type of model on the target evaluation sample, and further performing merge sorting on the model set based on the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample, includes: Based on a preset timing strategy, the annotation results of the first type of model pairs on the target evaluation sample are obtained periodically, and the current round of merge sorting of the model set is completed according to the annotation results. After the current round is completed, the first type of model pairs to be compared generated in the next round of merge sorting of the model set are triggered until the last round of merge sorting of the model set is completed, so as to obtain the ranking result of the model performance of the models in the model set on the target evaluation sample.

8. A model performance evaluation device, the device comprising: The acquisition module acquires a model set containing multiple models to be evaluated, and a sample set containing at least one evaluation sample. The first sorting module sequentially determines the evaluation samples in the sample set as target evaluation samples, and obtains the first type of model pairs to be compared during the merging and sorting process of the model set. The publishing module publishes the first type of model pair to the annotator, so that the annotator can annotate the first type of model pair according to the model output obtained by inputting the target evaluation sample into the two models in the first type of model pair respectively, and obtain the annotation result of the first type of model pair on the target evaluation sample; wherein, the annotation result is used to indicate the comparison result of the model performance of the two models in the first type of model pair on the target evaluation sample; The second sorting module obtains the annotation results of the first type of model on the target evaluation sample, and continues to perform merge sorting on the model set based on the annotation results to obtain the ranking results of the model performance of the models in the model set on the target evaluation sample.

9. An electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor implements the method as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Label sample determination method and device, machine readable medium and equipment

    CN112257812A

  • Deep learning test case sorting method based on variation analysis

    CN113128556A