Evaluation subset selection method and device, electronic equipment and storage medium
By using a method of selecting target evaluation subsets through multiple rounds of iterations and utilizing difference differences and annealing temperature adjustments, the problems of high computational overhead and cost in the evaluation of large language models are solved, achieving efficient and accurate evaluation results.
Patent Information
- Application Number
- CN202510776365.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-10
AI Technical Summary
The existing large-scale benchmark test set evaluation method for large language models has high computational overhead and cost, and lacks a systematic subset selection method to minimize evaluation errors and preserve semantics.
Through multiple rounds of iterative selection of target evaluation subsets, the difference difference is used to select the evaluation subset, combined with annealing temperature adjustment to avoid local optimal solutions, ensuring that the evaluation error is close to the benchmark test set while taking into account semantic distribution consistency.
The use of computing resources is reduced, the evaluation cost is lowered, and the evaluation efficiency and accuracy of the evaluation subset are improved, ensuring that the evaluation error is close to the evaluation result of the benchmark test set.
Smart Images

Figure CN120763013A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of model evaluation, and particularly relates to a test subset selection method and device, electronic equipment and a storage medium. BACKGROUND
[0002] With the continuous expansion of the scale of large language models (LLM), comprehensive evaluation of the performance of the large language models has become an important demand in the field of artificial intelligence.
[0003] In related technologies, the performance of the large language models is evaluated by using large-scale benchmark test sets (such as MMLU, HellaSwag, GSM8K, etc.) to statistically evaluate the accuracy of the models in multiple tasks and multiple dimensions.
[0004] However, the method of directly using the large-scale benchmark test sets to statistically evaluate the models in multiple tasks and multiple dimensions has a large computational overhead and high cost. SUMMARY
[0005] The present application provides a test subset selection method and device, electronic equipment and a storage medium, which at least partially overcome the problem of large computational overhead and high cost in related technologies.
[0006] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.
[0007] According to one aspect of the present application, a test subset selection method is provided, comprising: obtaining a score vector of a benchmark test set, the score vector of the benchmark test set being constructed based on score data obtained by predicting the benchmark test set by a plurality of language models; performing a preset number of selection processes on the benchmark test set, each selection process comprising: selecting a candidate test subset from the benchmark test set, and calculating a difference between the candidate test subset and the benchmark test set in the score vector to obtain a candidate difference; subtracting a reference difference in the current selection process from the candidate difference to obtain a difference difference value, the reference difference in the current selection process being a difference between a target test subset selected in a previous selection process and the benchmark test set in the score vector, and if the current selection process is the first selection process, the target test subset selected in the previous selection process is an initial test subset selected from the benchmark test set; and selecting a target test subset in the current selection process from the target test subset selected in the previous selection process and the candidate test subset based on the difference difference value, wherein when the difference difference value is not greater than 0, the probability of the candidate test subset being selected is negatively correlated with the difference difference value, and when the difference difference value is greater than 0, the probability of the candidate test subset being selected is 1.
[0008] By selecting the target evaluation subset in multiple rounds of selection on the benchmark test set, the target evaluation subset can be applied for test evaluation during model evaluation, thereby reducing the use of computing resources and reducing computing overhead and cost.
[0009] Further, in the process of selecting the final target evaluation subset through multiple iterations, the difference difference value of the difference between the two evaluation subsets is compared in each round, and the evaluation subset is selected based on the difference difference value, so that the evaluation error of the selected evaluation subset for evaluating the performance of the large language model is similar to the evaluation error when the performance of the large language model is evaluated using the benchmark test set, and the semantic distribution consistency of the subset and the benchmark test set is also considered.
[0010] In some embodiments, selecting the candidate evaluation subset from the benchmark test set comprises: randomly selecting a preset number of evaluation questions from the benchmark test set; replacing the same number of evaluation questions in the target evaluation subset selected in the last round with the preset number of evaluation questions, and taking the subset obtained after the replacement processing as the candidate evaluation subset.
[0011] In some embodiments, the preset number is 1.
[0012] Through experimental data comparison, when the preset number is 1, the target evaluation subset obtained after the same number of iterations of selection processing has a smaller evaluation error when evaluating the large language model compared with the benchmark test set.
[0013] In some embodiments, the probability of the candidate evaluation subset being selected is as follows:
[0014]
[0015] T = aT old ;
[0016] Wherein, AQ is the difference difference value; T is the annealing temperature in the current round of selection processing; T old is the annealing temperature in the last round of selection processing, and the current round of selection processing is the first round, then T old is a preset annealing temperature; a is a preset attenuation factor greater than 0 and less than 1.
[0017] Through the above formula design, when the difference of the candidate evaluation subset is greater than the difference of the last round of target evaluation subset, it still has a certain probability of being selected. This way can avoid the selection of the subset into a local optimal solution, and further explore more possible subset selections, which is beneficial to the final selected target evaluation subset to more comprehensively and accurately evaluate the performance of the large language model.
[0018] Furthermore, by designing a method in which the annealing temperature decreases with each round, it is possible to ensure that when selecting a subset, it is less likely to select a "poor" subset (one with a large evaluation error compared to the performance evaluation of the benchmark test set on the large language model) and this helps the iteratively selected subset to be closer to the performance evaluation of the benchmark test set on the large language model.
[0019] In some embodiments, it also includes: obtaining the semantic embedding representation of the target evaluation subset obtained after the preset round of selection processing to obtain a first embedding representation; obtaining the semantic embedding representation of the benchmark test set to obtain a second embedding representation; calculating the distance between the first embedding representation and the second embedding representation, and the distance is used to represent the semantic similarity between the target evaluation subset obtained after the preset round of selection processing and the benchmark test set.
[0020] According to another aspect of the present application, another evaluation subset selection method is provided, including: obtaining a score vector of a benchmark test set, wherein the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set by multiple language models; performing multiple rounds of random selection from the benchmark test set to obtain multiple alternative evaluation subsets; calculating the difference between each alternative evaluation subset and the benchmark test set in the score vector respectively to obtain multiple differences; and taking the alternative evaluation subset corresponding to the minimum difference among the multiple differences as the target evaluation subset.
[0021] By performing multiple rounds of selection in the benchmark test set, the target evaluation subset is finally determined. Then, when evaluating the model, the target evaluation subset can be directly used for evaluation testing, thereby reducing the use of computing resources and lowering computing overhead and costs.
[0022] Furthermore, by selecting the alternative evaluation subset with the smallest difference from multiple rounds of randomly selected alternative evaluation subsets as the target evaluation subset, it can ensure that the evaluation error is small when using the target evaluation subset to evaluate the performance of the large language model.
[0023] According to still another aspect of the present application, there is also provided an evaluation subset selection apparatus, comprising: a first obtaining module configured to obtain a score vector of a benchmark test set, the score vector of the benchmark test set being constructed based on score data predicted by a plurality of language models on the benchmark test set; and a first selecting module configured to perform a preset number of selection processes on the benchmark test set, each selection process comprising: selecting a candidate evaluation subset from the benchmark test set, and calculating a difference between the candidate evaluation subset and the benchmark test set in the score vector to obtain a candidate difference; subtracting a reference difference in a previous selection process from the candidate difference to obtain a difference difference value, the reference difference in the previous selection process being a difference between a target evaluation subset selected in a last selection process and the benchmark test set in the score vector, and if the current selection process is the first selection process, the target evaluation subset selected in the last selection process is an initial evaluation subset selected from the benchmark test set; and selecting a target evaluation subset in the current selection process from the target evaluation subset selected in the last selection process and the candidate evaluation subset based on the difference difference value, wherein when the difference difference value is not greater than 0, a probability of the candidate evaluation subset being selected is negatively correlated with the difference difference value, and when the difference difference value is greater than 0, the probability of the candidate evaluation subset being selected is 1.
[0024] In some embodiments, the first selecting module is configured to randomly select a preset number of evaluation questions from the benchmark test set; replace a same number of evaluation questions in the target evaluation subset selected in the last selection process with the preset number of evaluation questions, and take the subset obtained after the replacement as the candidate evaluation subset.
[0025] In some embodiments, the apparatus further comprises a semantic analysis module configured to obtain a semantic embedding representation of the target evaluation subset obtained after the preset number of selection processes, to obtain a first embedding representation; obtain a semantic embedding representation of the benchmark test set, to obtain a second embedding representation; and calculate a distance between the first embedding representation and the second embedding representation, the distance being used to represent a semantic similarity between the target evaluation subset obtained after the preset number of selection processes and the benchmark test set.
[0026] According to still another aspect of the present application, there is also provided another evaluation subset selection apparatus, comprising: a second obtaining module configured to obtain a score vector of a benchmark test set, the score vector of the benchmark test set being constructed based on score data predicted by a plurality of language models on the benchmark test set; and a second selecting module configured to perform a plurality of random selection processes on the benchmark test set to obtain a plurality of candidate evaluation subsets; and calculate a difference between each candidate evaluation subset and the benchmark test set in the score vector to obtain a plurality of differences; and take a candidate evaluation subset corresponding to a minimum difference in the plurality of differences as a target evaluation subset.
[0027] According to another aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned evaluation subset selection methods by executing the executable instructions.
[0028] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements any one of the above-mentioned evaluation subset selection methods.
[0029] According to another aspect of the present application, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned evaluation subset selection methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0031] Figure 1 A flow chart of a method for selecting an evaluation subset in one embodiment of the present application is shown;
[0032] Figure 2 A flow chart of a method for selecting an evaluation subset in another embodiment of the present application is shown;
[0033] Figure 3 A flowchart of semantic evaluation processing of a target evaluation subset in one embodiment of the present application is shown;
[0034] Figure 4 A schematic diagram showing the evaluation errors of evaluation subsets selected on the MMLU using different methods in one embodiment of the present application;
[0035] Figure 5 A schematic diagram showing the evaluation errors of evaluation subsets selected on HellaSwag using different methods in one embodiment of the present application;
[0036] Figure 6 A schematic diagram showing the evaluation errors of evaluation subsets selected on GSM8K using different methods in one embodiment of the present application;
[0037] Figure 7 A schematic diagram of an evaluation subset selection device in one embodiment of the present application is shown;
[0038] Figure 8A schematic diagram of an evaluation subset selection device in another embodiment of the present application is shown;
[0039] Figure 9 A structural block diagram of an electronic device in one embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0041] In addition, the accompanying drawings are merely schematic illustrations of the present application and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0042] As the scale of large language models continues to expand, a comprehensive evaluation of their performance has become an important need in the field of artificial intelligence. The evaluation methods in related technologies usually rely on large-scale benchmark test sets (such as MMLU, HellaSwag, GSM8K, etc.) to perform multi-task and multi-dimensional accuracy statistics on the model. However, although this full evaluation method is comprehensive, it is accompanied by huge computational overhead, including computing resource consumption (GPU / TPU (Tensor Processing Unit) duration), API call fees, and data labeling and management costs. Therefore, how to reduce the evaluation cost while ensuring the reliability of the evaluation has become an urgent problem to be solved.
[0043] To this end, some scholars have proposed methods based on subset evaluation. Typical technical paths include:
[0044] 1. Random sampling method: Randomly select a certain proportion of questions from the original test set as the evaluation subset.
[0045] 2. Clustering-based method: By embedding the questions, and then applying unsupervised methods such as K-Means clustering, representative questions are selected in the semantic space to construct a subset.
[0046] 3. Dimensionality reduction based clustering method: Considering that there is information redundancy in the high-dimensional embedding space, some studies propose to perform PCA (Principal Component Analysis) dimensionality reduction or use an autoencoder to compress the embedding vectors of the questions before clustering, in order to improve the representativeness and robustness of the subset.
[0047] It can be seen that the existing methods are mainly based on heuristic clustering or random sampling, and lack a systematic subset selection method that minimizes the evaluation error. Therefore, there is still a need for a new subset selection method that minimizes the overall evaluation error while considering semantic preservation, to effectively improve the efficiency and accuracy of large language model evaluation.
[0048] To this end, the present application proposes an evaluation subset selection scheme. In the process of selecting the final target evaluation subset through multiple iterations, the difference difference value between the two evaluation subsets is compared in each iteration, and the evaluation subset is selected based on the difference difference value. When the evaluation subset selected in this way is used to evaluate the performance of a large language model, the evaluation error is similar to the evaluation error when the benchmark test set is used to evaluate the performance of a large language model, and at the same time, the semantic distribution consistency of the subset and the benchmark test set is considered.
[0049] The specific implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0050] Figure 1 A flowchart of an evaluation subset selection method in an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the evaluation subset selection method provided in the embodiment of the present application includes the following S101 to S102. Figure 1
[0051] S101, obtain the score vector of the benchmark test set, the score vector of the benchmark test set being constructed based on score data obtained by predicting the benchmark test set by a plurality of language models.
[0052] The benchmark test set is a set of question data used to evaluate and test the performance of a large language model. For example, the benchmark test set can be one of MMLU (a publicly available set of question data commonly used for testing large language models), GSM8K (a publicly available set of question data commonly used for testing large language models), HellaSwag (a publicly available set of question data commonly used for testing large language models), etc.
[0053] In one embodiment, obtaining the score vector of the benchmark test set can include: obtaining score data obtained by predicting the benchmark test set by a plurality of language models; and constructing the score vector based on the score data.
[0054] The method for obtaining the score data obtained by predicting the benchmark test set using multiple language models can be to directly apply multiple different language models to the benchmark test set to obtain score data for each language model on the benchmark test set, and then integrate the score data of the multiple language models on the benchmark test set to obtain a score vector (score matrix). For example, if the benchmark test set is MMLU, the total number of questions in MMLU is 14042, and the multiple different language models are 100 different language models, then the size of the score matrix is 14042×100.
[0055] Another method for obtaining the score data obtained by multiple language models predicting the benchmark test set is to download the evaluation results of multiple models on the benchmark test set from a public pre-evaluation dataset (such as Open LLM Leaderboard), extract the predicted score for each question in the benchmark test set, and obtain the score data of the multiple language models predicting the benchmark test set.
[0056] It should be noted that the questions in the benchmark test set must be from authoritative sources and have high-quality content. The multiple language models used must cover diverse ability levels to ensure the representativeness of the scoring matrix. The answer format of the questions must be unified (for example, multiple-choice questions must use a 0 / 1 scoring system) to facilitate subsequent error calculations.
[0057] S102, performing a preset round of selection process from the benchmark test set, each round of selection process includes the following S1021-S1023.
[0058] S1021, selecting an alternative evaluation subset from the benchmark test set, and calculating the difference in score vector between the alternative evaluation subset and the benchmark test set to obtain an alternative difference.
[0059] In one embodiment, selecting the candidate evaluation subset from the benchmark test set may include randomly selecting the candidate evaluation subset from the benchmark test set according to a preset compression ratio. For example, if the total number of questions in the MMLU is 14,042 and the preset compression ratio is 0.19, approximately 2,668 questions (14,042×0.19) are selected from the 14,042 questions to obtain the candidate evaluation subset.
[0060] In another embodiment, selecting an alternative evaluation subset from the benchmark test set may include: randomly selecting a preset number of evaluation questions from the benchmark test set; using the preset number of evaluation questions to replace the same number of evaluation questions in the target evaluation subset selected in the previous round, and using the subset obtained after the replacement process as the alternative evaluation subset.
[0061] For example, the preset number of evaluation questions is 5, 5 evaluation questions are randomly selected from the benchmark test set, and then the 5 evaluation questions are replaced with the 5 evaluation questions in the target evaluation subset selected in the last round. Then, the target evaluation subset after the evaluation question replacement processing is taken as the candidate evaluation subset.
[0062] The preset number is not limited in the embodiments of the present application. For example, the preset number is 1, 2, 3, etc.
[0063] In one embodiment, the preset number is 1.
[0064] Through experimental data comparison, when the preset number is 1, the target evaluation subset obtained after the same number of iterations of the selection processing has a smaller evaluation error when evaluating the large language model compared with the benchmark test set.
[0065] In one embodiment, calculating the difference between the candidate evaluation subset and the benchmark test set in the score vector to obtain the candidate difference can include: extracting score data corresponding to the candidate evaluation subset from score data obtained by predicting the benchmark test set by the plurality of language models; constructing a score vector of the candidate evaluation subset according to the score data corresponding to the candidate evaluation subset; and calculating the distance between the score vector of the candidate evaluation subset and the score vector of the benchmark test set to obtain the candidate difference.
[0066] The embodiments of the present application do not limit the way of calculating the distance between the score vector of the candidate evaluation subset and the score vector of the benchmark test set. For example, the L2 distance (Euclidean distance) can be calculated to calculate the distance between the score vector of the candidate evaluation subset and the score vector of the benchmark test set. For another example, the cosine distance can also be calculated to calculate the distance between the score vector of the candidate evaluation subset and the score vector of the benchmark test set.
[0067] It should be noted that, for the convenience of describing the technical solutions, the L2 distance is taken as an example to calculate the distance between different score vectors in the subsequent content of the present application.
[0068] S1022, subtract the reference difference in the current round of selection processing from the candidate difference to obtain a difference difference value, the reference difference in the current round of selection processing is the difference between the target evaluation subset selected in the last round and the benchmark test set in the score vector, and if the current round of selection processing is the first round, the target evaluation subset selected in the last round is the initial evaluation subset selected from the benchmark test set.
[0069] It should be noted that the number of questions included in the initial evaluation subset is the same as the number of questions included in the candidate evaluation subset.
[0070] The difference between the target evaluation subset and the benchmark test set in the score vector is calculated in the same way as the difference between the alternative evaluation subset and the benchmark test set in the score vector, which can be referred to as S1021, and will not be described here.
[0071] The alternative difference can be represented as Wherein, S is the benchmark test set, S' new is the alternative evaluation subset, is a plurality of language models. The reference difference can be represented as Q S' old is the target evaluation subset selected in the last round. The difference difference value can be represented as AQ, and the calculation method of AQ can be shown in the following formula 1.
[0072]
[0073] In S1023, the target evaluation subset in the current round is selected from the target evaluation subset selected in the last round and the alternative evaluation subset based on the difference difference value. When the difference difference value is not greater than 0, the probability of the alternative evaluation subset being selected is negatively related to the difference difference value. When the difference difference value is greater than 0, the probability of the alternative evaluation subset being selected is 1.
[0074] In an embodiment, selecting the target evaluation subset in the current round from the target evaluation subset selected in the last round and the alternative evaluation subset based on the difference difference value can include: calculating the probability of the alternative evaluation subset being selected based on the difference difference value, and determining the probability of the target evaluation subset selected in the last round being selected according to the probability; selecting the target evaluation subset in the current round from the target evaluation subset selected in the last round and the alternative evaluation subset based on the probabilities of the target evaluation subset selected in the last round and the alternative evaluation subset being selected. The selection method can be random selection between the two according to the probabilities.
[0075] The embodiments of the present application do not limit how to calculate the probability of the alternative evaluation subset being selected based on the difference difference value, as long as the probability of the alternative evaluation subset being selected is negatively related to the difference difference value when the difference difference value is not greater than 0, and the probability of the alternative evaluation subset being selected is 1 when the difference difference value is greater than 0.
[0076] In an embodiment, the probability of the alternative evaluation subset being selected is shown in the following formula 2 and formula 3:
[0077]
[0078] T = aT old (3)
[0079] Wherein, AQ is the difference difference value; T is the annealing temperature in the current round of selection processing; T oldis the annealing temperature in the previous round of selection treatment, and this round of selection treatment is the first round, then T old is the preset annealing temperature; α is a preset attenuation factor greater than 0 and less than 1.
[0080] The specific value of the preset annealing temperature is not limited in the embodiments of the present application and can be set based on experience. The specific value of α is not limited in the embodiments of the present application and can be set based on experience.
[0081] The probability of the candidate evaluation subset being selected is p, and the probability of the target evaluation subset selected in the previous round being selected is 1-p.
[0082] Through the above formula design, when the difference of the alternative evaluation subset is greater than the difference of the target evaluation subset in the previous round, it still has a certain probability of being selected. This method can prevent the subset selection from entering the local optimal solution, and then explore the possibility of more subset selections, which is conducive to the final selected target evaluation subset to more comprehensively and accurately evaluate the performance of the large language model.
[0083] Furthermore, by designing a method in which the annealing temperature decreases with each round, it is possible to ensure that when selecting a subset, it is less likely to select a "poor" subset (one with a large evaluation error compared to the performance evaluation of the benchmark test set on the large language model) and this helps the iteratively selected subset to be closer to the performance evaluation of the benchmark test set on the large language model.
[0084] By performing multiple rounds of selection on the benchmark test set, a target evaluation subset is finally selected, so that the target evaluation subset can be applied for test evaluation during model evaluation, thereby reducing the use of computing resources and lowering computing overhead and costs.
[0085] Furthermore, in the process of selecting the final target evaluation subset through multiple rounds of iteration, the difference between the two evaluation subsets is compared in each round, and the evaluation subset is selected based on the difference. This ensures that the evaluation error of the selected evaluation subset when used to evaluate the performance of the large language model is close to the evaluation error when the benchmark test set is used to evaluate the performance of the large language model, while taking into account the semantic distribution consistency of the subset and the benchmark test set.
[0086] In one embodiment, Figure 2 As shown, another evaluation subset selection method in the embodiment of the present application may include: S201-S204.
[0087] S201, obtaining a score vector of a benchmark test set, where the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models.
[0088] For how to obtain the score vector of the benchmark test set, please refer to the above Figure 1 The corresponding embodiments will not be described in detail here.
[0089] S202, performing multiple rounds of random selection from the benchmark test set to obtain multiple candidate evaluation subsets.
[0090] Among them, each candidate evaluation subset is independently selected from the benchmark test set.
[0091] S203 , respectively calculating the difference in score vector between each candidate evaluation subset and the benchmark test set to obtain a plurality of differences.
[0092] For how to calculate the difference in score vectors between the candidate evaluation subset and the benchmark test set, please refer to the above Figure 1 The corresponding embodiments will not be described in detail here.
[0093] S204: taking the candidate evaluation subset corresponding to the minimum difference among the multiple differences as the target evaluation subset.
[0094] By performing multiple rounds of selection in the benchmark test set, the target evaluation subset is finally determined. Then, when evaluating the model, the target evaluation subset can be directly used for evaluation testing, thereby reducing the use of computing resources and lowering computing overhead and costs.
[0095] Furthermore, by selecting the alternative evaluation subset with the smallest difference from multiple rounds of randomly selected alternative evaluation subsets as the target evaluation subset, it can ensure that the evaluation error is small when using the target evaluation subset to evaluate the performance of the large language model.
[0096] In one embodiment, after multiple rounds of selection, the target evaluation subset is obtained, and the following steps can be performed: Figure 3 The semantic evaluation processing process S301-S303 of the target evaluation subset is shown to evaluate the semantic distribution of the target evaluation subset.
[0097] S301, obtaining a semantic embedding representation of a target evaluation subset obtained after a preset round of selection processing to obtain a first embedding representation.
[0098] The embodiments of this application do not limit the encoding method used to obtain the semantic embedding representation of the target evaluation subset. For example, Sentence-BERT (Sentence Embeddings using Siamese BERT-Networks) can be used to encode the target evaluation subset to obtain the first embedding representation.
[0099] S302: Obtain a semantic embedding representation of the benchmark test set to obtain a second embedding representation.
[0100] It should be noted that, when encoding the benchmark test set to obtain the second embedding representation, the encoding method needs to be the same as the encoding method of the target evaluation subset.
[0101] S303: Calculate the distance between the first embedding representation and the second embedding representation. The distance is used to represent the semantic similarity between the target evaluation subset obtained after the preset round of selection processing and the benchmark test set.
[0102] The embodiments of the present application do not limit how to calculate the distance between the first embedding representation and the second embedding representation. For example, the Wasserstein distance between the first embedding representation and the second embedding representation can be calculated, and the semantic similarity between the target evaluation subset and the benchmark test set can be represented by this distance. The larger the Wasserstein distance, the lower the semantic similarity between the target evaluation subset and the benchmark test set.
[0103] The semantic similarity between the target evaluation subset and the benchmark test set can be used to judge the quality of the target evaluation subset, and then determine whether the target evaluation subset needs to be adjusted so that the target evaluation subset can better represent the benchmark test set.
[0104] In order to better reflect the advantages of the technical solutions in the embodiments of the present application, the following will be described in conjunction with specific experimental results. Figure 4-Figure 6 As shown, Figure 4-Figure 6 The evaluation errors of different evaluation subset selection methods on three benchmark test sets, MMLU, HellaSwag, and GSM8K, are respectively shown. Among them, sample is a random sampling method to select the evaluation subset, pca and enc (autoencoder) both correspond to the embedding vector of the question and then clustering to select the evaluation subset. The dimensionality reduction method corresponding to pca is principal component analysis dimensionality reduction, and the dimensionality reduction method corresponding to enc is autoencoder dimensionality reduction. bestsample represents Figure 2 In the corresponding embodiment, the evaluation subset selection method, anneal represents Figure 1 The evaluation subset selection method in the corresponding embodiment. Copress Rate is the compression ratio. "Performance Est, Error" is the evaluation error of the performance estimation, which represents the difference in the score vector between the selected evaluation subset and the benchmark test set.
[0105] from Figure 4-Figure 6 As can be seen in this application Figure 1 and Figure 2 The evaluation subset selection method in the corresponding embodiment has lower evaluation error and higher stability.
[0106] The Wasserstein distances between the evaluation subsets selected by different evaluation subset selection methods and the benchmark test set at different compression ratios are shown in Table 1 below.
[0107] Table 1
[0108]
[0109] It can be seen from Table 1 that, under the premise of keeping the compression ratio consistent, the evaluation subset selected by the anneal method has the smallest Wasserstein distance when the compression ratio is 0.05. At compression ratios of 0.12, 0.19, 0.26, 0.33 and 0.40, the evaluation subset selected by the anneal method is slightly larger in Wasserstein distance than the evaluation subset selected by the samle method. However, judging from the Wasserstein distance under different compression ratios, the evaluation subset selected by the evaluation subset selection method in this application has a high similarity in semantic distribution with the benchmark test set, avoiding the semantic deviation problem caused by compression.
[0110] Based on the same inventive concept, the present application also provides two evaluation subset selection devices, as described in the following embodiments. Since the principles of the device embodiments are similar to those of the above-mentioned method embodiments, the implementation of the device embodiments can refer to the implementation of the above-mentioned method embodiments, and the repeated parts will not be repeated.
[0111] Figure 7 A schematic diagram of an evaluation subset selection device in an embodiment of the present application is shown as follows: Figure 7 As shown, the device includes: a first acquisition module 71, which is used to obtain a score vector of a benchmark test set, where the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models; a first selection module 72, which is used to perform a preset round of selection processing from the benchmark test set, where each round of selection processing includes: selecting an alternative evaluation subset from the benchmark test set, and calculating the difference in the score vector between the alternative evaluation subset and the benchmark test set to obtain an alternative difference; subtracting the reference difference in this round of selection processing from the alternative difference to obtain a difference difference value The reference difference in this round of selection is the difference in the score vector between the target evaluation subset selected in the previous round and the benchmark test set. If this round of selection is the first round, the target evaluation subset selected in the previous round is the initial evaluation subset selected from the benchmark test set. Based on the difference value, the target evaluation subset of this round is selected from the target evaluation subset selected in the previous round and the alternative evaluation subset. When the difference value is not greater than 0, the probability of the alternative evaluation subset being selected is negatively correlated with the difference value. When the difference value is greater than 0, the probability of the alternative evaluation subset being selected is 1.
[0112] In some embodiments, the first selection module 72 is used to randomly select a preset number of evaluation questions from the benchmark test set; use the preset number of evaluation questions to replace the same number of evaluation questions in the target evaluation subset selected in the previous round, and use the subset obtained after the replacement process as the alternative evaluation subset.
[0113] In some embodiments, the device also includes: a semantic analysis module, used to obtain the semantic embedding representation of the target evaluation subset obtained after a preset round of selection processing to obtain a first embedding representation; obtain the semantic embedding representation of the benchmark test set to obtain a second embedding representation; calculate the distance between the first embedding representation and the second embedding representation, and the distance is used to represent the semantic similarity between the target evaluation subset obtained after the preset round of selection processing and the benchmark test set.
[0114] Figure 8 A schematic diagram of another evaluation subset selection device in an embodiment of the present application is shown. Figure 8 As shown, the device includes: a second acquisition module 81, used to obtain a score vector of a benchmark test set, where the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set through multiple language models; a second selection module 82, used to perform multiple rounds of random selection from the benchmark test set to obtain multiple alternative evaluation subsets; respectively calculate the difference between each alternative evaluation subset and the benchmark test set in the score vector to obtain multiple differences; and use the alternative evaluation subset corresponding to the minimum difference among the multiple differences as the target evaluation subset.
[0115] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0116] Refer to the following Figure 9 hereinafter, an electronic device 900 according to this embodiment of the present application is described. Figure 9 The electronic device 900 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0117] like Figure 9 As shown, electronic device 900 is implemented as a general-purpose computing device. Components of electronic device 900 may include, but are not limited to, at least one processing unit 910, at least one storage unit 920, and a bus 930 connecting various system components (including storage unit 920 and processing unit 910).
[0118] The storage unit stores program code, which can be executed by the processing unit 910, so that the processing unit 910 performs the steps of various exemplary embodiments of the present application described in the "Exemplary Method" section above. For example, the processing unit 910 can perform the following steps of the above method embodiment: S101-S102, S201-S204, and S301-S303.
[0119] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 9201 and / or a cache memory unit 9202 , and may further include a read-only memory unit (ROM) 9203 .
[0120] The storage unit 920 may also include a program / utility 9204 having a set (at least one) of program modules 9205, such program modules 9205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0121] Bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0122] The electronic device 900 can also communicate with one or more external devices 940 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 900, and / or any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 950. Furthermore, the electronic device 900 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0123] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0124] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer program product, which includes: a computer program, which implements the above-mentioned evaluation subset selection method when executed by a processor.
[0125] In an exemplary embodiment of the present application, a computer-readable storage medium is further provided, which may be a readable signal medium or a readable storage medium. The computer-readable storage medium stores a program product capable of implementing the above-mentioned method of the present application.
[0126] In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present application described in the above "Exemplary Method" section of this specification.
[0127] More specific examples of computer-readable storage media in the present application may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0128] In this application, a computer-readable storage medium may include a data signal transmitted in baseband or as part of a carrier wave, which carries readable program code. Such a transmitted data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0129] Optionally, program code embodied on a computer readable storage medium can be transmitted by way of electromagnetic signals, such as using radio frequency (RF) signals, infrared signals, etc., and / or the like. Computer readable storage medium then comprises a non-transitory computer readable medium.
[0130] In one or more embodiments, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over a computer-readable medium as one or more instructions or code. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. In this manner, a computer-readable medium can take many forms, including but not limited to, a tangible floppy disk, a tangible compact disk, tangible tape, tangible memory chip, a tangible application- specific integrated circuit (ASIC), a tangible programmable logic device (PLD), a tangible ROM, a tangible RAM, a tangible flash memory, an electrical signal via a tangible electrical connection, a tangible electrical connection comprising a tangible electrical pin, a tangible optical fiber, and / or the like.
[0131] It should be noted that although several modules or units of devices for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. Indeed, features and functions of two or more modules or units described above can be embodied in one module or unit according to embodiments of the application. Conversely, features and functions of one module or unit described above can be further divided into multiple modules or units embodied.
[0132] Moreover, although the various steps of the methods in the present application are described in a particular order in the figures, this is not required or implied as to the order of execution of the steps, nor is it required that all of the steps be executed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, one step can be broken into multiple steps, etc.
[0133] From the above description of the embodiments, those skilled in the art will readily perceive that the example embodiments described herein can be practiced by various other methods than those specifically described. Thus, the present embodiments are not limited by the methods described above, which are presented as examples only but rather are limited only by the claims. Furthermore, from the above description of the embodiments, those skilled in the art will perceive various modifications and changes, which they are encouraged to effect and are within the spirit of the application. It is therefore desired that what is claimed be broadly construed, interpreted and determined.
[0134] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
Claims
1. A method for selecting an evaluation subset, characterized in that: include: Obtaining a score vector for a benchmark test set, where the score vector for the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models; Performing a predetermined round of selection process from the benchmark test set, each round of selection process comprising: Selecting an alternative evaluation subset from the benchmark test set, and calculating the difference in score vector between the alternative evaluation subset and the benchmark test set to obtain an alternative difference; Subtract the reference difference in the current round of selection from the candidate difference to obtain a difference difference value, where the reference difference in the current round of selection is the difference in score vectors between the target evaluation subset selected in the previous round and the benchmark test set. If this round of selection is the first round, the target evaluation subset selected in the previous round is the initial evaluation subset selected from the benchmark test set; Based on the difference value, the target evaluation subset of this round is selected from the target evaluation subset selected in the previous round and the alternative evaluation subset. When the difference value is not greater than 0, the probability of the alternative evaluation subset being selected is negatively correlated with the difference value. When the difference value is greater than 0, the probability of the alternative evaluation subset being selected is 1.
2. The evaluation subset selection method according to claim 1, characterized in that: The selecting of a candidate evaluation subset from the benchmark test set includes: Randomly selecting a preset number of evaluation questions from the benchmark test set; The preset number of evaluation topics is used to replace the same number of evaluation topics in the target evaluation subset selected in the previous round, and the subset obtained after the replacement process is used as the candidate evaluation subset.
3. The evaluation subset selection method according to claim 2, characterized in that: The preset number is 1.
4. The evaluation subset selection method according to claim 1, characterized in that: The probability of the candidate evaluation subset being selected is shown in the following formula: T=αT old ; Wherein, ΔQ is the difference; T is the annealing temperature in this round of selection process; T old is the annealing temperature in the previous round of selection treatment, and this round of selection treatment is the first round, then T old is the preset annealing temperature; α is a preset attenuation factor greater than 0 and less than 1.
5. The evaluation subset selection method according to claim 1, characterized in that: Also includes: Obtaining a semantic embedding representation of the target evaluation subset obtained after the preset round of selection processing to obtain a first embedding representation; Obtaining a semantic embedding representation of the benchmark test set to obtain a second embedding representation; A distance between the first embedding representation and the second embedding representation is calculated, where the distance is used to represent the semantic similarity between the target evaluation subset obtained after the preset round of selection processing and the benchmark test set.
6. A method for selecting an evaluation subset, characterized in that: include: Obtaining a score vector for a benchmark test set, where the score vector for the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models; Perform multiple rounds of random selection from the benchmark test set to obtain multiple candidate evaluation subsets; Calculating the difference in score vector between each candidate evaluation subset and the benchmark test set respectively to obtain a plurality of differences; The candidate evaluation subset corresponding to the minimum difference among the multiple differences is used as the target evaluation subset.
7. An evaluation subset selection device, characterized in that: include: A first acquisition module is configured to acquire a score vector of a benchmark test set, where the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models; The first selection module is used to perform preset rounds of selection processing from the benchmark test set, and each round of selection processing includes: selecting an alternative evaluation subset from the benchmark test set, and calculating the difference in score vector between the alternative evaluation subset and the benchmark test set to obtain an alternative difference; subtracting the reference difference in this round of selection processing from the alternative difference to obtain a difference difference value, where the reference difference in this round of selection processing is the difference in score vector between the target evaluation subset selected in the previous round and the benchmark test set. If this round of selection processing is the first round, the target evaluation subset selected in the previous round is the initial evaluation subset selected from the benchmark test set; based on the difference difference value, the target evaluation subset of this round is selected from the target evaluation subset selected in the previous round and the alternative evaluation subset. When the difference difference value is not greater than 0, the probability of the alternative evaluation subset being selected is negatively correlated with the difference difference value. When the difference difference value is greater than 0, the probability of the alternative evaluation subset being selected is 1.
8. An evaluation subset selection device, characterized in that: include: A second acquisition module is configured to acquire a score vector of a benchmark test set, where the score vector of the benchmark test set is constructed based on score data obtained by predicting the benchmark test set using multiple language models; A second selection module is used to perform multiple rounds of random selection from the benchmark test set to obtain multiple candidate evaluation subsets; Calculating the difference in score vector between each candidate evaluation subset and the benchmark test set respectively to obtain a plurality of differences; The candidate evaluation subset corresponding to the minimum difference among the multiple differences is used as the target evaluation subset.
9. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the evaluation subset selection method according to any one of claims 1 to 6 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the evaluation subset selection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Network detection method, device, equipment and medium
CN112995222A
Deep learning test input selection method based on multi-objective optimization
CN114721934A
Model parameter alignment device, method for the same and program
JP2013008200A