Large model evaluation method and device based on multi-judgment model, medium and equipment

By obtaining and optimizing the scoring data of the multi-referee model, determining the target answers and adjusting the weight of the referee model, the problem of inaccurate evaluation results in large-scale model evaluation is solved, and a more accurate and reliable comprehensive evaluation value is achieved.

CN119917829APending Publication Date: 2025-05-02CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411982072.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the evaluation of large-scale model, a single referee model cannot ensure the highest accuracy rate in all scenarios. The evaluation results of multiple referee models may differ, and simply taking the mean cannot accurately reflect the true ability of the large model.

Method used

By obtaining the scoring data set of the preferred referee model, the target answer that makes the score consistency less than the preset threshold is determined, and a new preferred referee model is introduced for scoring until the score consistency reaches or exceeds the threshold. Based on this, the comprehensive evaluation value of the large model to be evaluated is calculated.

Benefits of technology

It improves the accuracy and robustness of the large-scale model evaluation results, enhances the credibility of the evaluation results, and ensures that the evaluation results can automatically adapt to the fluctuations in the performance of the referee model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917829A_ABST
    Figure CN119917829A_ABST
Patent Text Reader

Abstract

The invention discloses a large model evaluation method and device based on a multi-judgment model, a medium and equipment. The method comprises the steps of obtaining a score data set corresponding to a preferred judgment model; according to the score data set, determining a target answer enabling the score consistency of the preferable judgment model to be smaller than a preset threshold value; aiming at the target answer, introducing a target number of newly added preferable judgment models for scoring until the score consistency of the newly added preferable judgment models is greater than or equal to a preset threshold value; based on the newly-added score data set of the newly-added optimal judgment model, determining a weight corresponding to the newly-added optimal judgment model and an optimal score data set of the to-be-evaluated large model; calculating a comprehensive evaluation value of the to-be-evaluated large model according to the optimal score data set and the weight of each newly added judgment model; in this way, scoring data with large errors can be eliminated, the model weight is adjusted to ensure that the evaluation result automatically adapts to the performance fluctuation of the judgment model, and the accuracy of the overall comprehensive evaluation value is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model evaluation, and more specifically, particularly relates to a large model evaluation method, device, medium and equipment based on a multi-referee model. Background Art

[0002] When evaluating the capabilities of a large model, a single referee model cannot guarantee the highest accuracy when evaluating each scenario. There may be inaccurate evaluation results in some scenarios. Introducing multiple referee models is a solution, but the evaluation performance of multiple referees may vary. Simply taking the average of their evaluation results may not accurately reflect the true capability level of the large model to be evaluated.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0004] The purpose of the present invention is to provide a large model evaluation method, device, medium and equipment based on a multi-referee model to solve the problem of inaccurate evaluation results in related technologies.

[0005] According to one aspect of an embodiment of the present application, a large model evaluation method based on a multi-referee model is provided, the method comprising:

[0006] Obtain a scoring data set corresponding to the preferred referee model; the scoring data set is a data set consisting of the scores of each answer data given by the preferred referee model;

[0007] Determine, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold;

[0008] For the target answer, introduce a target number of newly added preferred referee models for scoring until the scoring consistency of the newly added preferred referee models is greater than or equal to the preset threshold;

[0009] Based on the newly added scoring data set of the newly added optimal referee model, determine the weight corresponding to the newly added optimal referee model and the optimal scoring data set of the large model to be evaluated;

[0010] Based on the preferred scoring data set and the weights of each newly added referee model, the comprehensive evaluation value of the large model to be evaluated is calculated.

[0011] In some embodiments, based on the scoring data set, determining a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold includes: calculating the scoring consistency of the preferred referee model for each answer data based on the scoring data set; if the scoring consistency corresponding to a certain answer data is less than the preset threshold, determining the answer data as the target answer.

[0012] In some embodiments, based on the newly added scoring data set of the newly added preferred referee model, the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated are determined, including: based on the newly added scoring data set, calculating the absolute value error of the score of each newly added preferred referee model for each answer data; dividing the absolute value error of the score of each newly added preferred referee model by the absolute value error of the total score to obtain the error rate corresponding to each newly added preferred referee model; according to the error rate corresponding to the newly added preferred referee model, calculating the weight corresponding to each newly added preferred referee model; deleting the newly added preferred referee model with an error rate greater than the error threshold to obtain the target preferred referee model of the large model to be evaluated, and using the scoring data set of the target preferred referee model as the preferred scoring data set.

[0013] In some embodiments, the weight corresponding to each newly added preferred referee model is calculated based on the error rate corresponding to the newly added preferred referee model, including: according to the error rate corresponding to the newly added preferred referee model, the error rate of each newly added preferred referee model for the answer data is counted; the inverse of the error rate corresponding to each newly added preferred referee model is added to obtain the inverse of the total error rate; the inverse of the error rate of each newly added preferred referee model is divided by the inverse of the total error rate to obtain the weight corresponding to each newly added preferred referee model.

[0014] In some embodiments, a comprehensive evaluation value of the large model to be evaluated is calculated based on the preferred scoring data set and the weights of each newly added post-judgment model, including: calculating the weighted score corresponding to each answer data based on the preferred scoring data set and the weights of each newly added post-judgment model; calculating the average score of the weighted scores corresponding to all answer data, and using the average score as the comprehensive evaluation value of the large model to be evaluated.

[0015] In some embodiments, before obtaining the scoring data set corresponding to the preferred referee model, the method also includes: obtaining an evaluation data set of the large model to be evaluated; each evaluation data in the evaluation data set includes a question and a reference answer; inputting the question in the evaluation data set into the large model to be evaluated for reasoning to obtain an answer data set; randomly extracting an answer test subset from the evaluation data set and the answer data set; selecting at least one preferred referee model from the candidate referee models based on the scores of the answer test subsets by each candidate referee model in the large model to be evaluated; inputting the answer test subsets into each preferred referee model to obtain the score of each preferred referee model for each answer data; and collecting the scores to obtain a scoring data set for the preferred referee model.

[0016] In some embodiments, at least one preferred referee model is selected from the candidate referee models based on the scores of the answer test subsets given by each candidate referee model in the large model to be evaluated, including: calculating the model evaluation parameters of the candidate referee models based on the scores of the answer test subsets given by each candidate referee model in the large model to be evaluated; the model evaluation parameters include at least the accuracy, precision, recall and F1 value of the candidate referee models; sorting the candidate referee models according to the size of the model evaluation parameters to obtain a candidate model queue of the preferred referee model; and selecting a specified number of candidate referee models from the candidate model queue as preferred referee models in the order of the model evaluation parameters from large to small.

[0017] According to one aspect of an embodiment of the present application, a large model evaluation device based on a multi-referee model is provided, the device comprising:

[0018] The data set acquisition module is used to acquire the scoring data set corresponding to the preferred referee model; the scoring data set is a data set consisting of the scores of each answer data by the preferred referee model;

[0019] An answer determination module is used to determine a target answer that makes the score consistency of the preferred referee model less than a preset threshold value based on the score data set;

[0020] A model introduction module is used to introduce a target number of newly added preferred referee models for scoring the target answer until the score consistency of the newly added preferred referee models is greater than or equal to a preset threshold;

[0021] A parameter optimization module, used to determine the weight corresponding to the newly added optimal referee model and the optimal scoring data set of the large model to be evaluated based on the newly added scoring data set of the newly added optimal referee model;

[0022] The evaluation calculation module is used to calculate the comprehensive evaluation value of the large model to be evaluated based on the preferred scoring data set and the weights of each newly added referee model.

[0023] According to one aspect of an embodiment of the present application, a computer medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the large model evaluation method based on a multi-judge model provided in any embodiment of the present application is implemented.

[0024] According to one aspect of an embodiment of the present application, there is provided an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor executes the executable instructions to enable the electronic device to implement the large model evaluation method based on a multi-judge model provided in any embodiment of the present application.

[0025] In the technical solution of the present application, a scoring data set corresponding to a preferred referee model is obtained; based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold is determined; for the target answer, a target number of newly added preferred referee models are introduced for scoring until the scoring consistency of the newly added preferred referee models is greater than or equal to the preset threshold; based on the newly added scoring data set of the newly added preferred referee model, the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated are determined; based on the preferred scoring data set and the weights of each newly added referee model, a comprehensive evaluation value of the large model to be evaluated is calculated; in this way, a more accurately scored preferred referee model is introduced for the target answer, and scoring data with large errors are eliminated, and the weights of each preferred referee model are adjusted to ensure that the evaluation results can automatically adapt to the performance fluctuations of the referee model, and the comprehensive evaluation value of the large model is calculated based on more accurate scoring data and model weights, thereby improving the accuracy of the overall comprehensive evaluation value and enhancing the robustness and credibility of the large model evaluation results.

[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0028] Figure 1 The flowchart of the large model evaluation method based on the multi-judge model provided in one embodiment of the present application is schematically shown.

[0029] Figure 2 The scoring data table of the preferred judging model provided in one embodiment of the present application is schematically shown.

[0030] Figure 3The flowchart of the large model evaluation method based on the multi-judge model provided in one embodiment of the present application is schematically shown.

[0031] Figure 4 The flowchart of the large model evaluation method based on the multi-judge model provided in one embodiment of the present application is schematically shown.

[0032] Figure 5 The structural block diagram of a large model evaluation device based on a multi-judge model provided in one embodiment of the present application is schematically shown.

[0033] Figure 6 The structural block diagram of an electronic device provided by an embodiment of the present application is schematically shown.

[0034] Figure 7 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0036] In addition, the features, structures or characteristics described in the present application may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.

[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0039] The referee model mentioned in this application refers to a model that evaluates the quality of answers generated by the large model to be evaluated. The large model to be evaluated is a model that can generate answers to user questions. The method provided in this application evaluates the answer quality of the large model to be evaluated based on multiple referee models, making the evaluation results of the large model more accurate and solving the problem in related technologies that the real ability level of the large model to be evaluated cannot be accurately reflected. Figure 1 As shown, the present application provides a large model evaluation method based on a multi-referee model including S110 to S150, and the specific process is as follows.

[0040] S110, obtaining a scoring data set corresponding to the preferred referee model.

[0041] Specifically, the preferred referee model refers to a referee model selected from the candidate referee models based on the scoring accuracy of the referee model, and the candidate referee model is the referee model initially associated with a certain answer data, and the associated candidate referee model for each answer data can be set by the user based on experience. The answer data includes reference answers and output answers. The scoring data set is a data set composed of the scores of each answer data given by the preferred referee model. By processing the scores of the answer data by each referee model through the steps provided in this application, the scoring accuracy of each referee model can be determined, so as to select a more suitable referee model to evaluate the large model to be evaluated and obtain a more accurate evaluation result.

[0042] In some embodiments, before obtaining the scoring data set corresponding to the preferred referee model, the method also includes: obtaining an evaluation data set of the large model to be evaluated; each evaluation data in the evaluation data set includes a question and a reference answer; inputting the question in the evaluation data set into the large model to be evaluated for reasoning to obtain an answer data set; randomly extracting an answer test subset from the evaluation data set and the answer data set; selecting at least one preferred referee model from the candidate referee models based on the scores of the answer test subsets by each candidate referee model in the large model to be evaluated; inputting the answer test subsets into each preferred referee model to obtain the score of each preferred referee model for each answer data; and collecting the scores to obtain a scoring data set for the preferred referee model.

[0043] Specifically, the questions in the evaluation data set are pre-set questions, and the reference answers are accurate answers predetermined by the user based on the questions. When the questions are input into the large model to be evaluated for reasoning, the output answers can be obtained. All the output answers are put into a data set to form an answer data set. Whether the output answers are accurate depends on the performance of the large model to be evaluated. The answer test subset includes reference answers and output answers, where each reference answer and its corresponding output answer constitute an answer data; the reference answer and its corresponding output answer refer to the reference answer and output answer corresponding to a certain question. The difference in the scores of the candidate referee model on the answer data also reflects the evaluation accuracy of the referee model. Through the difference in the scores of the reference answer and the output answer in each answer data by each candidate referee model, some referee models with more accurate scores on the answer data can be preliminarily selected, thereby obtaining the preferred referee model. Then, the answer data in the answer test subset is input into the preferred referee model, and the scores of each answer data by the preferred referee model can be obtained again. The scores of each answer data by each preferred referee model are collected into a set to obtain the scoring data set.

[0044] In some embodiments, at least one preferred referee model is selected from the candidate referee models based on the scores of the answer test subsets given by each candidate referee model in the large model to be evaluated, including: calculating the model evaluation parameters of the candidate referee models based on the scores of the answer test subsets given by each candidate referee model in the large model to be evaluated; the model evaluation parameters include at least the accuracy, precision, recall and F1 value of the candidate referee models; sorting the candidate referee models according to the size of the model evaluation parameters to obtain a candidate model queue of the preferred referee model; and selecting a specified number of candidate referee models from the candidate model queue as preferred referee models in the order of the model evaluation parameters from large to small.

[0045] Specifically, the model evaluation parameters are used to reflect the accuracy of the candidate referee model in evaluating the large model. The accuracy of the candidate referee model refers to the accuracy of the candidate referee model in judging whether the output answer of the large model reasoning is correct or not. When the candidate referee model judges that the output answer of the large model reasoning is correct, it is recorded as 1, and when the candidate referee model judges that the output answer of the large model reasoning is wrong, it is recorded as 0. By comparing the accuracy of each candidate referee model, it can be used as a basis for selecting the preferred referee model. Accuracy of candidate referee models where n correct represents the number of times the candidate referee model judges the output answer to be correct, and N represents the number of output answers. The accuracy of the candidate referee model refers to the proportion of the output answers that are actually correct among the ones that the candidate referee model judges to be correct. Assume that the number of times the candidate referee model judges to be correct is n predict_true , where the number of correct answers is n true_positive , then the accuracy of the candidate referee model The recall rate of the candidate referee model refers to the proportion of all truly correct answers of the big model that the candidate referee model can correctly judge. Assume that the number of all truly correct big model answers is n actual_true , the number of times the referee model correctly judges as correct is n true_positive , then the recall rate The F1 value of the candidate referee model is the harmonic mean of precision and recall, which can comprehensively reflect the performance of the referee model. By combining the accuracy, precision, recall and F1 value of each candidate referee model, the model evaluation parameters of each candidate referee model can be obtained. Further, the weights of the above parameters can be set according to the degree of influence of accuracy, precision, recall and F1 value on the model evaluation parameters, and finally the model evaluation parameter Score = Accuracy × α1 + Precesion × α2 + Recall × α3 + F1 × α4 is obtained, where α1, α2, α3 and α4 represent the weights of accuracy, precision, recall and F1 value respectively. For example, the candidate referee models include Model 1, Model 2, Model 3 and Model 4, and the model evaluation parameters corresponding to the above models are 30, 34, 35 and 28. The model evaluation parameters are arranged from large to small to obtain: 35, 34, 30 and 28, then the candidate model queue is: Model 3, Model 2, Model 1 and Model 4. Assuming that the specified number is 3, the preferred referee models include: Model 3, Model 2 and Model 1.

[0046] S120. Determine, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold.

[0047] Specifically, in some embodiments, based on the scoring data set, determining the target answer that makes the scoring consistency of the preferred referee model less than a preset threshold includes: calculating the scoring consistency of the preferred referee model for each answer data based on the scoring data set; if the scoring consistency corresponding to a certain answer data is less than the preset threshold, determining the answer data as the target answer. Scoring consistency refers to the proportion of the same answer data scored by various preferred referee models. Specifically, the evaluation consistency of the preferred referee model can be determined by the variance, standard deviation or other parameters of the evaluation data difference of the scores of various preferred referee models for the same answer data. The preset threshold is the critical value for determining whether the scoring consistency meets the requirements. Figure 2 As shown, the scores of preferred referee models 1, 2 and 3 for answer 1 are 8, 7 and 8 respectively, and the score consistency of preferred referee model 1, preferred referee model 2 and preferred referee model 3 is 93%; for answer 3, the score consistency of preferred referee model 1, preferred referee model 2 and preferred referee model 3 is 76%, and the preset threshold is 90%, so the target answer is answer data 3.

[0048] S130. For the target answer, introduce a target number of newly added preferred referee models for scoring until the score consistency of the newly added preferred referee models is greater than or equal to a preset threshold.

[0049] Specifically, Figure 2 As shown, when the target answer is answer 3, based on the preferred referee models 1 to 3, a new preferred referee model 4 is introduced to score answer 3. However, the score consistency of referee models 1 to 4 is still less than the preset threshold, then the new preferred referee model 5 is introduced to score answer 3. If the score consistency of referee models 1 to 5 is greater than the preset threshold, the introduction of the new preferred referee model is stopped. At this time, the target number is 2.

[0050] S140. Based on the newly added scoring data set of the newly added preferred referee model, determine the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated.

[0051] Specifically, the newly added scoring dataset includes the scoring dataset and the newly added scoring dataset. Figure 2 As shown in the figure, although the scoring results of the priority referee model for the answer data are relatively accurate, there are still scoring errors. For example, the scoring of answer 3 by model 1 is quite different from that of other models. This application can reduce the impact of the misjudgment of these referee models on the final evaluation results by adjusting the weights of each newly added preferred referee model and optimizing the newly added scoring data set to obtain the preferred scoring data.

[0052] In some embodiments, Figure 3 As shown, based on the newly added scoring data set of the newly added preferred referee model, the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated are determined, including S1410 to S1440, and the specific process is as follows.

[0053] S1410. Based on the newly added scoring data set, calculate the absolute value error of the score of each newly added optimal referee model for each answer data.

[0054] For example, Figure 2 As shown, for answer 3, the newly added preferred referee models include referee models 1, 2, 3, 4 and 5. The average score of the above referee models is 4.8, and the absolute value errors between the scores of each referee model and the average score are |-3.8|, 1.2, 3.2, |-1.8| and 1.2 respectively.

[0055] S1420. Divide the absolute value error of the score of each newly added optimal referee model by the absolute value error of the total score to obtain the error rate corresponding to each newly added optimal referee model.

[0056] Specifically, assume that the absolute value error of the score of each newly added optimal referee model is x, and the absolute value error of the total score is x all , then the error rate is x / x all Based on this, it can be calculated that the error rates corresponding to each newly added optimal referee model are 0.79, 0.25, 0.67, 0.38 and 0.25 respectively.

[0057] S1430. According to the error rate corresponding to the newly added optimal referee model, the weight corresponding to each newly added optimal referee model is calculated.

[0058] Specifically, Figure 4 As shown, according to the error rate corresponding to the newly added preferred referee model, the weights corresponding to each newly added preferred referee model are calculated, including S1431 to S1433, and the specific process is as follows.

[0059] S1431. According to the error rate corresponding to the newly added optimal referee model, the error rate of each newly added optimal referee model on the answer data is counted.

[0060] Specifically, assuming that the error threshold is 0.3, for answer 3, the error rates of referee models 1, 3, and 4 are greater than the error threshold, so referee models 1, 3, and 4 misjudge answer 3. And for answer 4, according to the method described above, the average score of referee models 1, 2, 3, and 4 is calculated to be 6, and the absolute value errors of the scores of each referee model and the average score are 2, 0, 2, and 0, respectively, and the error rates are 0.33, 0, 0.33, and 0, respectively, so referee models 1 and 3 misjudge answer 4. Then it can be obtained that the number of misjudgements of referee models 1 to 5 is 2, 0, 2, 1, and 0, respectively. Combined with the number of times each newly added referee model scores the answer data, it can be obtained that the misjudge rates of referee models 1 to 5 are 2 / 4=0.5, 0, 2 / 4=0.5, 1 / 2=0.5, and 0.

[0061] S1432. Add the inverse of the error rate corresponding to each newly added optimal referee model to obtain the inverse of the total error rate.

[0062] Exemplarily, the inverses of the error rates of referee models 1 to 5 are 2, 1 / 0, 2, 2 and 1 / 0, respectively, where 1 / 0 is infinite. In actual situations, the denominator can be set to a larger number, such as 100, so the inverse of the total error rate of referee models 1 to 5 is 2+100+2+2+100=206.

[0063] S1433. Divide the inverse of the error rate of each newly added optimal referee model by the inverse of the total error rate to obtain the weight corresponding to each newly added optimal referee model.

[0064] Exemplarily, the inverse of each misjudgment rate is divided by the inverse of the total misjudgment rate, and the weights of referee models 1 to 5 are obtained as 2 / 206, 100 / 206, 2 / 206, 2 / 206 and 100 / 206.

[0065] S1440. Delete the newly added optimal referee model whose error rate is greater than the error threshold, obtain the target optimal referee model of the large model to be evaluated, and use the scoring data set of the target optimal referee model as the preferred scoring data set.

[0066] Specifically, according to Figure 2 By calculating the error rate of each referee model, it can be obtained that for answer 1, the error rates of referee models 1 to 3 are all less than the error threshold. For answer 2, the error rates of referee models 1 to 3 are also less than the error threshold. For answer 3, the error rates of referee models 2 and 5 are less than the error threshold, and the error rates of referee models 1, 3 and 4 are greater than the error threshold. Therefore, the target preferred referee model includes referee models 2 and 5, and the scoring data of answer 3 in the preferred scoring data set includes the scores of referee models 2 and 5. For answer 4, the error rates of referee models 2 and 4 are less than the error threshold, and the error rates of referee models 1 and 3 are greater than the error threshold. Therefore, the target preferred referee model includes referee models 2 and 4, and the scoring data of answer 4 in the preferred scoring data set includes the scores of referee models 2 and 4. It should be understood that for the priority referee model with an error rate less than or equal to the error threshold, it means that the scoring of each priority referee model for the answer meets the accuracy requirement, so the scoring data of the priority referee model that meets the requirements for the answer is used as the scoring data for calculating the large model to be evaluated.

[0067] S150. Calculate the comprehensive evaluation value of the large model to be evaluated based on the preferred scoring data set and the weights of each newly added referee model.

[0068] Specifically, in some embodiments, S150 specifically includes: calculating the weighted score corresponding to each answer data according to the preferred scoring data set M and each newly added referee model weight k, for example Figure 2 Based on the data, we can calculate that the weighted score Add1 of answer 1 includes: 8×(2 / 206), 7×(100 / 206) and 8×(2 / 206); the weighted score Add2 of answer 2 includes: 7×(2 / 206), 7×(100 / 206) and 9×(2 / 206); the weighted score Add3 of answer 3 includes: 6×(100 / 206) and 6×(100 / 206); the weighted score Add4 of answer 4 includes: 6×(100 / 206) and 6×(2 / 206). Finally, calculate the mean score of the weighted scores corresponding to all the answer data, and use the mean score as the comprehensive evaluation value of the large model to be evaluated. For example, the mean score Score of answers 1 to 4 above is 均值=(Add1+Add2+Add3+Add4) / 4, and the average score Score 均值 As the comprehensive evaluation value of the large model to be evaluated.

[0069] In the technical solution of the present application, a scoring data set corresponding to a preferred referee model is obtained; based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold is determined; for the target answer, a target number of newly added preferred referee models are introduced for scoring until the scoring consistency of the newly added preferred referee models is greater than or equal to the preset threshold; based on the newly added scoring data set of the newly added preferred referee model, the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated are determined; based on the preferred scoring data set and the weights of each newly added referee model, a comprehensive evaluation value of the large model to be evaluated is calculated; in this way, a more accurately scored preferred referee model is introduced for the target answer, and scoring data with large errors are eliminated, and the weights of each preferred referee model are adjusted to ensure that the evaluation results can automatically adapt to the performance fluctuations of the referee model, and the comprehensive evaluation value of the large model is calculated based on more accurate scoring data and model weights, thereby improving the accuracy of the overall comprehensive evaluation value and enhancing the robustness and credibility of the large model evaluation results.

[0070] like Figure 5 As shown, the present application provides a large model evaluation device based on a multi-referee model, which includes the following modules, as follows.

[0071] The data set acquisition module 510 is used to acquire the scoring data set corresponding to the preferred referee model; the scoring data set is a data set consisting of the scores of each answer data by the preferred referee model.

[0072] The answer determination module 520 is used to determine, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold.

[0073] The model introduction module 530 is used to introduce a target number of newly added preferred referee models for scoring the target answer until the score consistency of the newly added preferred referee models is greater than or equal to a preset threshold.

[0074] The parameter optimization module 540 is used to determine the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated based on the newly added scoring data set of the newly added preferred referee model.

[0075] The evaluation calculation module 550 is used to calculate the comprehensive evaluation value of the large model to be evaluated based on the preferred scoring data set and the weights of each newly added referee model.

[0076] In some embodiments, the answer determination module 520 includes: a scoring evaluation unit, which is used to calculate the scoring consistency of each answer data by the preferred referee model based on the scoring data set; and an answer determination unit, which is used to determine the answer data as the target answer if the scoring consistency corresponding to a certain answer data is less than a preset threshold.

[0077] In some embodiments, the parameter optimization module 540 includes: an error calculation unit, which is used to calculate the absolute value error of the score of each newly added preferred referee model for each answer data based on the newly added scoring data set; an error rate calculation unit, which is used to divide the absolute value error of the score of each newly added preferred referee model by the absolute value error of the total score to obtain the error rate corresponding to each newly added preferred referee model; a weight calculation unit, which is used to calculate the weight corresponding to each newly added preferred referee model according to the error rate corresponding to the newly added preferred referee model; a data set optimization unit, which is used to delete the newly added preferred referee model whose error rate is greater than the error threshold, obtain the target preferred referee model of the large model to be evaluated, and use the scoring data set of the target preferred referee model as the preferred scoring data set.

[0078] In some embodiments, the weight calculation unit is specifically used to count the error rates of the newly added preferred referee models for the answer data according to the error rates corresponding to the newly added preferred referee models; add the inverses of the error rates corresponding to each newly added preferred referee model to obtain the inverse of the total error rate; divide the inverse of the error rate of each newly added preferred referee model by the inverse of the total error rate to obtain the weights corresponding to each newly added preferred referee model.

[0079] In some embodiments, the evaluation calculation module 550 includes: a weighted calculation unit, which is used to calculate the weighted score corresponding to each answer data according to the preferred scoring data set and the weights of each newly added referee model; an evaluation value calculation unit, which is used to calculate the average score of the weighted scores corresponding to all answer data, and use the average score as the comprehensive evaluation value of the large model to be evaluated.

[0080] In some embodiments, the device also includes: an evaluation data set acquisition module, which is used to acquire an evaluation data set of the large model to be evaluated before acquiring a scoring data set corresponding to the preferred referee model; each evaluation data in the evaluation data set includes a question and a reference answer; an answer data set acquisition module, which is used to input the question in the evaluation data set into the large model to be evaluated for reasoning, and obtain an answer data set; an answer test subset acquisition module, which is used to randomly extract an answer test subset from the evaluation data set and the answer data set; a model selection module, which is used to select at least one preferred referee model from the candidate referee models according to the scores of the answer test subsets by each candidate referee model in the large model to be evaluated; a score determination module, which is used to input the answer test subset into each preferred referee model to obtain the score of each preferred referee model for each answer data; a data set collection module, which is used to collect the score data set of the preferred referee model.

[0081] In some embodiments, the model selection module includes: an evaluation parameter calculation unit, which is used to calculate the model evaluation parameters of the candidate referee model according to the scores of each candidate referee model in the large model to be evaluated on the answer test subset; the model evaluation parameters include at least the accuracy, precision, recall rate and F1 value of the candidate referee model; a model sorting unit, which is used to sort the candidate referee models according to the size of the model evaluation parameters to obtain a candidate model queue of the preferred referee model; a model selection unit, which is used to select a specified number of candidate referee models from the candidate model queue as the preferred referee model in the order of the model evaluation parameters from large to small.

[0082] It should be noted that the specific implementation content of the above-mentioned device has been explained in detail in the corresponding method embodiment and will not be repeated here.

[0083] The electronic equipment of this application is introduced below, such as Figure 6 As shown, the present application provides an electronic device 600, which includes: a processor 610 and a memory 620, the memory 620 is used to store executable instructions of the processor; the processor 610 executes the executable instructions to enable the electronic device to implement the large model evaluation method based on the multi-judge model provided by any embodiment of the present application.

[0084] Specifically, the large model evaluation method based on the multi-judge model provided by the present application is stored in the memory 620 of the electronic device, and the large model evaluation method based on the multi-judge model is executed by the processor 610 to obtain a scoring data set corresponding to the preferred judge model; based on the scoring data set, determine the target answer that makes the scoring consistency of the preferred judge model less than a preset threshold; for the target answer, introduce a target number of newly added preferred judge models for scoring until the scoring consistency of the newly added preferred judge models is greater than or equal to the preset threshold; based on the newly added scoring data set of the newly added preferred judge models, determine the weight corresponding to the newly added preferred judge models and the preferred scoring data set of the large model to be evaluated; according to the preferred scoring data set and the weights of each newly added judge model, calculate the comprehensive evaluation value of the large model to be evaluated; in this way, introduce a more accurate preferred judge model for the target answer, eliminate the scoring data with large errors, and adjust the weights of each preferred judge model to ensure that the evaluation result can automatically adapt to the performance fluctuations of the judge model, calculate the comprehensive evaluation value of the large model based on more accurate scoring data and model weights, improve the accuracy of the overall comprehensive evaluation value, and enhance the robustness and credibility of the large model evaluation result.

[0085] It should be known that the specific implementation content of the electronic device in the present application has been explained in detail in the corresponding method embodiment and will not be repeated here.

[0086] Figure 7 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.

[0087] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0088] like Figure 7 As shown, the computer system 700 includes a processor 701, which can be a CPU (Central Processing Unit) or an MCU (Microcontroller Unit). The processor 701 can perform various appropriate actions and processes according to the program stored in the read-only memory 702 (Read-Only Memory, ROM) or the program loaded from the storage part 708 to the random access memory 703 (Random Access Memory, RAM). In the random access memory 703, various programs and data required for system operation are also stored. The processor 701, the read-only memory 702 and the random access memory 703 are connected to each other through a bus 704. The input / output interface 705 (Input / Output interface, i.e., I / O interface) is also connected to the bus 704.

[0089] The following components are connected to the input / output interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.

[0090] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 709, and / or installed from a removable medium 711. When the computer program is executed by the processor 701, various functions defined in the system of the present application are executed.

[0091] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0092] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0093] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0094] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the implementation method according to the present application.

[0095] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.

[0096] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A large model evaluation method based on a multi-judge model, characterized in that: include: Obtain the scoring data set corresponding to the optimal referee model; The scoring data set is a data set consisting of the scores of each answer data given by the preferred referee model; Determining, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold; For the target answer, introduce a target number of newly added preferred referee models for scoring until the scoring consistency of the newly added preferred referee models is greater than or equal to a preset threshold; Based on the newly added scoring data set of the newly added preferred referee model, determine the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated; According to the preferred scoring data set and the weights of each newly added referee model, the comprehensive evaluation value of the large model to be evaluated is calculated.

2. The large model evaluation method based on the multi-judge model as claimed in claim 1, characterized in that: Determining, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold comprises: Calculating the consistency of the score of each answer data by the preferred referee model according to the score data set; If the score consistency corresponding to a certain answer data is less than a preset threshold, the answer data is determined as the target answer.

3. The large model evaluation method based on the multi-judge model as claimed in claim 1, characterized in that: The method of determining the weight corresponding to the newly added optimal referee model and the optimal scoring dataset of the large model to be evaluated based on the newly added scoring dataset of the newly added optimal referee model comprises: Based on the newly added scoring data set, calculate the absolute value error of the score of each newly added optimal referee model for each answer data; Divide the absolute value error of the score of each newly added optimal referee model by the absolute value error of the total score to obtain the error rate corresponding to each newly added optimal referee model; According to the error rate corresponding to the newly added optimal referee model, the weight corresponding to each newly added optimal referee model is calculated; The newly added preferred referee model whose error rate is greater than the error threshold is deleted to obtain the target preferred referee model of the large model to be evaluated, and the scoring data set of the target preferred referee model is used as the preferred scoring data set.

4. The large model evaluation method based on the multi-judge model as claimed in claim 3, characterized in that: The weights corresponding to the newly added optimal referee models are calculated based on the error rates corresponding to the newly added optimal referee models, including: According to the error rate corresponding to the newly added optimal referee model, the error rate of each newly added optimal referee model on the answer data is counted; Add the inverse of the error rate corresponding to each newly added optimal referee model to obtain the inverse of the total error rate; Divide the inverse of the error rate of each newly added optimal referee model by the inverse of the total error rate to obtain the weight corresponding to each newly added optimal referee model.

5. The large model evaluation method based on the multi-judge model according to claim 1, characterized in that: The comprehensive evaluation value of the large model to be evaluated is calculated based on the preferred scoring data set and the weights of each newly added referee model, including: Calculate the weighted score corresponding to each answer data according to the preferred scoring data set and each newly added referee model weight; Calculate the mean score of the weighted scores corresponding to all answer data, and use the mean score as the comprehensive evaluation value of the large model to be evaluated.

6. The large model evaluation method based on the multi-judge model as claimed in claim 1, characterized in that: Before obtaining the scoring data set corresponding to the preferred referee model, the method further includes: Obtaining an evaluation data set of a large model to be evaluated; each evaluation data in the evaluation data set includes a question and a reference answer; Input the questions in the evaluation data set into the large model to be evaluated for reasoning to obtain the answer data set; randomly extract an answer test subset from the evaluation data set and the answer data set; Selecting at least one preferred referee model from the candidate referee models according to the scores of the answer test subset by each candidate referee model in the large model to be evaluated; Inputting the answer test subset into each optimal referee model to obtain the score of each answer data by each optimal referee model; The scores are collected to obtain a score data set of the preferred referee model.

7. The large model evaluation method based on the multi-judge model as claimed in claim 6, characterized in that: The step of selecting at least one preferred referee model from the candidate referee models according to the scores of the answer test subset by the candidate referee models in the large model to be evaluated comprises: Calculate the model evaluation parameters of the candidate referee model according to the scores of the answer test subset by each candidate referee model in the large model to be evaluated; the model evaluation parameters include at least the accuracy, precision, recall and F1 value of the candidate referee model; Sort the candidate referee models according to the size of the model evaluation parameters to obtain a candidate model queue of the preferred referee model; A specified number of candidate referee models are selected from the candidate model queue as preferred referee models in descending order of model evaluation parameters.

8. A large model evaluation device based on a multi-judge model, characterized in that: include: A data set acquisition module is used to acquire a scoring data set corresponding to the optimal referee model; The scoring data set is a data set consisting of the scores of each answer data given by the preferred referee model; An answer determination module, used to determine, based on the scoring data set, a target answer that makes the scoring consistency of the preferred referee model less than a preset threshold; A model introduction module is used to introduce a target number of newly added preferred referee models for scoring the target answer until the score consistency of the newly added preferred referee models is greater than or equal to a preset threshold; A parameter optimization module, used to determine the weight corresponding to the newly added preferred referee model and the preferred scoring data set of the large model to be evaluated based on the newly added scoring data set of the newly added preferred referee model; The evaluation calculation module is used to calculate the comprehensive evaluation value of the large model to be evaluated based on the preferred scoring data set and the weights of each newly added referee model.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the large model evaluation method based on the multi-judge model as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; The processor executes the executable instructions to enable the electronic device to implement the large model evaluation method based on the multi-judge model as described in any one of claims 1 to 7 above.

Citation Information

Cited By

  • Multi-model comparison and tuning method and device for referee model

    CN121834260A