Model evaluation method and device, storage medium, equipment and program product

By adjusting the test text structure and combining open and multi-optional tasks, the performance of pre-trained models is accurately evaluated, which solves the problem of inaccurate evaluation in the existing technology, ensuring the accuracy of model selection and the efficiency of business execution.

CN120163245APending Publication Date: 2025-06-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510220536.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing technology cannot accurately evaluate the model performance in the pre-training stage, which may lead to high evaluation scores but poor actual performance or low evaluation scores but high actual performance, affecting subsequent model selection and user service execution.

Method used

A model evaluation method is proposed. By obtaining test requests, determining the test task, and obtaining the specified text and question-answer sample templates from the sample dataset according to the task type (open or multi-optional) and combining them into the target test text, and inputting the model to be tested for evaluation.

Benefits of technology

By adjusting the test text structure, the real performance of the model to be tested can be more accurately evaluated, and a clear basis for model selection can be provided to ensure that users obtain powerful artificial intelligence models and improve business execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163245A_ABST
    Figure CN120163245A_ABST
Patent Text Reader

Abstract

The invention discloses a model evaluation method and device, a storage medium, equipment and a program product. According to the scheme, the text structure of the text obtained from the sample data set can be adjusted according to the specific test task to obtain the target test text which is more convenient for the to-be-tested model in the pre-training stage to understand, so that the test efficiency is improved when the to-be-tested model is evaluated through the target test text. The real performance of the to-be-tested model can be accurately evaluated, so that a clear basis is provided for follow-up model selection, it can be guaranteed that an artificial intelligence model with strong performance is provided for a user in the follow-up process, and the service execution efficiency of the user based on the artificial intelligence model is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the fields of computer technology and artificial intelligence, and in particular, to a method, device, storage medium, device, and program product for model evaluation. Background Art

[0002] With the continuous development of computer technology, artificial intelligence technology has played a major role in many business scenarios. It not only brings great convenience to people's daily lives but also makes great contributions to ensuring the security of users' personal information and privacy data.

[0003] In order to provide users with a powerful artificial intelligence model, usually, it is necessary to first select a model with better performance from many models that have undergone the pre-training stage, and then fine-tune the selected model through a preset labeled data set, and finally obtain an artificial intelligence model for performing business.

[0004] In the process of selecting a model above, it is necessary to evaluate each model that has undergone the pre-training stage to determine the performance of each model. However, the current evaluation method cannot accurately evaluate the performance of the model in the pre-training stage. There may be a situation where the evaluation score is high but the actual performance of the model is poor, or the evaluation score is low but the actual performance of the model is high. This also makes it impossible to ensure that a powerful artificial intelligence model is provided to users subsequently, thus bringing inconvenience to users' business execution. Summary of the Invention

[0005] In view of this, one or more embodiments of this specification provide the following technical solutions:

[0006] According to a first aspect of one or more embodiments of this specification, a model evaluation method is proposed. The method is applied to evaluate at least one to-be-tested model that has only undergone pre-training, and includes:

[0007] Obtain a test request for the to-be-tested model;

[0008] Determine a test task to be performed on the to-be-tested model according to the test request;

[0009] If it is determined that the test task is an open-ended task, obtain a specified question text from a preset sample data set, and obtain a question-and-answer example template for the specified question text. The open-ended task is used to enable the to-be-tested model to output an answer based on a question text that does not contain answer options. The question-and-answer example template contains an example question text, an example analysis text, and an example answer text. The example analysis text is used to represent the reasoning process based on which the example question is answered;

[0010] Combine the Q&A example template and the specified question text to obtain a target test text;

[0011] Input the target test text into the model to be tested, obtain the output result of the model to be tested, and evaluate the model to be tested according to the output result.

[0012] Optionally, combining the Q&A example template and the specified question text to obtain a target test text specifically includes:

[0013] Combine the Q&A example template and the specified question text to obtain an initial test text;

[0014] Add an identification field to the initial test text to obtain the target test text, where a question identification field is added to the text head of the specified question text and the example question text, and the question identification field is used to identify that the text after the question identification field is the question text, and

[0015] Add an analysis identification field to the text head of the example analysis text, and the analysis identification field is used to identify that the text after the analysis identification field is the analysis text, and

[0016] Add an answer identification field to the text head of the example answer text, and the answer identification field is used to identify that the text after the answer identification field is the answer text.

[0017] Optionally, the target test text contains multiple specified question texts;

[0018] Inputting the target test text into the model to be tested to obtain the output result of the model to be tested specifically includes:

[0019] Input the target test text into the model to be tested, so that the model to be tested generates an answer text corresponding to each specified question text, and outputs the answer text corresponding to the specified question text after recognizing the question identification field included in the next specified question text.

[0020] Optionally, when evaluating multiple models to be tested, evaluating the models to be tested according to the output result specifically includes:

[0021] For each model to be tested, determine the evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested, where the number of training samples used by different models to be tested in the pre-training stage is different;

[0022] Determine the stability evaluation results corresponding to the to-be-tested models according to the evaluation scores corresponding to each to-be-tested model and the number of training samples used by each to-be-tested model in the pre-training stage. The stability evaluation results are used to indicate whether the magnitude of the number of training samples used by each to-be-tested model in the pre-training stage is positively correlated with the magnitude of the evaluation score corresponding to each to-be-tested model.

[0023] Optionally, determining the stability evaluation results corresponding to the to-be-tested models according to the evaluation scores corresponding to each to-be-tested model and the number of training samples used by each to-be-tested model in the pre-training stage specifically includes:

[0024] For any two to-be-tested models, if the magnitude relationship of the evaluation scores corresponding to the two to-be-tested models is the same as the magnitude relationship of the number of training samples used by the two to-be-tested models in the pre-training stage, determine that the two to-be-tested models are a consistent pair; otherwise, determine that the two to-be-tested models are an inconsistent pair.

[0025] Determine the stability evaluation results corresponding to the to-be-tested models according to the number of determined consistent pairs.

[0026] Optionally, when evaluating multiple to-be-tested models, evaluating the to-be-tested models according to the output results, specifically including:

[0027] For each to-be-tested model, determine the evaluation score corresponding to the to-be-tested model according to the deviation between the output result of the to-be-tested model and the labeled result corresponding to the to-be-tested model, and determine the evaluation score of the fine-tuning model according to the deviation between the output result of the fine-tuning model corresponding to the to-be-tested model and the labeled result corresponding to the fine-tuning model. Among them, the number of training samples used by different to-be-tested models in the pre-training stage is different, and the fine-tuning model corresponding to the to-be-tested model is obtained by fine-tuning on the basis of the to-be-tested model through a preset labeled data set.

[0028] Determine the consistency evaluation results for each to-be-tested model and the fine-tuning model corresponding to each to-be-tested model according to the evaluation score corresponding to each to-be-tested model and the evaluation score of the fine-tuning model corresponding to each to-be-tested model. The consistency evaluation results are used to indicate whether the magnitude of the evaluation score of the to-be-tested model is positively correlated with the magnitude of the evaluation score of the fine-tuning model corresponding to the to-be-tested model.

[0029] Optionally, determining the consistency evaluation results for each to-be-tested model and the fine-tuning model corresponding to each to-be-tested model according to the evaluation score corresponding to each to-be-tested model and the evaluation score of the fine-tuning model corresponding to each to-be-tested model specifically includes:

[0030] For any two models to be tested, if the size relationship of the evaluation scores corresponding to the two models to be tested is the same as the size relationship of the evaluation scores of the fine-tuned models corresponding to the two models to be tested, then determine that the two models to be tested are a consistent pair; otherwise, determine that the two models to be tested are an inconsistent pair.

[0031] According to the number of determined consistent pairs, determine the consistency evaluation results for each model to be tested and the fine-tuned models corresponding to each model to be tested.

[0032] According to the second aspect of one or more embodiments of this specification, a model evaluation method is proposed. The method is used to evaluate at least one model to be tested that has only undergone pre-training, and includes:

[0033] Obtain a test request for the model to be tested;

[0034] According to the test request, determine the test task to be performed on the model to be tested;

[0035] If it is determined that the test task is a multiple-choice task, obtain the original test text from a preset sample dataset. The original test text contains the specified question text that the model to be tested needs to answer, and each answer option corresponding to the specified question text. The multiple-choice task is used to enable the model to be tested to select the answer to the question from the preset answer options;

[0036] For each answer option, rewrite the specified question text through the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option as the rewritten text corresponding to the answer option;

[0037] Use the rewritten texts corresponding to each answer option as the target test text and input it into the model to be tested to obtain the output result of the model to be tested, and evaluate the model to be tested according to the output result.

[0038] Optionally, rewriting the specified question text through the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option specifically includes:

[0039] According to the semantics of the specified question text, determine the filling position of the answer text corresponding to the answer option after rewriting the specified question text into a fill-in-the-blank text;

[0040] Fill the answer text corresponding to the answer option into the filling position to obtain the rewritten text corresponding to the answer option.

[0041] Optionally, when evaluating multiple models to be tested, evaluating the model to be tested according to the output result specifically includes:

[0042] For each model to be tested, according to the deviation between the output result of the model to be tested and the corresponding labeled result of the model to be tested, determine the evaluation score corresponding to the model to be tested, where the number of training samples used by different models to be tested in the pre-training stage is different;

[0043] According to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, determine the stability evaluation result corresponding to each model to be tested, and the stability evaluation result is used to indicate whether the size of the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the size of the evaluation score corresponding to each model to be tested.

[0044] Optionally, according to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, determine the stability evaluation result corresponding to each model to be tested, specifically including:

[0045] For any two models to be tested, if the size relationship of the evaluation scores corresponding to the two models to be tested is the same as the size relationship of the number of training samples used by the two models to be tested in the pre-training stage, determine that the two models to be tested are a consistent pair, otherwise, determine that the two models to be tested are an inconsistent pair;

[0046] According to the number of determined consistent pairs, determine the stability evaluation result corresponding to each model to be tested.

[0047] Optionally, when evaluating multiple models to be tested, according to the output result, evaluate the model to be tested, specifically including:

[0048] For each model to be tested, according to the deviation between the output result of the model to be tested and the corresponding labeled result of the model to be tested, determine the evaluation score corresponding to the model to be tested, and according to the deviation between the output result of the fine-tuning model corresponding to the model to be tested and the corresponding labeled result of the fine-tuning model, determine the evaluation score of the fine-tuning model, where the number of training samples used by different models to be tested in the pre-training stage is different, and the fine-tuning model corresponding to the model to be tested is obtained by fine-tuning on the basis of the model to be tested through a preset labeled data set;

[0049] According to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuning model corresponding to each model to be tested, determine the consistency evaluation result for each model to be tested and the fine-tuning model corresponding to each model to be tested, and the consistency evaluation result is used to indicate whether the size of the evaluation score of the model to be tested is positively correlated with the size of the evaluation score of the fine-tuning model corresponding to the model to be tested.

[0050] Optionally, based on the evaluation scores corresponding to each model to be tested and the evaluation scores of the fine-tuned models corresponding to each model to be tested, determine the consistency evaluation results for each model to be tested and the fine-tuned models corresponding to each model to be tested, specifically including:

[0051] For any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the evaluation scores of the fine-tuned models corresponding to the two models to be tested, then determine that the two models to be tested are a consistent pair; otherwise, determine that the two models to be tested are an inconsistent pair;

[0052] Based on the number of determined consistent pairs, determine the consistency evaluation results for each model to be tested and the fine-tuned models corresponding to each model to be tested.

[0053] According to the third aspect of one or more embodiments of this specification, a model evaluation device is proposed. The device is used to evaluate at least one model to be tested that has only undergone pre-training, and includes:

[0054] A first acquisition module, configured to acquire a test request for the model to be tested;

[0055] A task determination module, configured to determine a test task to be executed for the model to be tested according to the test request;

[0056] A second determination module, configured to, if it is determined that the test task is an open-ended task, acquire a specified question text from a preset sample dataset, and acquire a question-and-answer example template for the specified question text. The open-ended task is used to enable the model to be tested to output an answer based on a question text that does not include answer options. The question-and-answer example template includes an example question text, an example analysis text, and an example answer text. The example analysis text is used to represent the reasoning process based on which the example question is answered;

[0057] A combination module, configured to combine the question-and-answer example template and the specified question text to obtain a target test text;

[0058] An evaluation module, configured to input the target test text into the model to be tested to obtain an output result of the model to be tested, and evaluate the model to be tested according to the output result.

[0059] According to the fourth aspect of one or more embodiments of this specification, a model evaluation device is proposed. The device is used to evaluate at least one model to be tested that has only undergone pre-training, and includes:

[0060] A first acquisition module, configured to acquire a test request for the model to be tested;

[0061] A task determination module, configured to determine a test task to be performed on the model under test according to the test request;

[0062] A second acquisition module, configured to, if it is determined that the test task is a multiple-choice task, acquire an original test text from a preset sample data set, where the original test text includes a specified question text that the model under test needs to answer, and each answer option corresponding to the specified question text, and the multiple-choice task is used to enable the model under test to select an answer to the question to be answered from the preset answer options;

[0063] A rewriting module, configured to, for each answer option, rewrite the specified question text through the answer text corresponding to the answer option to obtain a rewritten text including the answer text corresponding to the answer option as the rewritten text corresponding to the answer option;

[0064] An evaluation module, configured to input the rewritten texts corresponding to the answer options as target test texts into the model under test to obtain an output result of the model under test, and evaluate the model under test according to the output result.

[0065] According to a fifth aspect of one or more embodiments of the present specification, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor runs the executable instructions to implement the steps of the above model evaluation method.

[0066] According to a sixth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the above model evaluation method are implemented.

[0067] According to a seventh aspect of one or more embodiments of the present specification, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above model evaluation method are implemented.

[0068] As can be seen from the above embodiments, since the text structure of the text obtained from the sample data set can be adjusted according to the specific test task to obtain a target test text that is more convenient for the model under test in the pre-training stage to understand, when evaluating the model under test through the target test text, the true performance of the model under test can be accurately evaluated, thereby providing a clear basis for subsequent model selection, and ensuring that a powerful artificial intelligence model is provided to users in the subsequent process, and guaranteeing the business execution efficiency of users based on the artificial intelligence model. Description of the Drawings

[0069] Figure 1Schematic diagram of the relationship between the evaluation scores obtained by the current model evaluation method provided in this specification and the number of training samples used by the model in the pre-training stage;

[0070] Figure 2 Schematic diagram of the relationship between the evaluation scores obtained by the model in the pre-training stage using the current model evaluation method and the evaluation scores obtained by the fine-tuned model using the current model evaluation method provided in this specification;

[0071] Figure 3 Schematic flow diagram of a model evaluation method based on open-ended tasks provided in this specification;

[0072] Figure 4 Schematic diagram of an adjustment for the text corresponding to open-ended tasks provided in this specification;

[0073] Figure 5 Schematic flow diagram of a model evaluation method based on multiple-choice tasks provided in this specification;

[0074] Figure 6 Schematic diagram of an adjustment for the original test text corresponding to multiple-choice tasks provided in this specification;

[0075] Figure 7 Schematic diagram of the structure of a device provided in this specification;

[0076] Figure 8 Block diagram of a model evaluation device provided in this specification;

[0077] Figure 9 Block diagram of a model evaluation device provided in this specification. Detailed implementation manners

[0078] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0079] Currently, in order to ensure that artificial intelligence models can effectively and accurately execute services such as intelligent customer service, autonomous driving, information recommendation, intelligent healthcare, data prediction, etc., in addition to optimizing the artificial intelligence models themselves, it is particularly crucial to select models with better performance from numerous basic models in the pre-training stage to perform subsequent model optimization tasks in order to ultimately obtain artificial intelligence models with better performance.

[0080] Therefore, in practical applications, it is necessary to evaluate each basic model that has gone through the pre-training stage through some model evaluation methods, so as to select a model with better performance. (For the sake of convenience of description, the model that has only gone through the pre-training stage will be referred to as the basic model below).

[0081] However, the current model evaluation methods cannot accurately evaluate the true performance of each basic model, and there are often situations where the evaluation scores do not match the true performance of the basic model. Therefore, the evaluation scores obtained through the current model evaluation methods may be seriously misleading, resulting in the final selection of a basic model with poor performance as the object of subsequent model optimization. Furthermore, the artificial intelligence model obtained after optimization may not be able to give the accurate results required by users during business execution, thus affecting the business execution efficiency and experience of users.

[0082] The poor accuracy of the current model evaluation methods in model evaluation is mainly reflected in two aspects. The first is model stability, and the second is model consistency.

[0083] The first point:

[0084] The so-called model stability means that for a model, if the number of training samples used for the model in the pre-training stage is larger, then the understanding ability of the model should be stronger, and the performance of the model should also be higher. However, due to the fact that the current model evaluation methods cannot accurately give the evaluation scores of the models, the evaluation scores obtained after evaluation of the models with a higher number of training samples used in the pre-training stage may also be lower, as Figure 1 shown.

[0085] Figure 1 is a schematic diagram of the relationship between the evaluation scores obtained by using the current model evaluation methods provided in this specification and the number of training samples used by the model in the pre-training stage.

[0086] Figure 1 mainly uses three current conventional model evaluation methods, and the training steps mentioned in Figure 1 usually refer to the process in which the model uses a batch of training samples for one forward propagation and one backward propagation. In this process, the model calculates the loss value based on the current batch of data and updates the model parameters through the backpropagation algorithm. Therefore, each time such an operation is completed, it is counted as one training step.

[0087] Therefore, the more training steps the model executes, it means that the more batches of training samples are used in the pre-training stage, and the more training samples are used.

[0088] Furthermore, fromFigure 1 It can be clearly seen that regardless of the model evaluation method, the evaluation scores obtained do not show a positive correlation with the number of training steps performed by the model during the pre-training stage. For models that perform a higher number of training steps during the pre-training stage, the final evaluation scores obtained are not necessarily higher. On the contrary, for models that perform a lower number of training steps during the pre-training stage, the final evaluation scores are not necessarily lower. The reason for this situation is that the current model evaluation methods cannot accurately evaluate the true performance of the model.

[0089] Second point:

[0090] The so-called model consistency means that: if the performance of the base model after the pre-training stage is stronger, then the model obtained after fine-tuning it with the preset labeled dataset should also be stronger in performance. Put another way, assuming there are two base models after the pre-training stage, namely Model A and Model B, if the performance of Model A is higher than that of Model B, after fine-tuning these two models with the same labeled dataset, the performance of the model obtained after fine-tuning Model A should be higher than the performance of the model obtained after fine-tuning Model B.

[0091] However, since the current model evaluation methods cannot accurately give the evaluation scores of the models, it is possible that even if the base model after the pre-training stage has a relatively high corresponding evaluation score, the evaluation score of the model obtained after its fine-tuning may be relatively low, as Figure 2 shown.

[0092] Figure 2 This is a schematic diagram showing the relationship between the evaluation scores obtained for the base model using the current model evaluation method in this specification and the evaluation scores obtained for the fine-tuned model using the current model evaluation method.

[0093] Figure 2 A total of 6 models are involved. These 6 models can be understood as models with the same model architecture but different numbers of parameters. From Figure 2 it can be clearly seen that since the evaluation scores obtained by the existing model evaluation methods are not accurate, they cannot reflect the relationship between the base model and its corresponding fine-tuned model either. This is mainly because the existing model evaluation methods do not accurately evaluate the base model.

[0094] Take Figure 2Taking Model 4 and Model 5 as examples, both Model 4 and Model 5 are basic models. The evaluation score obtained after fine-tuning Model 4 is the highest. However, compared with other basic models that have only gone through the pre-training stage, the evaluation score of Model 4 is not the highest, but rather relatively low. The evaluation score obtained after fine-tuning Model 5 is not high compared with other fine-tuned models, but the evaluation score of Model 5 is the highest compared with other basic models that have only gone through the pre-training stage. The reason for this situation also lies in that the current model evaluation method cannot accurately evaluate the true performance of the model.

[0095] It can be seen from the above content that whether it is the model stability in the first point or the model consistency in the second point, the current model evaluation method cannot effectively evaluate them. The main reason is that: basic models that have only gone through the pre-training stage often use a large number of unlabeled training samples for training during the pre-training stage, and this process can enable the basic model to have good language expression ability.

[0096] However, since the basic model has not gone through the fine-tuning stage, for some specific forms of instruction texts, because these instruction texts are not in the conventional sentence patterns (the conventional sentence patterns mentioned here refer to sentences with common sentence pattern structures such as declarative sentences, interrogative sentences, imperative sentences, conditional sentences, etc.), when these specific forms of instruction texts are input into the basic model, the basic model cannot well understand their semantics.

[0097] For example, when the instruction text is a multiple-choice text, it contains a question text and multiple answer options. If this multiple-choice text is used as a test text and input into the basic model, since the multiple-choice text has various answer options, the basic model may piece together the question text + option number (such as A, B, C, etc.) + the answer text in the answer options into a text sequence, and calculate the perplexity corresponding to this text sequence to output the final result. And because this kind of text has options and does not belong to the conventional text sentence pattern, the option number may mislead the basic model in semantic understanding, resulting in the basic model not being able to give a result that conforms to its own performance.

[0098] In this specification, the so-called perplexity is used to measure the prediction ability of a language model for a set of text data. The lower the value, the better the performance of the language model. From another perspective, the lower the perplexity, the more accurate the prediction of the language model for the text. During the calculation of perplexity, perplexity can be regarded as a quantification of the "uncertainty" of the prediction probability distribution of the language model. Therefore, perplexity is calculated based on the prediction probability of each word by the language model.

[0099] For another example, when the instruction text is open-ended text, which includes the text of a specified question that the model needs to answer and the prompt text for prompting the model to answer the specified question, then the model can, based on the prompt of the prompt text, answer the answer corresponding to the specified question and output it.

[0100] However, from the perspective of the base model, different from the multiple-choice text, the open-ended text does not contain answer options, and the model needs to output the answer by itself according to the specified question text in the open-ended text. Therefore, in the absence of any reference, the base model often cannot effectively give the answer to the specified question text, but this does not mean that the language expression ability of the base model is very poor. In addition, after the base model outputs the answer to the specified question text, it may combine the output result with the content in the open-ended text and continue to output the result. Obviously, the base model will not only consume additional resources to give an answer that is not originally needed, but also the result continuously output by the base model may seriously deviate from the original specified question.

[0101] Based on the actual test requirements, the current model evaluation methods often use test texts in a specific form (i.e., the above-mentioned instruction texts) to evaluate the base model. The reason for using this is that most of the current test texts are derived from the sample data used for fine-tuning the base model, so there is no need to construct additional test texts specifically for the base model. However, this obviously also leads to the situation that when using these test texts to evaluate the base model, the base model cannot effectively understand or even misinterprets the text content of the test texts input into it, and thus cannot obtain accurate evaluation results.

[0102] Therefore, this specification provides a model evaluation method. Since this method can adjust the text structure of the text obtained from the sample data set according to the specific test task to obtain a target test text that is more convenient for the model under test that has only gone through the pre-training stage to understand, when the target test text is input into the model under test that has only gone through the pre-training stage, the model under test can more easily understand the semantics therein and can also give an output result that conforms to its own performance. Therefore, through the model evaluation method provided in this specification, the true performance of the model under test can be accurately evaluated, providing a clear basis for subsequent model selection, and ensuring that a powerful artificial intelligence model can be provided to users in the subsequent process, guaranteeing the business execution efficiency of users based on the artificial intelligence model.

[0103] As can be seen from the above, since there are two different forms of instruction texts involved in the model evaluation process, there are also two different test tasks in this specification. One is an open-ended task, mainly used to enable the model under test to output answers based on question texts that do not contain answer options. The other is a multiple-choice task, mainly used to enable the model under test to select the answer to the question to be answered from the preset answer options.

[0104] For different tasks, the adjustment methods for adjusting the text are also different. The adjustment methods for the two different tasks will be introduced separately below.

[0105] The technical solutions provided by the embodiments of this specification will be described in detail below with reference to the accompanying drawings.

[0106] I. Open-ended task:

[0107] Figure 3 It is a schematic flowchart of a model evaluation method based on an open-ended task provided by this specification.

[0108] S300: Obtain a test request for the model under test.

[0109] In this specification, when testing the model under test, the user can initiate a test request. Among them, the model under test mentioned in this specification is the basic model that has only gone through the pre-training stage mentioned above. During the process of performing model evaluation, the user can perform test operations on the terminal device used, and the terminal device generates a corresponding test request, and then executes the test request in the subsequent process.

[0110] The execution subject of the model evaluation method provided by this specification can be the terminal device mentioned above. The terminal device can refer to electronic devices such as desktop computers and laptop computers, or it can be the client installed in the terminal device for performing model evaluation. Of course, the execution subject can also be a server. For the convenience of description, in the following, only the terminal device will be used as the execution subject to describe the specific process of the model evaluation method provided by this specification.

[0111] In addition, the model under test (i.e., the basic model mentioned above) in this specification can be a large language model (Large Language Model, LLM). The model finally obtained after fine-tuning the model under test can provide functions such as intelligent customer service, knowledge answering, and information search for users in actual business applications. Of course, the model under test can also be other artificial intelligence models that can process text, and this specification does not strictly limit the specific form of the model under test.

[0112] S302: Determine the test task to be performed on the model under test according to the test request.

[0113] After obtaining the above test request, the terminal device can parse the test request to determine the test tasks to be performed on the model to be tested subsequently. In this process, the terminal device can determine which one or which models to be tested need to be evaluated based on the parsing result of the test request, and what specific types of test tasks to perform on each model to be tested.

[0114] In this embodiment, it mainly focuses on open-ended tasks. Therefore, subsequently, the terminal device needs to determine the target test text required for performing model evaluation based on this open-ended task.

[0115] In addition, from the above content, it can be seen that the model evaluation method provided in this specification does not necessarily only evaluate one model to be tested, but can also evaluate multiple models to be tested. When evaluating multiple models to be tested, these models to be tested can be basic models trained with different numbers of training samples in the pre-training stage. Then, in the subsequent process, through the evaluation of these models to be tested, it can be determined which one or which of the models to be tested have better performance and can be used in the subsequent fine-tuning process and business deployment.

[0116] S304: If it is determined that the test task is an open-ended task, obtain the specified question text from the preset sample data set, and obtain the Q&A example template for the specified question text.

[0117] After the terminal device determines the test task, it can determine the required target test text according to the task type of the test task. Among them, the text content in the target test text can use the current conventional sample data set, and there is no need to additionally construct a corresponding data set for the model to be tested. Therefore, the terminal device can first determine the corresponding sample data set according to the task type, and then determine the required specified question text from the sample data set. There are various ways to select the specified question text. For example, the specified question text can be randomly selected from the sample data set, or the appropriate specified question text can be selected according to the number of times the specified question text is used. Other methods will not be exemplified here one by one.

[0118] For open-ended tasks, in addition to obtaining the above specified question text, it is also necessary to obtain the answer example template for providing reference for the model to be tested. The Q&A example template includes example question text, example analysis text, and example answer text. The example analysis text is used to represent the reasoning process based on which the example question is answered.

[0119] For example, assume that a Q&A example template contains the following text content:

[0120] Question: John had a wrist-wrestling match with 20 people and won 80% of them. How many people did he lose to?

[0121] Let's think step by step. He defeated 20 × 0.8 = 16 people, so he lost to 20 - 16 = 4 people.

[0122] 4 people.

[0123] In the above example, "Question: John had a wrist wrestling match with 20 people and won 80%. How many people did he lose to?" is the example question text included in the Q&A example template, "Let's think step by step. He defeated 20 × 0.8 = 16 people, so he lost to 20 - 16 = 4 people" is the example analysis text included in the Q&A example template, and "4 people" is the example answer text included in the Q&A example template.

[0124] S306: Combine the Q&A example template and the specified question text to obtain the target test text.

[0125] After determining the above-mentioned specified question text and Q&A example template, the terminal device needs to combine the specified question text with the Q&A example template to obtain the target test text, and then evaluate the model to be tested through the target test text in the subsequent process.

[0126] From the above detailed analysis of the existing problems, it can be seen that the current test text does not provide a reasoning reference for the basic model, making it difficult for the basic model to smoothly infer the required answer. However, this does not mean that the language expression ability of the basic model is poor. Therefore, the currently used test text cannot well reflect the actual performance of the basic model. Therefore, after adding the above Q&A example template to the target test text, the model to be tested can refer to the reasoning process in the Q&A example template to effectively obtain the answer to the specified question text.

[0127] Furthermore, to prevent the situation where the model to be tested combines the output result with the input text and continues to output an answer, in this specification, the terminal device can first combine the Q&A example template and the specified question text to obtain the initial test text, and then add some identification fields to the initial test text to obtain the target test text. Among them, a question identification field can be added to the text head of the specified question text and the text head of the example question text, and this question identification field is used to identify that the text after the question identification field is the question text.

[0128] An analysis identification field can be added to the text head of the example analysis text, and the analysis identification field is used to identify that the text after the analysis identification field is the analysis text corresponding to the example question text.

[0129] An answer identification field is added to the text head of the example answer text, and the answer identification field is used to identify that the text after the answer identification field is the answer text.

[0130] Therefore, these identification fields added to the initial test text can actually be regarded as the division boundaries of different text contents.

[0131] To further illustrate the adjustment method corresponding to the open-ended task, a specific example will be used to explain it below, as Figure 4 shown.

[0132] Figure 4 It is a schematic diagram for adjusting the text corresponding to the open-ended task provided in this specification.

[0133] Figure 4 It shows the specific forms of the original test text and the target test text. Among them, "Given the following questions, please reason step by step and give the answer at the end" is the prompt text used to prompt the model under test to answer the specified questions when introducing the above instruction text. It can be seen from the original test text that it does not contain the Q&A example template for providing reasoning references to the model under test. Therefore, it is difficult for the model under test to effectively output the answer to the specified question text without being fine-tuned.

[0134] And in Figure 4 the shown target test text, it not only contains the Q&A example template, but also adds identification fields at the text headers of different text contents. Among them, "Problem" is the question identification field in the example question text and the specified question text, "Solution" is the parsing identification field in the example parsing text, and "The finalanswer" is the answer identification field.

[0135] It can be seen that these identification fields clearly divide the text content in the target test text, that is, it indicates which part belongs to the question text, which part belongs to the parsing text, and which part belongs to the answer text. Therefore, the model under test can deepen its understanding of the text content in the target test text by recognizing these identification fields, so as to obtain accurate output results.

[0136] In addition, in Figure 4 the example, in addition to the parsing identification field of the example parsing text being included in the Q&A example template, a parsing identification field can also be set after the specified question text, which can enable the model under test to output the reasoning process for the specified question corresponding to the specified question text after recognizing this parsing identification field. Of course, an answer identification field corresponding to the final output result can also be added to the target test text ( Figure 4 not shown in the figure), which enables the model under test to output the answer to the specified question corresponding to the specified question text after recognizing this answer identification field.

[0137] It should be noted that in the above example, the target test text only contains one Q&A example template and one specified question text. However, in actual applications, the target test text can contain multiple Q&A example templates and multiple specified question texts that need to be answered by the model under test.

[0138] S308: Input the target test text into the model under test to obtain the output result of the model under test, and evaluate the model under test based on the output result.

[0139] After obtaining the above target test text, the terminal device can input it into the model under test, and the model under test will perform semantic understanding on the target test text to obtain the corresponding output result.

[0140] In this process, since the target test text corresponding to the open-ended task not only contains Q&A example templates, but also various identification fields added to the target test text, when the model under test generates the answer text corresponding to each specified question text in the target test text and then recognizes the question identification field included in the next specified text question, it will stop continuing to output the answer to the specified question text and output the already generated answer text.

[0141] It can be seen that each question identification field actually divides the answering scope of each specified question text. Once the model under test recognizes a question identification field, it will not continue to output the answer to the previous specified question in combination with the text content corresponding to the next specified question, thus effectively ensuring the rationality of the output result of the model under test.

[0142] For the process of evaluating the model under test based on the output result of the model under test, whether it is an open-ended task or a multiple-choice task, the general process is basically the same. Therefore, this process will be introduced together later, and the following will focus on introducing the adjustment method for the text corresponding to the multiple-choice task, as Figure 5 shown.

[0143] II. Multiple-choice task:

[0144] Figure 5 is a schematic flowchart of a model evaluation method provided in this specification based on a multiple-choice task.

[0145] S500: Obtain a test request for the model under test.

[0146] S502: Determine the test task to be performed on the model under test according to the test request.

[0147] In this specification, the terminal device can parse the obtained test request to determine the test task to be performed subsequently. This test task determines which model or models to be tested the terminal device needs to perform model evaluation for, and determines which task text needs to be adjusted subsequently. The specific contents of these two steps are basically the same as the first two steps involved in the above-mentioned open task, so they will not be repeated here.

[0148] S504: If it is determined that the test task is a multiple-choice task, an original test text is obtained from a preset sample data set.

[0149] The terminal device can obtain the original test text from the existing sample data set. For a multiple-choice task, the corresponding original test text includes the specified question text that the model to be tested needs to answer and the answer options corresponding to the specified question text.

[0150] For example, suppose an original test text corresponding to a multiple-choice task contains the following text content:

[0151] Please answer which is the highest mountain in the world?

[0152] A. Mount Everest;

[0153] B. K2;

[0154] C. Kangchenjunga.

[0155] “Please answer which is the highest mountain in the world?” is the designated question text contained in the original test text, and “A. Mount Everest”, “B. K2” and “C. Kangchenjunga” are the three answer options contained in the original test text.

[0156] In addition, the answer options contained in the original test text corresponding to the multiple-choice task are not necessarily in the form of the above examples. The options such as "right" and "wrong", "yes" and "no" selected for judgment questions are also answer options.

[0157] S506: For each answer option, rewrite the specified question text using the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option as the rewritten text corresponding to the answer option.

[0158] Furthermore, when the test task is a multiple-choice task, the terminal device can rewrite the specified question text based on each answer option to obtain a rewritten text containing the answer text corresponding to the answer option, wherein each answer option can correspond to a rewritten text.

[0159] This process can actually be understood as rewriting the original question statement into a declarative sentence stating the answer. However, since each answer option contains incorrect answers, the rewritten text actually contains declarative sentences with factual statement errors.

[0160] In the specific execution process, for each answer option, the terminal device can, according to the semantics of the specified question text in the original test text, determine the filling position of the answer text corresponding to this answer option in the specified question text after rewriting the specified question text into a fill-in-the-blank text. Then, by filling the answer text corresponding to this answer option into this filling position, the rewritten text corresponding to this answer option can be obtained.

[0161] Among them, since each answer option corresponds to the same specified question text, the filling positions of these answer options in the specified question text can actually be the same (of course, they can also be different. If the filling positions are different, the word order of the specified question text needs to be adjusted according to the semantics of the specified question text. In the case of different word orders, the filling positions will also be different. However, it is necessary to ensure that the expressed meaning is always the same, that is, no matter how it is adjusted, the consistency of the specified question needs to be ensured). Then, for each answer option, according to the determined filling position, the answer text contained in this answer option can be filled into the specified question text, so as to obtain the rewritten text corresponding to this answer option. Finally, the rewritten text corresponding to each answer option can be used as the target test text.

[0162] For the determination of the filling position, the terminal device needs to, according to the semantics of the specified question text, identify the fields (including punctuation marks) used to replace the answers in the specified question text, and then remove these fields to obtain the above-mentioned filling position.

[0163] For the convenience of understanding, the following will use a specific example to illustrate the adjustment method of the original test text corresponding to the multiple-choice task, as Figure 6 shown.

[0164] Figure 6 FIG. is a schematic diagram for adjusting the original test text corresponding to the multiple-choice task provided in this specification.

[0165] From Figure 6 the example, it can be seen that the specified question text in the original test text is: "What is the capital of Country A?", and it contains two answer options: "A: City X" and "B: City Y". The terminal device can, through semantic recognition of the specified question text, determine that "What?" is the field to be replaced, and then use the position of this field as the filling position where the answer needs to be filled in subsequently.

[0166] The terminal device fills the answer texts of the two answer options into the specified question text according to the determined filling positions respectively, so as to obtain Figure 6 two texts in the target test text in Figure 6 . It can be seen that these two texts are no longer in the form of interrogative sentences, but are formed by splicing "What is the capital of Country A" in the specified question text and the answer text (removing the option number), forming two declarative sentences. Subsequently, the model to be tested can calculate the perplexity of these two texts and select the text that is considered correct (i.e., the text with the smallest perplexity) from these two texts, so as to output the final result.

[0167] S508: Use the rewritten text corresponding to each answer option as the target test text and input it into the model to be tested, obtain the output result of the model to be tested, and evaluate the model to be tested according to the output result.

[0168] The terminal device can input the target test text determined by the above method into the model to be tested, and evaluate the model to be tested based on the obtained output result. Since the process of model evaluation is basically the same for both open-ended tasks and multiple-choice tasks, it will be introduced together below.

[0169] As mentioned above, the terminal device can actually evaluate multiple models to be tested. Then, for each model to be tested, the terminal device can determine the evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested. Since the above-mentioned specified question text is determined from the existing sample dataset, there is a corresponding labeled result for the specified question text, and this labeled result can be understood as the correct answer to the specified question corresponding to the specified question text. Therefore, if the deviation between the output result of the model to be tested and the corresponding labeled result is larger, it means that the result output by the model to be tested is farther from the correct answer, and the evaluation score is lower; if the deviation between the output result of the model to be tested and the corresponding labeled result is smaller, it means that the result output by the model to be tested is closer to the correct answer, and the evaluation score is higher.

[0170] Then, the terminal device can determine the evaluation results for evaluating each model to be tested according to the evaluation scores corresponding to each model to be tested. Among them, there are two types of evaluation results. One is used to reflect the model stability, and its result can be called the stability evaluation result, which is used to indicate whether the size of the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the size of the evaluation score corresponding to each model to be tested. The other is used to reflect the model consistency, and its result can be called the consistency evaluation result, which is used to indicate whether the size of the evaluation score of the model to be tested is positively correlated with the size of the evaluation score of the fine-tuning model corresponding to the model to be tested. The model stability and the model consistency have been clearly explained in the above content and will not be elaborated here.

[0171] For model stability, the terminal device can determine the stability evaluation results corresponding to each model to be tested according to the evaluation scores corresponding to each model to be tested and the number of training samples used by each model to be tested during the pre-training phase. Among them, for any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the number of training samples used by the two models to be tested during the pre-training phase, then the two models to be tested can be determined to be a consistent pair; otherwise, the two models to be tested can be determined to be an inconsistent pair. The terminal device can determine the stability evaluation results corresponding to each model to be tested according to the number of determined consistent pairs.

[0172] For example, there are the following 5 models to be tested, and their corresponding evaluation scores and the number of training samples used during the pre-training phase are shown in the following table.

[0173] Table 1

[0174] Model to be tested Evaluation score Number of training samples used Model A 50 20000 Model B 60 30000 Model C 70 25000 Model D 80 40000 Model E 90 35000

[0175] In Table 1, the evaluation scores of models A to E increase in sequence. However, obviously, the number of training samples used during the pre-training phase does not increase in sequence. Therefore, for the above 5 models combined in pairs, there are a total of 10 combinations. For the two models to be tested, model A and model B, the evaluation score of model A (50 points) is less than the evaluation score of model B (60 points), and the number of training samples used by model A during the pre-training phase (20,000) is less than the number of training samples used by model B during the pre-training phase (30,000). Therefore, whether from the perspective of the evaluation score or the number of training samples used during the pre-training phase, the relationship between the two is the same, that is, model A < model B. Therefore, model A and model B are a consistent pair.

[0176] For models B and C, the evaluation score of model B (60 points) is less than the evaluation score of model C (70 points). However, the number of training samples used by model B during the pre-training phase (30,000) is greater than the number of training samples used by model C during the pre-training phase (25,000). Therefore, from the perspective of the evaluation score, model B < model C, and from the perspective of the number of training samples used during the pre-training phase, model B > model C. The relationship between the two is not consistent. Therefore, model B and model C are an inconsistent pair.

[0177] According to the above rules, the terminal device can determine the number of consistent pairs, and then can calculate the Kendall coefficient through the following formula to determine the stability evaluation results of each model to be tested.

[0178]

[0179] In the above formula, P is used to represent the number of consistent pairs, n is the total number of models to be tested, and n(n - 1) / 2 is used to represent the number of pairs of models to be compared (such as the 10 combinations in Table 1 example above). τ is the calculated stability evaluation result. Among them, the value range of the stability evaluation result is between -1 and 1. If τ = -1, it means that the relationship between the evaluation scores of the models to be tested is completely inconsistent with the relationship between the number of training samples used by the models to be tested in the pre-training stage. If τ = 1, it means that the relationship between the evaluation scores of the models to be tested is completely consistent with the relationship between the number of training samples used by the models to be tested in the pre-training stage. And if τ = 0, it means that there is no correlation between the two relationships of the evaluation scores of the models to be tested and the number of training samples used in the pre-training stage.

[0180] It should be noted that in the above example, the terminal device calculates the stability evaluation result by determining the number of consistent pairs. In actual applications, the terminal device can also calculate the stability evaluation result by determining the number of inconsistent pairs, or calculate the stability evaluation result by the number of consistent pairs and the number of inconsistent pairs. Therefore, the above formula is not unique and can be determined according to actual needs. However, it needs to satisfy the relationship that the higher the number of consistent pairs, the larger the value of τ, and the higher the number of inconsistent pairs, the smaller the value of τ.

[0181] For model consistency, the terminal device can specifically determine the evaluation score corresponding to each model to be tested according to the deviation between the output result of the model to be tested and the corresponding labeled result of the model to be tested, and, determine the evaluation score of the corresponding fine-tuned model according to the deviation between the output result of the fine-tuned model corresponding to the model to be tested and the corresponding labeled result of the fine-tuned model.

[0182] Then, the terminal device can determine the consistency evaluation results for each model to be tested and the corresponding fine-tuned models of each model to be tested according to the evaluation scores corresponding to each model to be tested and the evaluation scores of the corresponding fine-tuned models of each model to be tested. In this process, for any two models to be tested, if the size relationship of the evaluation scores corresponding to the two models to be tested is the same as the size relationship of the evaluation scores of the two corresponding fine-tuned models of the two models to be tested, then it can be determined that the two models to be tested are consistent pairs. Otherwise, it is determined that the two models to be tested are inconsistent pairs. Subsequently, the terminal device can determine the consistency evaluation results for each model to be tested and the corresponding fine-tuned models of each model to be tested according to the number of consistent pairs determined.

[0183] For example, there are 5 models to be tested below, and the corresponding evaluation scores and the evaluation scores of the corresponding fine-tuned models of these 5 models to be tested are shown in the following table.

[0184] Table 2

[0185] Model to be tested Evaluation score Fine-tuned model Evaluation score Model a 50 Model a' 70 Model b 60 Model b' 66 Model c 70 Model c' 84 Model d 80 Model d' 90 Model e 90 Model e' 99

[0186] In the above Table 2, the evaluation scores of the to-be-tested models a - e increase successively. However, the evaluation scores of the fine-tuned models corresponding to these 5 to-be-tested models do not increase successively. Among them, the evaluation score of model c (70 points) is less than the evaluation score of model d (80 points), and the evaluation score of the fine-tuned model corresponding to model c: model c' (84 points) is also less than the evaluation score of the fine-tuned model corresponding to model d: model d' (90 points). Therefore, the magnitude relationship between the evaluation scores of model c and model d is consistent with the magnitude relationship between the evaluation scores of the fine-tuned models corresponding to these two models. Thus, model c and model d belong to a consistent pair.

[0187] For models a and b, the evaluation score of model a (50 points) is less than the evaluation score of model b (60 points), while the evaluation score of the fine-tuned model corresponding to model a: model a' (70 points) is greater than the evaluation score of the fine-tuned model corresponding to model b: model b' (66 points). Therefore, the magnitude relationship between the evaluation scores of model a and model b is inconsistent with the magnitude relationship between the evaluation scores of the fine-tuned models corresponding to these two models. Thus, model a and model b belong to an inconsistent pair.

[0188] According to the above rules, the terminal device can determine the number of consistent pairs, and then determine the consistency evaluation result of each to-be-tested model by calculating the Kendall coefficient. Among them, the terminal device can also use the above formula to determine the consistency evaluation result.

[0189] Moreover, in addition to calculating the consistency evaluation result through the determined number of consistent pairs, the terminal device can also calculate the consistency evaluation result by determining the number of inconsistent pairs, or can calculate the consistency evaluation result through the number of consistent pairs and the number of inconsistent pairs.

[0190] As can be seen from the above content, the model evaluation method provided in this specification does not aim to establish a proprietary test data set for the to-be-tested models. Instead, on the basis of the original sample data set, the text determined in the sample data set is adjusted in text structure to obtain the target test text. Therefore, the model evaluation method provided in this specification actually still uses the currently used test text, but it is necessary to adjust the text structure according to the specific task, which not only reduces the data maintenance cost but also ensures that the actual performance of the to-be-tested models can be accurately tested.

[0191] Moreover, since the model evaluation method adopted in this specification can effectively determine the actual performance of the model to be tested, the obtained evaluation scores are more accurate. Based on the more accurate evaluation scores, the terminal device can accurately determine the above-mentioned stability evaluation result and consistency evaluation result. In this way, based on the two determined evaluation results, a basic model with better performance can be determined from many models to be tested, and it can be fine-tuned in the subsequent process, so as to finally obtain an artificial intelligence model for executing actual services and deploy it. The artificial intelligence model finally obtained based on the model evaluation method provided in this specification can output accurate service results for users, thereby ensuring the user's service execution efficiency while significantly improving the user experience of using the artificial intelligence model to execute services, and providing more service conveniences for users.

[0192] Figure 7 is a schematic structural diagram of a device provided in this specification. Please refer to Figure 7 , at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710. Of course, there may also be other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0193] Please refer to Figure 8 , a model evaluation device provided in this specification can be applied to a device as shown in Figure 7 to implement the technical solution of this specification. Among them, the model evaluation device may include:

[0194] A first acquisition module 800, configured to acquire a test request for the model to be tested;

[0195] A task determination module 802, configured to determine a test task to be executed for the model to be tested according to the test request;

[0196] A second acquisition module 804, configured to, if it is determined that the test task is an open-ended task, acquire a specified question text from a preset sample dataset, and acquire a Q&A example template for the specified question text, where the open-ended task is used to enable the model under test to output an answer based on a question text that does not include answer options, and the Q&A example template includes an example question text, an example analysis text, and an example answer text, and the example analysis text is used to represent the reasoning process for answering the example question;

[0197] A combination module 806, configured to combine the Q&A example template and the specified question text to obtain a target test text;

[0198] An evaluation module 808, configured to input the target test text into the model under test to obtain an output result of the model under test, and evaluate the model under test according to the output result.

[0199] Optionally, the combination module 806 is specifically configured to combine the Q&A example template and the specified question text to obtain an initial test text;

[0200] Add an identification field to the initial test text to obtain a target test text, where a question identification field is added to the text head of the specified question text and the text head of the example question text, and the question identification field is used to identify that the text after the question identification field is the question text, and

[0201] Add an analysis identification field to the text head of the example analysis text, and the analysis identification field is used to identify that the text after the analysis identification field is the analysis text, and

[0202] Add an answer identification field to the text head of the example answer text, and the answer identification field is used to identify that the text after the answer identification field is the answer text.

[0203] Optionally, the target test text includes multiple specified question texts;

[0204] The evaluation module 808 is specifically configured to input the target test text into the model under test, so that the model under test generates an answer text corresponding to each specified question text for each specified question text, and outputs the answer text corresponding to the specified question text after recognizing the question identification field included in the next specified question text.

[0205] Optionally, when evaluating multiple models to be tested, the evaluation module 808 is specifically configured to, for each model to be tested, determine the evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested, where the number of training samples used by different models to be tested in the pre-training stage is different; determine the stability evaluation result corresponding to each model to be tested according to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, and the stability evaluation result is used to indicate whether the magnitude of the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the magnitude of the evaluation score corresponding to each model to be tested.

[0206] Optionally, the evaluation module 808 is specifically configured to, for any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the number of training samples used by the two models to be tested in the pre-training stage, determine that the two models to be tested are a consistent pair, otherwise, determine that the two models to be tested are an inconsistent pair; determine the stability evaluation result corresponding to each model to be tested according to the number of determined consistent pairs.

[0207] Optionally, when evaluating multiple models to be tested, the evaluation module 808 is specifically configured to, for each model to be tested, determine the evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested, and determine the evaluation score of the fine-tuning model according to the deviation between the output result of the fine-tuning model corresponding to the model to be tested and the labeled result corresponding to the fine-tuning model, where the number of training samples used by different models to be tested in the pre-training stage is different, and the fine-tuning model corresponding to the model to be tested is obtained by fine-tuning on the basis of the model to be tested through a preset labeled data set; determine the consistency evaluation result for each model to be tested and the fine-tuning model corresponding to each model to be tested according to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuning model corresponding to each model to be tested, and the consistency evaluation result is used to indicate whether the magnitude of the evaluation score of the model to be tested is positively correlated with the magnitude of the evaluation score of the fine-tuning model corresponding to the model to be tested.

[0208] Optionally, the evaluation module 808 is specifically configured to, for any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the evaluation scores of the fine-tuning models corresponding to the two models to be tested, determine that the two models to be tested are a consistent pair, otherwise, determine that the two models to be tested are an inconsistent pair; determine the consistency evaluation result for each model to be tested and the fine-tuning model corresponding to each model to be tested according to the number of determined consistent pairs.

[0209] Please refer to Figure 9, a model evaluation device provided in this specification can be applied to devices such as Figure 7 shown to implement the technical solutions of this specification. Among them, the model evaluation device may include:

[0210] A first acquisition module 900, configured to acquire a test request for a model to be tested;

[0211] A task determination module 902, configured to determine a test task to be executed for the model to be tested according to the test request;

[0212] A second acquisition module 904, configured to, if it is determined that the test task is a multiple-choice task, acquire original test texts from a preset sample dataset, where the original test texts include designated question texts that the model to be tested needs to answer, and each answer option corresponding to the designated question text, and the multiple-choice task is used to enable the model to be tested to select an answer to the question to be answered from the preset answer options;

[0213] A rewriting module 906, configured to, for each answer option, rewrite the designated question text through the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option as the rewritten text corresponding to the answer option;

[0214] An evaluation module 908, configured to input the rewritten texts corresponding to each answer option as target test texts into the model to be tested to obtain an output result of the model to be tested, and evaluate the model to be tested according to the output result.

[0215] Optionally, the rewriting module 906 is specifically configured to determine the filling position of the answer text corresponding to the answer option after rewriting the designated question text into a fill-in-the-blank text according to the semantics of the designated question text; fill the answer text corresponding to the answer option into the filling position to obtain the rewritten text corresponding to the answer option.

[0216] Optionally, when evaluating multiple models to be tested, the evaluation module 908 is specifically configured to, for each model to be tested, determine an evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested, where the number of training samples used by different models to be tested in the pre-training stage is different; determine the stability evaluation results corresponding to the models to be tested according to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, and the stability evaluation results are used to indicate whether the magnitude of the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the magnitude of the evaluation score corresponding to each model to be tested.

[0217] Optionally, the evaluation module 908 is specifically configured to, for any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the number of training samples used by the two models to be tested in the pre-training stage, determine that the two models to be tested are a consistent pair; otherwise, determine that the two models to be tested are an inconsistent pair; and determine the stability evaluation result corresponding to each model to be tested according to the number of determined consistent pairs.

[0218] Optionally, when evaluating multiple models to be tested, the evaluation module 908 is specifically configured to, for each model to be tested, determine the evaluation score corresponding to the model to be tested according to the deviation between the output result of the model to be tested and the labeled result corresponding to the model to be tested, and determine the evaluation score of the fine-tuned model according to the deviation between the output result of the fine-tuned model corresponding to the model to be tested and the labeled result corresponding to the fine-tuned model, where the number of training samples used by different models to be tested in the pre-training stage is different, and the fine-tuned model corresponding to the model to be tested is obtained by fine-tuning on the basis of the model to be tested through a preset labeled data set; determine the consistency evaluation result for each model to be tested and the fine-tuned model corresponding to each model to be tested, and the consistency evaluation result is used to indicate whether the magnitude of the evaluation score of the model to be tested is positively correlated with the magnitude of the evaluation score of the fine-tuned model corresponding to the model to be tested.

[0219] Optionally, the evaluation module 908 is specifically configured to, for any two models to be tested, if the magnitude relationship of the evaluation scores corresponding to the two models to be tested is the same as the magnitude relationship of the evaluation scores of the fine-tuned models corresponding to the two models to be tested, determine that the two models to be tested are a consistent pair; otherwise, determine that the two models to be tested are an inconsistent pair; and determine the consistency evaluation result for each model to be tested and the fine-tuned model corresponding to each model to be tested according to the number of determined consistent pairs.

[0220] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor runs the executable instructions to implement the steps of the method described in any of the above embodiments.

[0221] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by the processor, the steps of the method described in any of the above embodiments are implemented.

[0222] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, which when executed by a processor, implement the steps of the method described in any of the above embodiments.

[0223] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0224] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for related content.

[0225] The above are only embodiments of this specification and are not used to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A model evaluation method, the method being used to evaluate at least one pre-trained model to be tested, comprising: Get the test request for the model to be tested; Determining, according to the test request, a test task to be performed on the model to be tested; If it is determined that the test task is an open task, a specified question text is obtained from a preset sample data set, and a question-answering example template for the specified question text is obtained, wherein the open task is used to enable the model to be tested to output an answer based on the question text that does not contain answer options, and the question-answering example template contains an example question text, an example parsing text, and an example answer text, and the example parsing text is used to represent the reasoning process based on which the answer to the example question is based; Combine the question and answer example template and the specified question text to obtain a target test text; The target test text is input into the model to be tested, an output result of the model to be tested is obtained, and the model to be tested is evaluated according to the output result.

2. The method according to claim 1, combining the question and answer example template and the specified question text to obtain a target test text, specifically comprising: Combine the question and answer example template and the specified question text to obtain an initial test text; adding an identification field to the initial test text to obtain a target test text, wherein a question identification field is added to the text header of the specified question text and the text header of the sample question text, and the question identification field is used to identify that the question text follows the question identification field, and Adding a parsing identification field to the text header of the example parsed text, wherein the parsing identification field is used to identify that the parsing text follows the parsing identification field, and An answer identification field is added to the text header of the example answer text, and the answer identification field is used to identify that the answer text follows the answer identification field.

3. The method according to claim 2, wherein the target test text includes a plurality of specified question texts; Inputting the target test text into the model to be tested to obtain the output result of the model to be tested specifically includes: The target test text is input into the model to be tested so that the model to be tested generates an answer text corresponding to each specified question text, and after identifying the question identification field contained in the next specified question text, outputs the answer text corresponding to the specified question text.

4. The method according to claim 1, when evaluating multiple models to be tested, evaluating the models to be tested according to the output results, specifically comprising: For each model to be tested, the evaluation score corresponding to the model to be tested is determined according to the deviation between the output result of the model to be tested and the annotation result corresponding to the model to be tested, wherein different models to be tested use different numbers of training samples in the pre-training stage; According to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, the stability evaluation result corresponding to each model to be tested is determined. The stability evaluation result is used to indicate whether the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the evaluation score corresponding to each model to be tested.

5. The method according to claim 4, determining the stability evaluation results corresponding to each model to be tested according to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, specifically comprising: For any two models to be tested, if the size relationship between the evaluation scores corresponding to the two models to be tested is the same as the size relationship between the numbers of training samples used by the two models to be tested in the pre-training stage, then the two models to be tested are determined to be consistently paired; otherwise, the two models to be tested are determined to be inconsistently paired; According to the number of determined consistent pairs, the stability evaluation results corresponding to the models to be tested are determined.

6. The method according to claim 1, when evaluating multiple models to be tested, evaluating the models to be tested according to the output results, specifically comprising: For each model to be tested, the evaluation score corresponding to the model to be tested is determined according to the deviation between the output result of the model to be tested and the annotation result corresponding to the model to be tested, and the evaluation score of the fine-tuned model is determined according to the deviation between the output result of the fine-tuned model corresponding to the model to be tested and the annotation result corresponding to the fine-tuned model, wherein different models to be tested use different numbers of training samples in the pre-training stage, and the fine-tuned model corresponding to the model to be tested is obtained after fine-tuning the model to be tested through a preset annotated data set on the basis of the model to be tested; According to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuned model corresponding to each model to be tested, the consistency evaluation results for each model to be tested and the fine-tuned model corresponding to each model to be tested are determined. The consistency evaluation results are used to indicate whether the evaluation score of the model to be tested is positively correlated with the evaluation score of the fine-tuned model corresponding to the model to be tested.

7. The method according to claim 6, determining the consistency evaluation results for each model to be tested and the fine-tuned model corresponding to each model to be tested according to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuned model corresponding to each model to be tested, specifically comprising: For any two models to be tested, if the size relationship between the evaluation scores corresponding to the two models to be tested is the same as the size relationship between the evaluation scores of the fine-tuning models corresponding to the two models to be tested, then the two models to be tested are determined to be consistently paired; otherwise, the two models to be tested are determined to be inconsistently paired; According to the number of determined consistent pairs, the consistency evaluation results for each model to be tested and the fine-tuning model corresponding to each model to be tested are determined.

8. A model evaluation method, the method being used to evaluate at least one pre-trained model to be tested, comprising: Get the test request for the model to be tested; Determining, according to the test request, a test task to be performed on the model to be tested; If it is determined that the test task is a multiple-choice task, an original test text is obtained from a preset sample data set, wherein the original test text includes a specified question text that the model to be tested needs to answer, and answer options corresponding to the specified question text, and the multiple-choice task is used to enable the model to be tested to select an answer to the question to be answered from the preset answer options; For each answer option, rewrite the specified question text using the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option as the rewritten text corresponding to the answer option; The rewritten text corresponding to each answer option is input into the model to be tested as the target test text to obtain the output result of the model to be tested, and the model to be tested is evaluated based on the output result.

9. The method according to claim 8, rewriting the specified question text by using the answer text corresponding to the answer option to obtain a rewritten text containing the answer text corresponding to the answer option, specifically comprising: According to the semantics of the specified question text, determining the filling position of the answer text corresponding to the answer option after rewriting the specified question text into a fill-in-the-blank text; The answer text corresponding to the answer option is filled into the filling position to obtain the rewritten text corresponding to the answer option.

10. The method according to claim 8, when evaluating multiple models to be tested, evaluating the models to be tested according to the output results, specifically comprising: For each model to be tested, the evaluation score corresponding to the model to be tested is determined according to the deviation between the output result of the model to be tested and the annotation result corresponding to the model to be tested, wherein different models to be tested use different numbers of training samples in the pre-training stage; According to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, the stability evaluation result corresponding to each model to be tested is determined. The stability evaluation result is used to indicate whether the number of training samples used by each model to be tested in the pre-training stage is positively correlated with the evaluation score corresponding to each model to be tested.

11. The method according to claim 10, wherein the stability evaluation result corresponding to each model to be tested is determined according to the evaluation score corresponding to each model to be tested and the number of training samples used by each model to be tested in the pre-training stage, and specifically comprises: For any two models to be tested, if the size relationship between the evaluation scores corresponding to the two models to be tested is the same as the size relationship between the numbers of training samples used by the two models to be tested in the pre-training stage, then the two models to be tested are determined to be consistently paired; otherwise, the two models to be tested are determined to be inconsistently paired; According to the number of determined consistent pairs, the stability evaluation results corresponding to the models to be tested are determined.

12. The method according to claim 1, when evaluating multiple models to be tested, evaluating the models to be tested according to the output results, specifically comprising: For each model to be tested, the evaluation score corresponding to the model to be tested is determined according to the deviation between the output result of the model to be tested and the annotation result corresponding to the model to be tested, and the evaluation score of the fine-tuned model is determined according to the deviation between the output result of the fine-tuned model corresponding to the model to be tested and the annotation result corresponding to the fine-tuned model, wherein different models to be tested use different numbers of training samples in the pre-training stage, and the fine-tuned model corresponding to the model to be tested is obtained after fine-tuning the model to be tested through a preset annotated data set on the basis of the model to be tested; According to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuned model corresponding to each model to be tested, the consistency evaluation results for each model to be tested and the fine-tuned model corresponding to each model to be tested are determined. The consistency evaluation results are used to indicate whether the evaluation score of the model to be tested is positively correlated with the evaluation score of the fine-tuned model corresponding to the model to be tested.

13. The method according to claim 12, wherein the consistency evaluation results for each model to be tested and the fine-tuned model corresponding to each model to be tested are determined according to the evaluation score corresponding to each model to be tested and the evaluation score of the fine-tuned model corresponding to each model to be tested, specifically comprising: For any two models to be tested, if the size relationship between the evaluation scores corresponding to the two models to be tested is the same as the size relationship between the evaluation scores of the fine-tuning models corresponding to the two models to be tested, then the two models to be tested are determined to be consistently paired; otherwise, the two models to be tested are determined to be inconsistently paired; According to the number of determined consistent pairs, the consistency evaluation results for each model to be tested and the fine-tuning model corresponding to each model to be tested are determined.

14. A model evaluation device, the device being used to evaluate at least one pre-trained model to be tested, comprising: A first acquisition module is used to obtain a test request for a model to be tested; A task determination module, used to determine the test task to be performed on the model to be tested according to the test request; A second determination module is used for obtaining a specified question text from a preset sample data set and obtaining a question-answering example template for the specified question text if it is determined that the test task is an open task, wherein the open task is used to enable the model to be tested to output an answer based on the question text that does not include answer options, and the question-answering example template includes an example question text, an example parsing text, and an example answer text, wherein the example parsing text is used to represent the reasoning process based on which the answer to the example question is based; A combination module, used for combining the question and answer example template and the specified question text to obtain a target test text; The evaluation module is used to input the target test text into the model to be tested, obtain the output result of the model to be tested, and evaluate the model to be tested according to the output result.

15. A model evaluation device, the device being used to evaluate at least one pre-trained model to be tested, comprising: A first acquisition module is used to obtain a test request for a model to be tested; A task determination module, used to determine the test task to be performed on the model to be tested according to the test request; A second acquisition module is used to acquire an original test text from a preset sample data set if it is determined that the test task is a multiple-choice task, wherein the original test text includes a specified question text that the model to be tested needs to answer, and answer options corresponding to the specified question text, and the multiple-choice task is used to enable the model to be tested to select an answer to the question to be answered from the preset answer options; A rewriting module, for rewriting the specified question text according to the answer text corresponding to each answer option, to obtain a rewritten text containing the answer text corresponding to the answer option as the rewritten text corresponding to the answer option; The evaluation module is used to input the rewritten text corresponding to each answer option as the target test text into the model to be tested, obtain the output result of the model to be tested, and evaluate the model to be tested based on the output result.

16. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 7 or 8 to 13 by running the executable instructions.

17. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method as claimed in any one of claims 1 to 7 or 8 to 13.

18. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7 or 8 to 13.

Citation Information

Cited By

  • Large model automatic evaluation method, device and equipment and readable storage medium

    CN122021937A

  • Large model automatic evaluation method, device and equipment and readable storage medium

    CN122021937B