Method, device and equipment for model evaluation and storage medium
By providing sample input for the evaluation set of machine learning models to obtain predicted outputs and determining performance levels based on errors, the problem of difficulty in comprehensively and accurately evaluating machine learning models in the prior art is solved, especially in code Q&A tasks, and higher evaluation accuracy and comprehensiveness are achieved.
Patent Information
- Application Number
- CN202311745475.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to comprehensively and accurately evaluate machine learning models, especially in code Q&A tasks. The lack of suitable evaluation sets and methods leads to inaccurate evaluation results.
By obtaining a set of evaluations for the target model, sample input is provided for different evaluation tasks to obtain the predicted output, and model performance levels are determined based on the error between the sample output and the predicted output.
It realizes comprehensive and accurate evaluation of the performance of machine learning models, especially in code Q&A tasks, which improves the comprehensiveness and accuracy of the evaluation.
Smart Images

Figure CN120179540A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computer technologies, and particularly to methods, apparatuses, devices, and computer-readable storage media for model evaluation. Background Art
[0002] As machine learning and deep learning technologies have been widely applied in many fields. A machine learning model can be configured and trained to generate corresponding model outputs based on given model inputs. The machine learning model can be applied to fields such as natural language processing, machine translation, speech synthesis, image generation, etc., and in the future, it can be applied to more fields, such as medical, financial, educational, etc. After the model is trained, it is usually desirable to accurately evaluate the performance of the model to determine the training effect, subsequent application methods, etc. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for model evaluation is provided. The method includes: obtaining an evaluation set for a target model, where the target model is at least configured to generate code-related answers based on questions, and the evaluation set includes sample inputs and sample outputs for the target model; for each evaluation task in at least one evaluation task, respectively providing the sample inputs in the evaluation set corresponding to the respective evaluation tasks to the target model to obtain predicted outputs of the target model, where at least one evaluation task includes at least one of the following: a first evaluation task for determining the ability of the target model to predict another part of a sample code generation scheme given a sample question and a part of the sample code generation scheme, and a second evaluation task for determining the ability of the target model to generate function code related to the sample question given the sample question; and for each evaluation task in at least one evaluation task, determining a performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for model evaluation. The apparatus includes: an evaluation set acquisition module configured to obtain an evaluation set for a target model, the target model being at least configured to generate code-related answers based on questions, and the evaluation set including sample inputs and sample outputs for the target model; an input providing module configured to, for each evaluation task in at least one evaluation task, respectively provide the sample inputs corresponding to the respective evaluation tasks in the evaluation set to the target model to obtain predicted outputs of the target model, where at least one evaluation task includes at least one of the following: a first evaluation task for determining the ability of the target model to predict another part in a sample code generation scheme given a sample question and a part of the sample code generation scheme, and a second evaluation task for determining the ability of the target model to generate function code related to the sample question given the sample question; and a performance determination module configured to, for each evaluation task in at least one evaluation task, determine the performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0007] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0008] In combination with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;
[0010] Figure 2 A block diagram showing a model evaluation architecture according to some embodiments of the present disclosure;
[0011] Figure 3Shows a schematic diagram of a model evaluation process according to some embodiments of the present disclosure;
[0012] Figure 4 Shows a flowchart of a process for model evaluation according to some embodiments of the present disclosure;
[0013] Figure 5 Shows a block diagram of an apparatus for model evaluation according to some embodiments of the present disclosure; and
[0014] Figure 6 Shows an electronic device in which one or more embodiments of the present disclosure can be implemented. Detailed Description of the Embodiments
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.
[0017] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to relevant laws and regulations.
[0019] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solution of the present disclosure according to the prompt message.
[0020] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way to send a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It can be understood that the above-mentioned notification and user authorization process are only illustrative and do not constitute a limitation on the implementation of the present disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0022] As used herein, the term "model" can learn the corresponding relationship between input and output from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides corresponding output by using multiple processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.
[0023] A "neural network" is a machine learning network based on deep learning. A neural network can process input and provide corresponding output, and it usually includes an input layer, an output layer, and one or more hidden layers between the input layer and the output layer. The neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.
[0024] Generally, machine learning can roughly include three stages, namely the training stage, the testing stage, and the application stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data, continuously iterating and updating the parameter values until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered capable of learning the association from input to output (also known as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. Sometimes the testing stage can be incorporated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained from training and determine the corresponding model output.
[0025] Model evaluation is mainly to test the performance or ability of the trained machine learning model to determine whether to continue to adjust the model and whether and / or how to apply the model.
[0026] Language models (LMs), especially large language models (LLMs), have certain question - answering capabilities. They can output corresponding answers based on the input questions. Code large language models can specialize in code - related tasks, such as code completion, natural language to code, etc., and they can help improve the efficiency of writing code. Traditionally, it is only possible to evaluate the performance of code large language models in tasks such as code completion and natural language to code, but it is impossible to evaluate the performance of code large language models in code question - answering tasks (which can also be called free - form question - answering tasks). In code question - answering tasks, the questions can include any form of content such as natural language and programming languages, and the corresponding model outputs (i.e., answers) can also include any form of content such as natural language and programming languages (e.g., code).
[0027] Some automated evaluation schemes have been proposed. These schemes mainly rely on test sets for various tasks, and the test samples in these test sets include model inputs and corresponding reference model outputs. For example, some classic test sets for language models will adopt multiple - choice questions or question - answering forms (including short answers), where the questions are the model inputs and the answers are the model outputs. The performance of the model is evaluated by calculating the accuracy of the model on the test set or the best linear unbiased estimator (BLUE) value between the generated answer and the standard answer. However, this evaluation scheme based on classic test sets will face the problem that the tasks covered by the test sets are not comprehensive enough, especially the lack of test sets for evaluating code question - answering tasks. In code question - answering tasks, there is no standard and unique reference output corresponding to the questions. If the evaluation method is to compare the answers output by the model with the reference answers in the test set, then the evaluation results are difficult to be accurate and comprehensive.
[0028] There are also some solutions that propose to use well-known advanced generative models to test other models. Specifically, a test set including model inputs and reference outputs is provided to the advanced model, and the advanced model is allowed to determine whether the output of the model to be evaluated conforms to the reference output provided by the test set. However, such an evaluation solution is still based on a test set in the <model input, reference model output> format. Therefore, this solution will also face the above problems, that is, the correct model output is usually not unique, and it is insufficient to compare it only with the reference model output. In addition, the accuracy of the results of the automatic evaluation by the advanced generative model also depends on the performance of the model itself, the setting of the prompt words provided to the model, etc. Because as a generative model, different questioning schemes may lead to different judgment accuracies.
[0029] In an embodiment of the present disclosure, an improved model evaluation scheme is proposed. In this scheme, an evaluation set for the target model is obtained. For each evaluation task in at least one evaluation task, the sample inputs corresponding to the respective evaluation tasks in the evaluation set are respectively provided to the target model to obtain the predicted output of the target model. For each evaluation task in at least one evaluation task, based on the error between the sample output of the respective evaluation task and the obtained predicted output, the performance level of the target model is determined. In this way, a test set that can be used for automatic evaluation can be established for various evaluation tasks, without being limited by the complexity of the model to be evaluated, input requirements, etc. For each evaluation task, the model performance of the target model can be evaluated based on the corresponding evaluation rules. The comprehensiveness of the evaluation of the code question-answering model and the accuracy of the automatic evaluation of the model can be improved.
[0030] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0031] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In the environment 100, the model evaluation system 110 is configured to perform an evaluation of the model performance level of the target model 120 to determine the quality of the model output generated by the target model 120.
[0032] The target model 120 is configured to provide a corresponding output 124 by processing an input 122. Depending on the model configuration, the input 122 to the target model 120 may include content in the text modality (i.e., text segments), and the output 124 is also content in the text modality. For example, the target model 120 may be based on a language model and can be used as a natural language processing model or a translation model. Such a model is capable of generating an output text sequence based on the input text sequence. In some of the example embodiments below, for the purpose of discussion, a text generation model is used as an example for illustration. In some embodiments, the target model 120 may be configured to support an interaction in the form of code question and answer, where the input 122 may be considered as a question, and the output 124 is considered as the answer generated by the target model 120 for the input question.
[0033] In some embodiments, the input 122 to the target model 120 may also include other modalities, such as the image modality, the video modality, the audio modality, etc., and the output 124 may also include other modalities in addition to the text modality, such as the image modality, the video modality, the audio modality. For example, the target model 120 may be an image generation model capable of generating a new image based on the input text and / or image. The target model 120 may also be a speech synthesis model capable of generating speech based on the input text. The target model 120 may be any suitable single-modal or multi-modal content generation model.
[0034] When performing model evaluation, the model evaluation system 110 needs to utilize a test set 130 for evaluation, which includes a plurality of test samples 132-1, 132-2,... 132-N (collectively or individually referred to as test samples 132). Depending on the applied evaluation scheme, the test set 130 may be different. The test set adopted by the model evaluation scheme proposed in the embodiments of the present disclosure will be described in detail below.
[0035] In the environment 100, the model evaluation system 110 may be any type of terminal device or server device. Examples of terminal devices include mobile terminals, fixed terminals or portable terminals, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the model evaluation system 110 is also capable of supporting any type of interface for users (such as "wearable" circuits, etc.). Examples of server devices include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0036] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.
[0037] Figure 2 A block diagram of a model evaluation architecture 200 according to some embodiments of the present disclosure is shown. The model evaluation architecture 200 can be implemented in or included in Figure 1 the model evaluation system 110. The model evaluation architecture 200 involves a data source 210, an evaluation set generation unit 220, an evaluation execution unit 230, and a performance level determination unit 240 to evaluate the performance of a target model 120 to be evaluated.
[0038] The target model 120 is at least configured to generate code-related answers based on questions, where the model input can be considered as the question, and the model output is considered as the answer to the input question. The target model 120 may include a code large model. In an application scenario of code question and answer, the target model 120 is expected to be trained to support free-form code question and answer. The questions provided to the target model 120 may include a mixture of natural language and programming language texts, including code text fragments corresponding to the programming language and natural language text fragments. This can support users to ask questions in the form of natural language mixed with code. In some cases, depending on the specific question, the answers provided by the target model 120 may also include a mixture of natural language and programming language texts. In this way, the target model 120 can use the form of natural language mixed with code to make the answer content more clearly understood by users. Such a code question and answer form is called free-form code question and answer. Of course, in some cases, the questions of the target model 120 may only contain natural language or only contain code, or the answers only contain natural language or only contain code. In addition, on the basis of supporting free-form code question and answer, the target model 120 can also be trained to support traditional tasks such as code completion and natural language to code.
[0039] The evaluation set generation unit 220 is configured to generate an evaluation set for evaluating the performance of the target model 120 based on the data source 210. The data source 210 is a data source related to code question and answer. As shown in the figure, the data source 210 includes one or more data sources 210-1, 210-2, 210-3, ……, 210-N. Different data sources in the data source 210 may come from different fields, such as Front-End, Back-End, Data Science and Machine Learning (DS&ML), Mobile&Desktop, and Information Technology Operations (IT Ops). The data source 210 may include data in any appropriate natural language and any appropriate programming language. Thus, the diversity of the data source can be ensured, and the comprehensiveness of subsequent evaluations can be improved.
[0040] The data source 210 may include a plurality of questions and a plurality of answers corresponding to each question. The evaluation set generation unit 220 may collect a plurality of sample questions from the data source 210. For each sample question among the plurality of sample questions, the evaluation set generation unit 220 may also collect a plurality of candidate answers to the corresponding question from the data source 210. The evaluation set generation unit 220 may determine the quality score of each of the plurality of candidate answers, and based on the quality scores of each of the plurality of candidate answers, select a sample answer that matches the corresponding sample question from the plurality of candidate answers. Both the sample questions and the sample answers here may include text fragments in natural language and text fragments in programming language.
[0041] Regarding the specific manner of determining the quality score of a certain candidate answer (for example, the first candidate answer) here, in some embodiments, the data source 210 also includes interactive feedback information (such as like, comment, forward, view, etc.) for different questions and answers. The evaluation set generation unit 220 may determine the quality score of the first candidate answer based on the interactive feedback information corresponding to the first candidate answer in the data source 210. Exemplarily, the more interactions indicated by the interactive feedback information (such as the higher the like, the more forwards, the more comments, the more views, etc.), the higher the quality score of the corresponding first candidate answer. Another example is that the closer the interaction time corresponding to the latest interaction (such as the latest like, the latest comment, etc.) indicated by the interactive feedback information is to the current time, the higher the quality score of the corresponding first candidate answer. It can be understood that the evaluation set generation unit 220 may determine the quality scores of each of the plurality of candidate answers based on a similar manner. The evaluation set generation unit 220 may then determine the candidate answer with the highest corresponding quality score as the sample answer that matches the corresponding sample question.
[0042] Exemplarily, Figure 3 A schematic diagram of a model evaluation process 300 according to some embodiments of the present disclosure is shown. The model evaluation process 300 may be implemented in or included inFigure 1 In the model evaluation system 110. The model evaluation system 110 can collect sample questions 310 from a data source (such as the data source 210). The model evaluation system 110 can also collect multiple candidate answers 320 corresponding to the sample questions 310 from the data source (for example, at least including candidate answer 320-1 and candidate answer 320-2). The model evaluation system 110 can determine the quality scores of the multiple candidate answers 320 respectively. For example, the model evaluation system 110 can determine the quality scores of the multiple candidate answers 320 respectively based on the voting numbers corresponding to the multiple candidate answers 320, where the candidate answer with a larger voting number has a higher corresponding quality score. The model evaluation system 110 can then select the sample answer 320-1 that matches the corresponding sample question 310 from the multiple candidate answers 320 based on the quality scores of the multiple candidate answers 320, and construct a sample output 325 in the evaluation set based on the sample answer.
[0043] Return reference Figure 2 After the evaluation set generation unit 220 obtains multiple sample questions and the sample answers corresponding to each sample question, it can generate multiple sample question-sample answer pairs. In some embodiments, the evaluation set generation unit 220 can obtain at least one evaluation task 202, and for each evaluation task 202, generate an evaluation set corresponding to the evaluation task 202 according to the multiple sample question-sample answer pairs. The evaluation set includes sample inputs and sample outputs for the target model 120 to be evaluated.
[0044] The at least one evaluation task 202 here can include the first evaluation task 202-1. The first evaluation task 202-1 can also be referred to as a filling task. The first evaluation task 202-1 is used to determine the ability of the target model to predict another part of the sample code generation scheme given a sample question and a part of the sample code generation scheme. For the first evaluation task 202-1, the evaluation set generation unit 220 can determine a first sample response for the first evaluation task 202-1 based on the first sample question and the first sample answer in a certain sample question-sample answer pair (such as the first sample question-first sample answer pair). The first sample response can include at least part of the first sample answer, for example.
[0045] The evaluation set generation unit 220 can further determine a sample input corresponding to the first evaluation task 202-1 based on the first sample question and a part of the first sample answer, and determine a sample output corresponding to the first evaluation task 202-1 based on the other parts of the first sample answer except those included in the sample input. For example, in the first prediction task, the sample input can be, for example, "What is the right way to disable back button influtter when one has reached the login page after logging out? … specifically, please don't add other text and repeat the following paragraph with [blank] filled: Use [blank] instead of [blank]", where [blank] indicates the part that needs to be predicted and filled by the target model. The sample output can be, for example, "Use pushAndRemoveUntil / pushAndRemoveUntil() instead of pop / pop()", which can be filled into the [blank] part of the sample input.
[0046] It should be noted that for the first sample question - first sample answer pair, the evaluation set generation unit 220 can determine multiple sample inputs and multiple sample outputs, and these multiple sample inputs can include different parts of the first answer. Thus, the evaluation set generation unit 220 can determine the sample input and sample output corresponding to the first evaluation task 202-1 based on the first sample question - first sample answer pair. The evaluation set generation unit 220 can also determine other sample inputs and other sample outputs corresponding to the first evaluation task 202-1 based on other sample question - sample answer pairs in the multiple sample question - sample answer pairs. The evaluation set generation unit 220 can further determine a first sample set corresponding to the first evaluation task 202-1 for the target model 120 based on all the sample inputs and all the sample outputs corresponding to the first evaluation task 202-1.
[0047] At least one of the evaluation tasks 202 here can also include a second evaluation task 202-2. The second evaluation task 202-2 is used to determine the ability of the target model to generate function code related to a given sample question. The second evaluation task 202-2 is also called a function prediction task, or a unit prediction task.
[0048] For the second evaluation task 202-2, the evaluation set generation unit 220 can obtain function signature information related to the second sample question in a certain sample question-sample answer pair (such as the second sample question-second sample answer pair). The function signature information can include, for example, the function name, function inputs, output formats, and so on. The evaluation set generation unit 220 can determine the sample input corresponding to the second evaluation task 202-2 based on the second sample question and the function signature information related to the second sample question, and determine the sample output corresponding to the second evaluation task 202-2 based on the test cases that match the function signature information. For example, in the second prediction task, the sample input can be, for example, "Complete the function ```vector<pair<int,int>>origin_to_goal(vector <int>origin)```.The function create…Example:…”。Sample output could be, for example, "int main(){vector <int>origin =... vector<pair<int, int>> goal = origin_to_goal(origin); / / assert......”.
[0049] Similarly, the evaluation set generation unit 220 can also determine other sample inputs and other sample outputs corresponding to the second evaluation task 202-2 based on other sample question-sample answer pairs among the multiple sample question-sample answer pairs, function signature information respectively related to the multiple other sample questions, and multiple test cases respectively matching the multiple function signature information. Furthermore, the evaluation set generation unit 220 can determine a second sample set corresponding to the second evaluation task 202-2 for the target model 120 based on all the sample inputs and all the sample outputs corresponding to the second evaluation task 202-2.
[0050] At least one of the evaluation tasks 202 here can also include a third evaluation task 202-3. The third evaluation task 202-3 is used to evaluate the matching degree between the keywords in the answer generated by the target model for a given question and the sample keywords. The third evaluation task 202-3 is also called the keyword matching task. Under the third evaluation task, the performance of the target model can be evaluated from the dimension of keyword matching. When at least one sample keyword in the third sample answer in a certain sample question-sample answer pair (such as the third sample question-third sample answer pair) is labeled, the evaluation set generation unit 220 can determine the sample input corresponding to the third evaluation task 202-3 based on the third sample question, and determine the sample output corresponding to the third evaluation task 202-3 based on the corresponding third sample answer. For example, under the third prediction task, the sample input can be, for example, "How can I increase the laravel 8dd() limitations? I have lotsof nested relationship and I can’t view most of them due to many data.”. The corresponding sample output can identify keywords such as "override" and "dumper / dump" in the sample answer corresponding to the sample question.
[0051] Similarly, when at least one sample keyword in the sample answers of other sample question-sample answer pairs among multiple sample question-sample answer pairs is labeled, the evaluation set generation unit 220 can also determine other sample inputs and other sample outputs corresponding to the third evaluation task 202-3 based on the other sample question-sample answer pairs. Further, the evaluation set generation unit 220 can determine a third sample set for the target model 120 corresponding to the third evaluation task 202-3 based on all the sample inputs and all the sample outputs corresponding to the third evaluation task 202-3.
[0052] At least one of the evaluation tasks 202 here may further include a fourth evaluation task 202-4. The fourth evaluation task 202-4 is used to determine the text similarity between the answer generated by the target model for a given question and the sample answer. The evaluation set generation unit 220 can directly determine the sample input corresponding to the fourth evaluation task 202-4 based on the fourth sample question in a certain sample question-sample answer pair (such as the fourth sample question-fourth sample answer pair), and determine the sample output corresponding to the fourth evaluation task 202-4 based on the corresponding fourth sample answer. For example, in the fourth prediction task, the sample input may be, for example, the sample question "How to enable Dev Tools project on IntelliJ 2021.2 using maven and observe the changes in code without having to restart the Tomcat server?". The sample output may be the text content of the corresponding sample answer.
[0053] Similarly, the evaluation set generation unit 220 can also determine other sample inputs and other sample outputs corresponding to the fourth evaluation task 202-4 based on the other sample question-sample answer pairs among multiple sample question-sample answer pairs. Further, the evaluation set generation unit 220 can determine a fourth sample set for the target model 120 corresponding to the fourth evaluation task 202-4 based on all the sample inputs and all the sample outputs corresponding to the fourth evaluation task 202-4.
[0054] It can be understood that at least one of the evaluation tasks 202 may include a combination of one or more of the above evaluation tasks, and the present disclosure does not limit this. Correspondingly, for at least one evaluation task 202, the evaluation set generation unit 220 can generate at least one sample set corresponding to each evaluation task. In some embodiments, the evaluation set generation unit 220 can generate one sample set, and this sample set includes at least one sub-sample set corresponding to at least one evaluation task 202.
[0055] The evaluation set generation unit 220 can provide the generated evaluation set to the evaluation execution unit 230. For each evaluation task 202 in at least one evaluation task 202, the evaluation execution unit 230 can provide the sample inputs corresponding to the respective evaluation tasks 202 in the evaluation set to the target model 120 to obtain the prediction outputs of the target model 120. In some embodiments, since both the sample problems and the sample answers can include text segments in natural language and text segments in programming languages, the sample inputs and sample outputs in the evaluation set generated by the evaluation set generation unit 220 can also include text segments in natural language and text segments in programming languages. It can be understood that some of the inputs of the target model include text segments in natural language and text segments in programming languages, and / or some of the outputs of the target model include text segments in natural language and text segments in programming languages.
[0056] Exemplarily, as Figure 2 shown, for the first evaluation task 202-1, the evaluation execution unit 230 can provide the first sample input 204-1 corresponding to the first evaluation task 202-1 to the target model 120 and obtain the first prediction output 206-1 of the target model 120 for the first sample input 204-1. Exemplarily, if the first sample input 204-1 is a sample input determined by the evaluation set generation unit 220 based on a first sample problem and a part of the first sample answer, the first prediction output 206-1 can include a prediction of the other parts of the first sample answer except those included in the sample input. For example, if the first sample problem is A, the first sample answer is BCD, and a part of the first sample answer is BD, the first sample input 204-1 can be A-B_D. The first prediction output 206 can be regarded as the target model 120 filling in the blank of "_" in the first sample answer.
[0057] For the second evaluation task 202-2, the evaluation execution unit 230 can provide the second sample input 204-2 corresponding to the second evaluation task 202-2 to the target model 120 and obtain the second prediction output 206-2 of the target model 120 for the second sample input 204-2. Exemplarily, if the second sample input 204-2 is a sample input determined by the evaluation set generation unit 220 based on a second sample problem in a second sample problem - second sample answer pair, function signature information related to the second sample problem, and test cases matching the second function signature information, the second prediction output 206-2 can include a predicted function code generated based on the second sample problem and the function signature information. The second prediction output 206-2 can be regarded as a predicted function generated by the target model 120 based on the second sample input 204-2.
[0058] For the third evaluation task 202-3, the evaluation execution unit 230 can provide the third sample input 204-3 corresponding to the third evaluation task 202-3 to the target model 120, and obtain the third prediction output 206-3 of the target model 120 for the third sample input 204-3. For the fourth evaluation task 202-4, the evaluation execution unit 230 can provide the fourth sample input 204-4 corresponding to the fourth evaluation task 202-4 to the target model 120, and obtain the fourth prediction output 206-4 of the target model 120 for the fourth sample input 204-4.
[0059] For each evaluation task in at least one evaluation task, the prediction output output by the target model 120 and the sample output corresponding to the evaluation task will be provided to the performance level determination unit 240. The performance level determination unit 240 can determine the performance level of the target model 120 based on the error between the sample output of the corresponding evaluation task and the obtained prediction output. It can be understood that the smaller the error, the higher the performance level of the target model 120.
[0060] Exemplarily, as Figure 3 shown, the model evaluation system 110 can determine the sample inputs corresponding to the target model 120 and at least one evaluation task respectively based on the sample question 310 and the corresponding sample answer 325. For a certain evaluation task, the model evaluation system 110 can provide the sample inputs corresponding to the evaluation task to the target model 120 respectively to obtain the prediction output 330 of the target model 120. The model evaluation system 110 can then determine the performance level 340 (e.g., performance level: X) of the target model 120 based on the prediction output 330 and the sample output corresponding to the evaluation task. For example, for the third prediction task, if the prediction output 330 does not contain the keyword 325 in the sample answer, then the score of the target model 120 in the third prediction task will be low.
[0061] It can be understood that for different evaluation tasks, the performance level determination unit 240 can determine the performance level of the target model 120 based on different evaluation rules. For the first evaluation task 202-1, the first predicted output 206-1 is the prediction of the target model 120 for the parts other than those included in the sample input in the first sample answer, and the first sample output includes the parts other than those included in the sample input in the first sample answer. The performance level determination unit 240 can determine the first error between the two by comparing the first predicted output 206-1 and the first sample output (which can also be understood as comparing the filled content output by the target model 120 and the real content). The performance level determination unit 240 can determine the first error, for example, based on matching keywords, regular expressions, recursive logical expressions constructed based on the matching results, etc. The first error can, for example, indicate the instruction following ability of the target model 120. The performance level determination unit 240 can then determine the performance level of the target model 120 for the first evaluation task 202-1 based on the first error.
[0062] For the second evaluation task 202-2, the second predicted output 206-2 is the predicted function code output by the target model 120 based on the second sample question and function signature information, and the second sample output includes test cases matching the function signature information. The performance level determination unit 240 can determine the second error between the two by comparing the second predicted output 206-2 and the second sample output (which can also be understood as comparing the function code generated by the target model 120 and the test cases). The second error can, for example, indicate the response generation ability of the target model 120. In some embodiments, the error between the sample output corresponding to the second evaluation task and the predicted output is based on whether the predicted function code can run correctly on the test cases. That is, the sample output is a test case for the function, and when evaluating the model, it is verified whether the test case can run correctly on the predicted function code. If the predicted function code can run correctly on the test cases, the second predicted output 206-2 and the second sample output output by the target model 120 are similar; if the predicted function code cannot run correctly on the test cases, the second predicted output 206-2 and the second sample output output by the target model 120 are not similar. The performance level determination unit 240 can then determine the performance level of the target model 120 for the second evaluation task 202-1 based on the second error.
[0063] For the third evaluation task 202-3 and the fourth evaluation task 202-4, the prediction outputs (i.e., the third prediction output 206-3 and the fourth prediction output 206-4) output by the target model 120 can both include the predicted answers corresponding to the sample questions in the sample input. For the third evaluation task 202-3, the performance level determination unit 240 can detect whether at least one keyword in the sample answers exists in the third prediction output 206-3 to determine the performance level of the target model 120 for the third evaluation task 202-3. It can be understood that the more the number of at least one keyword existing in the third prediction output 206-3, the higher the performance level of the target model 120 for the third evaluation task 202-3. For the fourth evaluation task 202-4, the performance level determination unit 240 can directly compare the fourth prediction output 206-4 with the corresponding sample output to determine the performance level of the target model 120 for the fourth evaluation task 202-4. It can be understood that the more similar the fourth prediction output 206-4 is to the corresponding sample output, the higher the performance level of the target model 120 for the fourth evaluation task 202-4.
[0064] In some embodiments, for each evaluation task, the performance level determination unit 240 can generate a prompt input provided to the third machine learning model based on the evaluation rules, the model input, and the model output generated by the target model 120, so as to correctly guide another machine learning model (e.g., the second machine learning model) outside the target model 120 to provide a quality evaluation related to the model output as expected. The performance level determination unit 240 can construct the prompt input provided to the second machine learning model according to any appropriate prompt engineering technique. In some embodiments, the prompt input can also specify the format of the result of its output to the second machine learning model, that is, the representation method of the quality of the model output under the evaluation rules. For example, it can be required that the second machine learning model determine whether the model output is the correct model output corresponding to the model input or the wrong model output based on the evaluation rules. For another example, it can be required that the second machine learning model score the model output based on the evaluation rules and give constraints on the score range, and so on.
[0065] In summary, in the embodiments of the present disclosure, a test set for automatic evaluation can be established for various evaluation tasks without being limited by the complexity of the model to be evaluated, input requirements, etc. For each evaluation task, the model performance of the target model can be evaluated based on the corresponding evaluation rules. The comprehensiveness of the evaluation and the accuracy of automatic evaluation of the model can be improved.
[0066] Figure 4 The flowchart of a process 400 for model evaluation according to some embodiments of the present disclosure is shown. The process 400 can be implemented in Figure 1 the model evaluation system 110 or Figure 2 at the model evaluation architecture 200. In the following, for purposes of explanation, reference is made to Figure 1 to describe the process 400.
[0067] In block 410, the model evaluation system 110 obtains an evaluation set for a target model, the target model being at least configured to generate code-related answers based on questions, the evaluation set including sample inputs and sample outputs for the target model.
[0068] In block 420, for each evaluation task in at least one evaluation task, the model evaluation system 110 provides the sample inputs corresponding to the respective evaluation tasks in the evaluation set to the target model respectively to obtain predicted outputs of the target model. The at least one evaluation task includes at least one of the following: a first evaluation task for determining the ability of the target model to predict another part in a sample code generation scheme given a sample question and a part of the sample code generation scheme; a second evaluation task for determining the ability of the target model to generate function code related to a sample question given the sample question.
[0069] In block 430, for each evaluation task in at least one evaluation task, the model evaluation system 110 determines a performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
[0070] In some embodiments, the sample input corresponding to the first evaluation task includes a first sample question and a part of a first sample answer, the sample output corresponding to the first evaluation task includes the other part of the first sample answer except for what is included in the sample input, and the predicted output corresponding to the first evaluation task includes a prediction of the other part of the first sample answer except for what is included in the sample input.
[0071] In some embodiments, the sample input corresponding to the second evaluation task includes a second sample question and function signature information related to the second sample question, the sample output corresponding to the second evaluation task includes a test case matching the function signature information, and the predicted output corresponding to the second evaluation task includes predicted function code generated based on the second sample question and the function signature information. The error between the sample output and the predicted output corresponding to the second evaluation task is based on whether the predicted function code runs correctly on the test case.
[0072] In some embodiments, the at least one evaluation task further includes at least one of the following: a third evaluation task for evaluating the matching degree between keywords in the answer generated by the target model for a given question and sample keywords; a fourth evaluation task for determining the text similarity between the answer generated by the target model for a given question and a sample answer.
[0073] In some embodiments, the sample input corresponding to the third evaluation task includes a third sample question, and the sample output corresponding to the third evaluation task includes a third sample answer that matches the third sample question, and at least one sample keyword in the third sample answer is marked.
[0074] In some embodiments, the sample input corresponding to the fourth evaluation task includes a fourth sample question, and the sample output corresponding to the fourth evaluation task includes a fourth sample answer that matches the fourth sample question.
[0075] In some embodiments, a partial input of the target model includes text fragments in natural language and text fragments in programming language, and / or a partial output of the target model includes text fragments in natural language and text fragments in programming language.
[0076] In some embodiments, multiple sample inputs and multiple sample outputs in the evaluation set are determined based on multiple sample questions and multiple sample answers, and the multiple sample questions and multiple sample answers are obtained in the following manner: collecting multiple sample questions from a data source related to code question and answer; for each sample question in the multiple sample questions, collecting multiple candidate answers corresponding to the question from the data source; determining the quality score of each of the multiple candidate answers; and based on the quality scores of each of the multiple candidate answers, selecting a sample answer that matches the corresponding sample question from the multiple candidate answers.
[0077] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. Figure 5 The block diagram of an apparatus 500 for model evaluation according to some embodiments of the present disclosure is shown. The apparatus 500 can be implemented as or included in Figure 1 the model evaluation system 110 of Figure 2 or the model evaluation architecture 200 of
[0078] As shown in the figure, the apparatus 500 includes an evaluation set acquisition module 510 configured to obtain an evaluation set for a target model, where the target model is at least configured to generate code-related answers based on questions, and the evaluation set includes sample inputs and sample outputs for the target model. The apparatus 500 further includes an input providing module 520 configured to provide, for each evaluation task in at least one evaluation task, the sample inputs corresponding to the respective evaluation tasks in the evaluation set to the target model to obtain predicted outputs of the target model, where the at least one evaluation task includes at least one of the following: a first evaluation task for determining the ability of the target model to predict another part in a sample code generation scheme given a sample question and a part of the sample code generation scheme; a second evaluation task for determining the ability of the target model to generate function code related to a sample question given the sample question. The apparatus 500 further includes a performance determination module 530 configured to determine, for each evaluation task in at least one evaluation task, the performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
[0079] In some embodiments, the sample input corresponding to the first evaluation task includes a first sample question and a part of a first sample answer, the sample output corresponding to the first evaluation task includes the other parts of the first sample answer except those included in the sample input, and the predicted output corresponding to the first evaluation task includes a prediction of the other parts of the first sample answer except those included in the sample input.
[0080] In some embodiments, the sample input corresponding to the second evaluation task includes a second sample question and function signature information related to the second sample question, the sample output corresponding to the second evaluation task includes a test case that matches the function signature information, and the predicted output corresponding to the second evaluation task includes predicted function code generated based on the second sample question and the function signature information. The error between the sample output and the predicted output corresponding to the second evaluation task is based on whether the predicted function code runs correctly on the test case.
[0081] In some embodiments, the at least one evaluation task further includes at least one of the following: a third evaluation task for evaluating the matching degree between keywords in the answer generated by the target model for a given question and sample keywords; a fourth evaluation task for determining the text similarity between the answer generated by the target model for a given question and a sample answer.
[0082] In some embodiments, the sample input corresponding to the third evaluation task includes a third sample question, the sample output corresponding to the third evaluation task includes a third sample answer that matches the third sample question, and at least one sample keyword in the third sample answer is labeled.
[0083] In some embodiments, the sample input corresponding to the fourth evaluation task includes a fourth sample question, and the sample output corresponding to the fourth evaluation task includes a fourth sample answer that matches the fourth sample question.
[0084] In some embodiments, a partial input of the target model includes text fragments in natural language and text fragments in a programming language, and / or a partial output of the target model includes text fragments in natural language and text fragments in a programming language.
[0085] In some embodiments, the multiple sample inputs and multiple sample outputs in the evaluation set are determined based on multiple sample questions and multiple sample answers. The apparatus 500 further includes a sample acquisition module configured to obtain the multiple sample questions and multiple sample answers in the following manner: collect multiple sample questions from a data source related to code Q&A; for each sample question in the multiple sample questions, collect multiple candidate answers for the corresponding question from the data source; determine the quality score of each of the multiple candidate answers; and based on the quality scores of each of the multiple candidate answers, select a sample answer that matches the corresponding sample question from the multiple candidate answers.
[0086] The units and / or modules included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 500 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0087] Figure 6 The block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 6 The electronic device 600 shown is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to implement Figure 1 of Figure 1 the model evaluation system 110, Figure 2 the model evaluation architecture 200, or Figure 5 the apparatus 500.
[0088] As Figure 6 As shown, the electronic device 600 is in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 600.
[0089] The electronic device 600 generally includes multiple computer storage media. Such media can be any available media accessible to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data and can be accessed within the electronic device 600.
[0090] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6 , a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.
[0091] The communication unit 640 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 600 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 600 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0092] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) as needed through the communication unit 640. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0093] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0094] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0095] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0096] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of this technology without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art in the field to understand the various implementation manners disclosed herein.< / int> < / int>
Claims
1. A method for model evaluation, comprising: Obtain an evaluation set for a target model, where the target model is at least configured to generate code-related answers based on questions, and the evaluation set includes sample inputs and sample outputs for the target model; For each evaluation task among at least one evaluation task, respectively provide the sample inputs in the evaluation set corresponding to the respective evaluation tasks to the target model to obtain the predicted outputs of the target model, where the at least one evaluation task includes at least one of the following: A first evaluation task for determining the ability of the target model to predict another part in the sample code generation scheme given a sample question and a part of the sample code generation scheme, The second evaluation task for determining the ability of the target model to generate function code related to the sample question given the sample question; And For each evaluation task among the at least one evaluation task, determine the performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
2. The method according to claim 1, wherein the sample input corresponding to the first evaluation task includes a first sample question and a part of a first sample answer, the sample output corresponding to the first evaluation task includes the other part of the first sample answer except for the part included in the sample input, and wherein the predicted output corresponding to the first evaluation task includes a prediction of the other part of the first sample answer except for the part included in the sample input.
3. The method according to claim 1, wherein the sample input corresponding to the second evaluation task includes a second sample question and function signature information related to the second sample question, the sample output corresponding to the second evaluation task includes a test case matching the function signature information, and wherein the predicted output corresponding to the second evaluation task includes a predicted function code generated based on the second sample question and the function signature information, and the error between the sample output and the predicted output corresponding to the second evaluation task is based on whether the predicted function code runs correctly on the test case.
4. The method according to claim 1, wherein the at least one evaluation task further includes at least one of the following: A third evaluation task for evaluating the matching degree between keywords in the answer generated by the target model for a given question and sample keywords; A fourth evaluation task for determining the text similarity between the answer generated by the target model for a given question and a sample answer.
5. The method according to claim 4, wherein the sample input corresponding to the third evaluation task includes a third sample question, the sample output corresponding to the third evaluation task includes a third sample answer matching the third sample question, and at least one sample keyword in the third sample answer is labeled.
6. The method according to claim 4, wherein the sample input corresponding to the fourth evaluation task includes a fourth sample question, and the sample output corresponding to the fourth evaluation task includes a fourth sample answer matching the fourth sample question.
7. The method according to claim 1, wherein a partial input of the target model includes text segments in natural language and text segments in programming language, and / or wherein a partial output of the target model includes text segments in natural language and text segments in programming language.
8. The method according to claim 1, wherein a plurality of sample inputs and a plurality of sample outputs in the evaluation set are determined based on a plurality of sample questions and a plurality of sample answers, and the plurality of sample questions and the plurality of sample answers are obtained by the following means: Collecting a plurality of sample questions from data sources related to code question and answer; For each of the multiple sample questions, Collect multiple candidate answers for the corresponding question from the data source; Determine the quality scores of the multiple candidate answers respectively; And Based on the quality scores of the multiple candidate answers, select a sample answer that matches the corresponding sample question from the multiple candidate answers.
9. An apparatus for model evaluation, comprising: An evaluation set acquisition module configured to obtain an evaluation set for a target model, where the target model is at least configured to generate code-related answers based on questions, and the evaluation set includes sample inputs and sample outputs for the target model; An input providing module configured to, for each evaluation task among at least one evaluation task, respectively provide the sample inputs in the evaluation set corresponding to the respective evaluation tasks to the target model to obtain the predicted outputs of the target model, where the at least one evaluation task includes at least one of the following: a first evaluation task for determining the ability of the target model to predict another part in the sample code generation scheme given a sample question and a part of the sample code generation scheme, the second evaluation task for determining the ability of the target model to generate function code related to the sample question given the sample question; And A performance determination module configured to, for each evaluation task among the at least one evaluation task, determine the performance level of the target model based on the error between the sample output of the respective evaluation task and the obtained predicted output.
10. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the device to perform the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Model evaluation method, device and platform and computer storage medium
CN121436024A