Model Evaluation Method, Device, Storage Medium and Program Product
By identifying the type of question and calling the appropriate lightweight or large language model for AI model evaluation, the problem of large resource consumption in the existing technology is solved, and the balance of resource, accuracy and time efficiency is achieved.
Patent Information
- Application Number
- CN202510008540.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-03
AI Technical Summary
When the existing AI model evaluation method is called on a large number of evaluation questions, the resource consumption is large, and when evaluating questions with high complexity, it is necessary to use a model with large parameters, resulting in waste of resources.
By identifying the type of question and calling the appropriate lightweight or large language models for evaluation, the lightweight model is used for objective questions with lower complexity, and the large language model is used for subjective questions with higher complexity, thus finding a balance between resource consumption and assessment accuracy.
It effectively reduces the resource cost required for AI model evaluation, improves the accuracy and efficiency of evaluation, and achieves a balance between resource consumption, accuracy and time efficiency.
Smart Images

Figure CN119416830B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model evaluation method, device, storage medium, and program product. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, AI technology is increasingly applied to various complex decision-making tasks. For example, in multiple industries such as education, exam evaluation, public affairs, and law, a large number of questions and answers can be learned based on AI technology, and an AI model capable of answering specified questions can be trained, such as a question-and-answer model, an expert model, a customer service model, etc.
[0003] To improve the performance of the AI model in the question-answering task, an evaluation model is usually used to evaluate the AI model. A commonly used evaluation method is to obtain the response results output by the AI model on a large number of evaluation questions, and assemble the evaluation questions and response results into a prompt, and call the evaluation model in the form of a dialogue question and answer. The evaluation model can execute the evaluation task according to the prompt and output the evaluation result. Usually, the evaluation model is a high-performance model with more parameters, and when the evaluation model is called on a large number of evaluation questions, it often consumes more resources. Therefore, there is a need to propose a new solution. Summary of the Invention
[0004] Multiple aspects of this application provide a model evaluation method, device, storage medium, and program product to reduce the resource cost required for testing the AI model.
[0005] An embodiment of this application provides a model evaluation method, including: in response to a model evaluation operation, obtaining a response result output by a target model for a target question; obtaining the question type corresponding to the target question; calling a target evaluation model adapted to the question type among multiple evaluation models to evaluate the response result, and obtaining an evaluation result of the target model on the target question; among the multiple evaluation models, different evaluation models correspond to different question types, and the number of parameters of different evaluation models is different.
[0006] Optionally, calling a target evaluation model adapted to the question type among multiple evaluation models to evaluate the response result includes: if the question type is an objective question, calling a first evaluation model to evaluate the response result corresponding to the target question, where the first evaluation model is a lightweight model with the number of parameters less than a set first threshold.
[0007] Optionally, call the first evaluation model to evaluate the response result corresponding to the target question, and obtain the evaluation result of the target model on the target question, including: calling the first evaluation model, and evaluating the response result corresponding to the target question according to the rule engine corresponding to the objective question, so as to obtain the evaluation result of the target model on the target question; the rule engine is used to match the response result with the reference answer of the target question, and calculate the score of the response result according to the matching result and the predefined scoring rules.
[0008] Optionally, call a target evaluation model adapted to the type of the target question to evaluate the response result, including: if the question type is a subjective question, call the second evaluation model to evaluate the response result corresponding to the target question, and the second evaluation model is a large language model with the number of parameters greater than a set second threshold.
[0009] Optionally, call the second evaluation model to evaluate the response result corresponding to the target question, including: generating a prompt word adapted to the question type according to the question type, the target question, and the corresponding response result; the structures and / or contents of the prompt words corresponding to different question types are different; calling the second evaluation model to evaluate the response result corresponding to the target question according to the prompt word.
[0010] Optionally, generating a prompt word adapted to the question type according to the question type, the target question, and the corresponding response result, including: if the question type is a subjective question of the knowledge Q&A type, generating a prompt word adapted to the subjective question of the knowledge Q&A type according to the target question, the response result, the reference answer of the question, and the scoring rules; if the question type is a subjective question of the content creation type, generating a prompt word for the subjective question of the content creation type according to the creation question, the creation result output by the target model for the creation question, the creation requirements, and the scoring rules.
[0011] Optionally, obtaining the question type corresponding to the target question, including: obtaining the question type corresponding to the target question according to the question type setting information of the target question; or calling a question type classifier to identify the question type corresponding to the target question, and the question type classifier is trained on a question type training set using a machine learning algorithm.
[0012] Optionally, before calling the target evaluation model adapted to the question type among multiple evaluation models to evaluate the response result, it further includes: obtaining the matching degree between the response result corresponding to the target question and the target question; if the matching degree is greater than a set matching degree threshold, then call the target evaluation model adapted to the question type among the multiple evaluation models to evaluate the response result.
[0013] An embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the method provided by the embodiment of the present application.
[0014] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the steps in the method provided by the embodiment of the present application.
[0015] An embodiment of the present application further provides a computer program product, including: a computer program / instructions, and when the computer program / instructions are executed by a processor, it can implement the steps in the method provided by the embodiment of the present application.
[0016] In an embodiment of the present application, during the process of evaluating a target model, after obtaining the response result output by the target model for a target question, the question type corresponding to the target question can be obtained, and the target evaluation model adapted to the question type among multiple evaluation models can be called to evaluate the response result. In this implementation manner, among the multiple evaluation models, different evaluation models correspond to different question types, and the number of parameters of different evaluation models is different. Based on this, a solution for collaboratively evaluating the response results output by the target model for different types of questions based on evaluation models with different numbers of parameters is realized, which facilitates using an evaluation model with a lower order of magnitude of the number of parameters to evaluate the response results of some questions with lower complexity, effectively reducing resource consumption, and facilitating using an evaluation model with a higher number of parameters to evaluate the response results of some questions with higher complexity, improving the evaluation accuracy, thereby achieving a balance of accuracy, cost, and time efficiency to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 It is a schematic flowchart of a model evaluation method provided by an exemplary embodiment of the present application;
[0019] Figure 2 It is a schematic structural diagram of an evaluation system provided by an exemplary embodiment of the present application;
[0020] Figure 3 It is a schematic flowchart of a model evaluation method provided by another exemplary embodiment of the present application;
[0021] Figure 4Schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0023] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0024] It should be understood that the term "and / or" used herein is only a relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0025] It should also be noted that the term "comprises", "comprising", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a product or system including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such product or system. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the product or system including the said element.
[0026] The purpose of evaluating an AI model is to assess the performance of the AI model on a specific task to determine whether its task execution ability meets expectations. One way to evaluate an AI model is to obtain the response results output by the AI model on a large number of evaluation questions, and assemble the evaluation questions and response results into a prompt, and call the evaluation model in the form of a dialogue and answer. The evaluation model can execute the evaluation task according to the prompt and output the evaluation result. When the evaluation model executes the task according to the prompt, it needs to split the prompt into small processing units (tokens). The longer the prompt, the more tokens are obtained by splitting. For the evaluation model, it takes more time and resources to process these tokens. In this evaluation method, converting evaluation questions with different question types and their response results into question-and-answer prompts will generate a large number of tokens, thus affecting the evaluation efficiency of the evaluation model and increasing the resource consumption required for evaluation.
[0027] Secondly, in this evaluation method, in order to evaluate the problem-solving ability of the AI model for questions with higher complexity, the number of parameters of the evaluation model is usually increased. The larger the number of parameters of the model, the more resources are consumed to run the model. When using an evaluation model with a larger number of parameters to evaluate the problem-solving ability of the AI model for questions with different complexities, even when evaluating the problem-solving ability of the AI model for questions with lower complexity, it also requires a large amount of resource cost. In addition, in some scenarios, when evaluating some AI models with poor capabilities, such models may not be able to accurately understand or follow the given instructions, resulting in generating interference information unrelated to the questions during the problem-solving process. If an evaluation model with a larger number of parameters is also used to evaluate this AI model in this case, it will cause more unnecessary resource consumption.
[0028] In view of the above technical problems, in some embodiments of the present application, a solution is provided. The following will describe in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.
[0029] Figure 1 is a schematic flowchart of a model evaluation method provided by an exemplary embodiment of the present application. The method may include steps as Figure 1 shown:
[0030] Step 101, in response to a model evaluation operation, obtain the response result output by the target model for the target question.
[0031] Step 102, obtain the question type corresponding to the target question.
[0032] Step 103: Invoke the target evaluation model that matches the question type among multiple evaluation models to evaluate the response result, and obtain the evaluation result of the target model on the target question. Among the multiple evaluation models, different evaluation models correspond to different question types, and the number of parameters of different evaluation models is different.
[0033] The execution entity of this embodiment is an electronic device, which can be, for example, a terminal device such as a mobile phone, a computer, or a tablet computer, or a server device such as a physical server or a cloud server. This embodiment does not make any restrictions. An evaluation tool can run on the electronic device, and this evaluation tool can be a plugin or an independent application program. In this embodiment, the evaluation tool can invoke other evaluation models to perform the evaluation operation on the target model. The following will give an exemplary description in combination with the above steps.
[0034] In step 101, the target model can be any AI model that can output a response result according to the target question. The target model can also be referred to as the participating model or the model to be tested. In different application scenarios, the implementation form of the target model is different. For example, in some application scenarios, the target model can be a knowledge Q&A model, a translation model, a customer service model, an intelligent problem-solving model, a teaching assistant model, a Q&A model related to public affairs, a Q&A model related to enterprise services, and so on.
[0035] In this embodiment, the model evaluation operation can be initiated by a user, or can be initiated by other application programs or system timing events. This embodiment does not make any restrictions. In some embodiments, the evaluation tool can respond to the model evaluation operation and obtain the response result output by the target model for the given evaluation question through the interface provided by the target model. In some evaluation scenarios, the evaluation tool can obtain one or more evaluation sets, and different evaluation sets can correspond to different application fields. For example, in some embodiments, evaluation questions applied in different fields such as academia, engineering, and business can be obtained to generate multiple evaluation sets corresponding to different fields. The evaluation questions in any evaluation set can cover multiple different question types, and each type of evaluation question can cover different difficulty levels to improve the coverage rate of the evaluation set. In some embodiments, for some evaluation questions with standard answers, the standard answers of such evaluation questions can be marked by manual marking or automatic marking, so as to use the standard answers as reference answers in the subsequent evaluation link.
[0036] The evaluation tool can input the evaluation questions in the evaluation set into the target model to obtain the response result output by the target model for the evaluation questions. The evaluation tool can score the response results of the evaluation questions to test the performance of the target model on the problem-solving tasks corresponding to the evaluation questions.
[0037] In this embodiment, any evaluation question will be taken as an example to exemplarily illustrate the model evaluation process. For the convenience of description and distinction, this any evaluation question is described as the target question. The target lateral evaluation question can be any evaluation question in the evaluation set, or can be any question obtained by the target model during operation. In some embodiments, when the target model is implemented as an intelligent problem-solving model or a teaching tutoring model, the target question can be test questions of various different question types. When the question types of the target questions are different, the response results output by the target model are also different. For example, the response result of a multiple-choice question is the target option, the response result of a writing question is a text paragraph, and the response result of a painting question is a picture.
[0038] After obtaining the response result output by the target model for the target question, in step 102, the evaluation tool can obtain the question type corresponding to the target question. In some optional embodiments, the question types that the evaluation tool can recognize may include traditional question types such as objective questions and subjective questions. Among them, an objective question is a question that requires a problem-solving model to make a judgment or choice based on facts or data. For example, it can be a multiple-choice question, a true or false question, etc. Multiple-choice questions can include single-choice questions and multiple-choice questions. A subjective question is a question that requires a problem-solving model to answer based on its own opinions or experiences. For example, knowledge answering questions, writing questions, decision-making guidance questions, etc. Optionally, in addition to traditional question types, the question types that the evaluation tool can recognize may also include some non-traditional question types. For example, scenario questions that make decisions or solve problems based on given information (such as questions related to the return process in the intelligent customer service scenario), matching questions that match different objects, sorting questions that sort processes or timelines, chart analysis questions that analyze charts or graphs, case study questions that deeply analyze cases in combination with case backgrounds, programming questions that use code to solve specific problems, debate questions that debate based on specified controversial topics, etc. This embodiment includes but is not limited to this.
[0039] In some optional embodiments, the evaluation tool can obtain the question type corresponding to the target question according to the question type setting information of the target question. Among them, the question type setting information can be set by a user (such as a tester), or can be marked in the evaluation set. This embodiment does not make any restrictions.
[0040] In some other alternative embodiments, the evaluation tool may adopt a question type classifier to identify the question type corresponding to the target question. Optionally, the question type classifier may be trained on a question type training set using a machine learning algorithm, and the machine learning algorithm may include, but is not limited to: Naive Bayes classifier, Support Vector Machine, etc. In addition to the machine learning method, the question type classifier may also be trained using a rule-based method, such as a classification algorithm based on keyword matching. When training the question type classifier, a large number of questions of various different question types can be collected as sample data, and the sample data is divided into a question type training set and a question type test set. Among them, the question type training set is used to train the question type classifier, and the question type test set is used to evaluate the accuracy rate and recall rate of the question type classifier, which will not be elaborated here.
[0041] After obtaining the question type based on the above implementation manner, in step 103, the evaluation tool may call a target evaluation tool adapted to the question type to evaluate the response result corresponding to the target question, and obtain the evaluation result of the target model on the target question. In this embodiment, the evaluation tool may provide different evaluation models for different question types, and the number of parameters of different evaluation models is different. The parameters of a model refer to the adjustable variables used to learn the data pattern inside the model. Generally, the more parameters a model has, the greater the capacity of the model to capture complex patterns, and thus the better it performs in complex tasks. Correspondingly, the more parameters a model has, the greater the consumption of resources required for the model to run, such as computing resources and storage resources.
[0042] In some alternative embodiments, the number of parameters of different evaluation models is different, which may include: the order of magnitude of the number of parameters of different evaluation models is different. Among them, the order of magnitude of the number of parameters refers to the order of magnitude of the number of parameters in the model, and is usually used to describe the scale and complexity of the model. A model with a high order of magnitude of parameters is more complex and can capture more details and patterns in the data. A model with a lower number of parameters is relatively simple and has good performance in simple tasks. When running models with different orders of magnitude of parameters, a model with a high order of magnitude of parameters requires more computing resources, and a model with a low order of magnitude of parameters requires less computing resources and has a smaller running cost.
[0043] In this embodiment, after identifying the question type corresponding to the target question, the evaluation tool can select an evaluation model adapted to the question type from multiple evaluation models to evaluate the target question and its response results. Among them, the evaluation model being adapted to the question type means that the evaluation model can perform well on the evaluation tasks corresponding to the question type. Among the multiple evaluation models available for the evaluation tool to select, any test model is a high-performance model with a rich knowledge base, excellent reasoning ability, and scalability. The multiple evaluation models may include, but are not limited to, evaluation models respectively corresponding to any two or more of the question types such as objective questions, subjective questions, scenario questions, matching questions, sorting questions, chart analysis questions, case study questions, programming questions, and debate questions. For example, the multiple evaluation models that the evaluation tool can select may include: a first evaluation model corresponding to objective questions, a second evaluation model corresponding to subjective questions, a third evaluation model corresponding to scenario questions, and a fourth evaluation model corresponding to matching questions, etc., which will not be listed one by one.
[0044] In this embodiment, considering that the complexities of the evaluation tasks corresponding to different question types are different, different evaluation models among the multiple evaluation models can be set as evaluation models with different numbers of parameters to balance the resource cost required for model operation and the evaluation accuracy. Some evaluation models with lower orders of magnitude of parameters have lower operation costs and can have better execution effects on tasks with lower complexities. Based on this, evaluation models with lower orders of magnitude of parameters can be used to execute the evaluation tasks of question types with lower complexities, such as the evaluation tasks of objective questions, matching questions, sorting questions, etc., thereby effectively reducing resource consumption. Some evaluation models with higher orders of magnitude of parameters have higher operation costs and can have better execution effects on tasks with higher complexities. Evaluation models with higher orders of magnitude of parameters can be used to execute the evaluation tasks of question types with higher complexities, such as the evaluation tasks of subjective questions, chart analysis questions, case study questions, programming questions, and debate questions, etc., to ensure the evaluation accuracy. For example, in some embodiments, the problem-solving difficulties of objective questions, matching questions, scenario questions, and subjective questions increase in sequence. The number of parameters of the first evaluation model corresponding to objective questions can be set to be less than the number of parameters of the fourth evaluation model corresponding to matching questions, the number of parameters of the fourth evaluation model is less than the number of parameters of the third evaluation model corresponding to scenario questions, and the number of parameters of the third evaluation model is less than the number of parameters of the second evaluation model corresponding to subjective questions.
[0045] The following will take subjective questions and objective questions as examples for further illustrative explanation.
[0046] In some alternative embodiments, when the evaluation tool calls the target evaluation model adapted to the question type corresponding to the target question to evaluate the response result of the target question, if the question type is an objective question, the evaluation tool can call the first evaluation model to evaluate the response result corresponding to the target question, and obtain the evaluation result of the target model on the target question. The first evaluation model is a lightweight model with a parameter quantity less than a set first threshold.
[0047] Optionally, the first threshold can be a threshold in the order of hundreds of billions or tens of billions. For example, in some embodiments, the first evaluation model can be a model of the 7B / 14B level. A model of the 7B / 14B level refers to a model with approximately 7 billion parameters, and a model of the 14B level refers to a large language model with approximately 14 billion parameters. Compared with larger-scale models, the 7B or 14B model is more lightweight, requires less computing resources for training and inference, and has a lower running cost.
[0048] Based on this implementation method, by identifying the question type and calling the first evaluation model according to the question type, the call instruction provided to the first evaluation model can be made clearer. Furthermore, the first evaluation model can accurately understand the call instruction and provide a more accurate evaluation result.
[0049] Objective questions may include, but are not limited to, single-choice questions, multiple-choice questions, true or false questions, etc. This type of question is mainly scored using rules. When the evaluation tool calls the first evaluation model to evaluate objective questions, it can call the first evaluation model and, according to the rule engine corresponding to the objective questions, evaluate the response result corresponding to the target question to obtain the evaluation result of the target model on the target question. Optionally, the first evaluation model can preprocess the response result of the target question to make the description format of the response result more standardized. Then, the first evaluation model can call the rule engine to perform rule scoring on the preprocessed response result corresponding to the target question. In this embodiment, the rule engine is used to match the response result of the target question with the reference answer of the target question and calculate the score of the response result according to the matching result and the predefined scoring rules. Optionally, the rule engine can use at least one of the judgment methods of regular expressions, edit distance, and semantic relevance to match the reference answer and the response result to judge the accuracy of the response result. Optionally, for multiple-choice questions among objective questions, the rule engine can use regular expressions to match the response result of the multiple-choice question with the reference answer to evaluate whether the response result of the multiple-choice question is accurate; for true or false questions among objective questions, the rule engine can use regular expressions to match the response result of the true or false question with the reference answer to evaluate whether the response result of the true or false question is accurate. For fill-in-the-blank questions among objective questions, the rule engine can use methods such as edit distance and semantic relevance to match the response result of the fill-in-the-blank question with the reference answer to evaluate whether the response result of the fill-in-the-blank question is accurate.
[0050] For example, taking a single-choice question as an example, assuming the reference answer is option A, then the rule engine can construct the following regular expression according to the reference answer: reference_pattern_single = r"^[A]$", where the symbol "^" represents the start position of the matching string, and the symbol "$" represents the end position of the matching string to ensure that the entire answer only contains "A". Assuming the user's response result user_answer_single = "A", the rule engine can use the match() function or fullmatch() function to compare the string in the regular expression reference_pattern_single = r"^[A]$" with the string in the response result user_answer_single = "A" to determine whether the string in the response result completely matches the string in the regular expression. Similarly, for multiple-choice questions and true or false questions, the rule engine can construct regular expressions according to the reference answers of multiple-choice questions and true or false questions and perform the above matching method according to the regular expressions and the response results, which will not be elaborated here.
[0051] For another example, taking fill-in-the-blank questions as an example, after the rule engine obtains the reference answers and response results of the fill-in-the-blank questions, it can preprocess the reference answers and response results to obtain the strings corresponding to the reference answers and the strings corresponding to the response results. Among them, the preprocessing includes but is not limited to: case conversion, removing extra spaces, word segmentation, removing stop words, etc. After that, the rule engine can use the edit distance algorithm to quantify the character-level differences between the string corresponding to the response result and the string corresponding to the reference answer, and obtain the edit distance between the response result and the reference answer. Among them, the lower the value of the edit distance, the more similar the two strings are. If the edit distance between the response result and the reference answer is less than a set minimum threshold, the rule engine can determine that the response result is accurate.
[0052] Among them, the scoring rules for different types of objective questions are different. For example, for single-choice questions, if the response result is accurate, 1 point is added. For multiple-choice questions, if all the response results are correct, 3 points are added, and if some are correct, 1 point is added. In some alternative embodiments, the standard answers of the target questions are provided in the test set as reference answers. In this case, the rule engine can directly perform the scoring operation according to the differences between the reference answers and the response results, in combination with the scoring rules. In some other alternative embodiments, when the reference answers of the target questions are not provided in the test set, the first evaluation model can analyze the target questions and generate reference answers that conform to common sense corresponding to the target questions. Optionally, the first evaluation model can parse out the entities in the target question and the relationships between the entities according to the target question and its context information. The first evaluation model can query the knowledge information related to the parsed entities and the relationships between the entities in a rich knowledge base, and based on excellent reasoning ability, generate reference answers according to the parsing results and the queried knowledge information. Furthermore, the rule engine can perform the scoring operation according to the differences between the reference answers generated by the first evaluation model and the response results, in combination with the scoring rules.
[0053] Based on this implementation method, by combining the first evaluation model and the rule engine, it is possible to accurately obtain the evaluation scores even when the number of parameters of the first evaluation model is small, effectively taking into account both the evaluation efficiency and the evaluation cost.
[0054] In some alternative embodiments, when the evaluation tool calls the target evaluation model adapted to the question type corresponding to the target question to evaluate the response result of the target question, if the question type is a subjective question, the evaluation tool can call the second evaluation model to evaluate the response result corresponding to the target question to obtain the evaluation result of the target model on the target question. The second evaluation model is a large language model with the number of parameters greater than a set second threshold.
[0055] Among them, the large language model is a natural language processing (NLP) model trained on a large scale. The large language model is usually built based on deep learning technology and trained on a large-scale training dataset, so as to show powerful performance when processing natural language tasks. Given a piece of text (i.e., context), the large language model will try to predict the most likely next word. The number of parameters of the large language model is greater than a set threshold, and this set threshold is usually on the order of millions or billions. During the training process, these parameters are continuously adjusted and optimized according to the difference between the prediction result and the actual result of the large language model to improve the prediction accuracy of the large language model. In some embodiments, the large language model usually adopts an advanced neural network architecture, such as the Transformer architecture, to build the model structure, which enables the large language model to capture complex patterns in the text and handle long-distance dependencies. After sufficient training, the large language model has powerful generation capabilities and can generate coherent and contextually appropriate text content according to the given prompts.
[0056] In this embodiment, a large language model pre-trained on a large number of general datasets can be used as the base model of the second evaluation model. To adapt to the specific application scenario of this application, that is, the evaluation scenario of the target model for automatic problem-solving, the pre-trained large language model can be fine-tuned on a dataset formed by a large number of various question types and their answer results, so that the fine-tuned large language model is more suitable for performing the task of evaluating the problem-solving ability of the target model. Furthermore, the evaluation tool can obtain more accurate evaluation results based on the advantages of the fine-tuned large language model in the evaluation task.
[0057] Subjective questions may include, but are not limited to: knowledge Q&A and content creation. Among them, knowledge Q&A is an objective question with a reference answer, and content creation is an objective question without a reference answer. For example, content creation can be topics such as writing articles, reports, and drawing. The evaluation focuses of the above two objective questions are different. When evaluating questions related to knowledge Q&A, the evaluation focus is to check the difference between the answer result and the reference answer to evaluate the accuracy of the answer result of the target model, so as to facilitate debugging the target model according to the evaluation result to reduce the hallucination degree of the target model. When evaluating questions related to content creation, the evaluation focus is: evaluating the writing content with a relatively open mind according to the question requirements, writing ideas, and industry general knowledge.
[0058] In some alternative embodiments, when the reference answer of the target question is provided in the evaluation set, the second evaluation model can directly perform the scoring operation according to the difference between the reference answer and the response result, in combination with the pre-learned scoring rules. In some alternative embodiments, when the reference answer of the target question is not provided by the rule engine, the second evaluation model can analyze the target question and generate a reference answer corresponding to the target question. Optionally, the second evaluation model can analyze the entities in the target question and the relationships between the entities according to the target question and its context information, and query the knowledge information related to the analyzed entities and the relationships between the entities in the knowledge base. According to the analysis result and the knowledge information, the second evaluation model can generate a reference answer by using the pre-learned content generation ability. Furthermore, the rule engine can perform the scoring operation according to the difference between the reference answer generated by the second evaluation model and the response result, in combination with the scoring rules.
[0059] Optionally, when performing the scoring operation, the second evaluation model can adopt a rule-based scoring method or a deep learning-based scoring method. For example, the second evaluation model can use regular expressions to match the reference answer and the response result, and output a score according to the matching result. Alternatively, the second evaluation model can use a pre-trained sub-model and a fine-tuned sub-model to score based on the difference between the reference answer and the response result. Taking the rule-based scoring method based on regular expressions as an example, assume that the reference answer to a knowledge-based question is paragraph T0. The second evaluation model can perform word segmentation on paragraph T0 to obtain text units. When constructing a regular expression, the second evaluation model can use the text units obtained by word segmentation as the values of the regular expression, and appropriate wildcards or quantifiers can be added to some text units to allow reasonable variations. For example, the symbol "\s*" can be used to match any number of whitespace characters, or the symbol "[\w\s]*" can be used to match a combination of words and spaces. If synonym replacement is allowed for knowledge-based questions, corresponding synonyms can be added to some text units with synonyms. After obtaining the response result of the knowledge-based question, the second evaluation model can use the constructed regular expression to match the response result to obtain a matching result. The matching result can include the number of overlapping text units between the response result and the values of the regular expression. Among them, the more the number of overlapping text units, the higher the score of the response result. Optionally, the second evaluation model can further calculate the semantic similarity score between the reference answer and the response result, and determine the consistency between the response result and the reference answer according to the matching result of the regular expression and the semantic similarity score. In this case, the first evaluation model can separately set the weight coefficients of the matching result and the semantic similarity score, and calculate the comprehensive score of the response result based on the number of overlapping text units in the matching result, the semantic similarity score, and the weight coefficients. In some alternative embodiments, in order to accurately evaluate the response results corresponding to different subjective questions, the test tool can construct prompt words to prompt the second evaluation model to distinguish different subjective questions when performing the evaluation task, and perform the evaluation tasks corresponding to different subjective questions based on different evaluation methods with different focuses.
[0060] Optionally, the evaluation tool can generate prompt words with different content and / or structures according to the question type to effectively guide the evaluation model to generate evaluation results that better meet the task requirements. Continuing with the target question as an example, optionally, when the evaluation tool calls the second evaluation model to evaluate the response result corresponding to the target question, it can generate prompt words adapted to the question type based on the question type, the target question, and the corresponding response result, and call the second evaluation model to evaluate the response result corresponding to the target question according to the prompt words. Among them, the adaptation of the prompt words to the question type means that the structure of the prompt words is adapted to the question type, and / or the content included in the prompt words is adapted to the task instructions corresponding to the question type, so that the evaluation model can execute the evaluation task corresponding to the question type in the expected manner.
[0061] Optionally, when the question type corresponding to the target question is a subjective knowledge quiz question, the structure of the prompt words can be composed of the target question, the response result of the target model to this question, the reference answer of the target question, and the scoring rules. For example, the prompt words can be: {Please score the response result according to the differences between the following response result and the reference answer; Question = List three renewable energy sources, Response result = Solar energy, wind energy, geothermal energy, Reference answer = Any three of solar energy, wind energy, water energy, biomass energy, ocean energy, geothermal energy, hydrogen energy, Scoring rules = 1 point can be obtained for answering any one correctly}.
[0062] Optionally, when the question type is a subjective content creation question, the structure of the prompt words can be composed of the creation question, the creation result output by the target model for this creation question, the creation requirements, and the scoring rules, etc. For example, the prompt words can be: {Please score the creation result according to the creation result and creation requirements of the following creation question; Creation question = Write an argumentative essay about renewable energy, Creation result = xxx, Creation requirements = Meet the requirements of the question, reliable citation of facts, logical rigor, complete structure, fluent language, conform to the style of an argumentative essay, not less than 1000 words; Scoring rules = Meet the requirements of the question 0–20 points, reliable citation of facts 0 - 20, logical rigor 0 - 20, complete structure 0 - 10, fluent language 0 - 10, conform to the style of an argumentative essay 0 - 10, not less than 1000 words 0 - 10}.
[0063] After inputting the above prompt words corresponding to knowledge quiz and content creation respectively into the second evaluation model, the second evaluation model can execute different scoring logics according to the instructions and rules included in the prompt words. That is to say, different scoring logics with different focuses can be executed on different sub-categories under the same question type, making the evaluation results more reliable.
[0064] In some alternative embodiments, the second evaluation model may provide one or more expert roles, and different expert roles have different focuses when performing tasks. The evaluation tool may also add information about the expert roles adapted to the question type to the prompt words, so that the second evaluation model can call the sub-models corresponding to different expert roles according to the prompt words to perform the evaluation task. Optionally, the sub-models corresponding to different expert roles are respectively good at evaluating different subjective questions. For example, the knowledge Q&A expert in the second evaluation model is good at evaluating knowledge Q&A questions, and the writing expert is good at evaluating content creation questions. Optionally, multiple expert roles may also correspond to multiple different field types. For example, the industrial knowledge Q&A expert in the second evaluation model is good at evaluating knowledge Q&A questions in the industrial field, the computer knowledge Q&A expert is good at evaluating knowledge Q&A questions in the field of computer technology, the industrial writing expert is good at evaluating content creation questions in the industrial field, and the computer writing expert is good at evaluating content creation questions in the field of computer technology.
[0065] Based on this implementation method, by constructing prompt words adapted to different question types, the second evaluation model can be prompted to have different focuses when performing evaluation tasks for different subjective questions, so as to facilitate the second evaluation model to accurately understand the evaluation instructions and give accurate and reliable evaluation results.
[0066] The above-mentioned implementation methods can be combined and executed. For the evaluation tool, when it is determined that the corresponding question type of the target question is an objective question, the lightweight first evaluation model is called to evaluate the answer result corresponding to the target question. When it is determined that the question type is a subjective question, the second evaluation model is called to evaluate the answer result corresponding to the target question, realizing a "one large and one small" collaborative evaluation scheme. In this collaborative evaluation scheme, the first evaluation model has good performance in the evaluation task of objective questions with lower complexity, and the parameter order of the first evaluation model is relatively low, which can effectively reduce resource consumption. The second evaluation model has good performance in the evaluation task of subjective questions with higher complexity, effectively improving the evaluation accuracy of subjective questions, thus achieving a balance among accuracy, cost, and time efficiency to a certain extent.
[0067] It should be noted that in some alternative embodiments, before the evaluation tool performs the above step 103, it may further obtain the matching degree between the response result of the target question and the target question. Among them, the matching degree between the question and its response result can be calculated through the text similarity between the question and its response result. If the matching degree is less than or equal to the set matching degree threshold, it can be considered that the target model fails to accurately understand the question or there is a model hallucination in the process of the target model solving the problem. In this case, the evaluation tool can directly output a lower evaluation score to quickly obtain the evaluation result. If the matching degree is greater than the set matching degree threshold, then step 103 is executed, and the target evaluation model adapted to the question type in the evaluation model is called to evaluate the response result. Based on this implementation method, the evaluation tool can preliminarily filter the evaluation questions and evaluation results before calling the target evaluation model, which can reduce the unnecessary call frequency of the evaluation model and is beneficial to further reducing the resource cost consumed by running the model.
[0068] The following will combine the attached Figure 2 and the attached Figure 3 , taking the participating model for solving problems as an example, to further exemplarily illustrate the model evaluation method provided by the embodiments of the present application. Among them, the response result output by the participating model for solving problems for the target question can be described as the problem-solving result.
[0069] Figure 2 is a schematic diagram of an evaluation system provided by an exemplary embodiment of the present application. As Figure 2 shown, the evaluation system includes: a participating model, a question type classifier, and multiple evaluation models. As Figure 2 shown, during the evaluation process, the evaluation tool can take out the evaluation questions from the evaluation set. In some embodiments, the question types are marked on a part of the evaluation questions. When the question types are not marked on another part of the evaluation questions, the evaluation tool can input this part of the evaluation questions into Figure 2 the question type classifier shown to obtain the corresponding question types. Then, the evaluation tool can input the evaluation questions into the participating model to obtain the problem-solving result output by the participating model.
[0070] After the evaluation tool obtains the problem-solving result of the evaluation question by the participating model and the question type corresponding to the evaluation question, if it is determined that the evaluation question is an objective question, the evaluation tool can call Figure 2 the first evaluation model shown to evaluate the problem-solving result corresponding to the evaluation question. As Figure 2 shown, the first evaluation model can preprocess the evaluation question to convert the evaluation question into a specified standardized format. Through the preprocessing operation, the first evaluation model can convert the evaluation question into a single-choice question, a multiple-choice question, or a true-false question with a standard format. Then, the first evaluation model can call the rule engine to score the problem-solving result of the evaluation question according to the rules. If the evaluation question is a subjective question, the evaluation tool can callFigure 2 The second evaluation model shown scores the solution results corresponding to the evaluation questions. Among them, for the short-answer questions with reference answers, the evaluation tool can construct prompt words corresponding to the short-answer questions, reference answers, and scoring rules, and call the second evaluation model according to the prompt words. The second evaluation model can use a sub-model adapted to the short-answer questions to score the solution results. For the creative questions without reference answers, the evaluation tool can construct prompt words corresponding to the creative questions and scoring rules, and call the second evaluation model according to the prompt words, so that the second evaluation model uses a sub-model adapted to the creative questions to score the solution results.
[0071] In some other embodiments, the evaluation tool can also call the question type classifier to classify the evaluation questions after obtaining the solution results output by the participating model for the evaluation questions. As Figure 3 shown, after the evaluation tool obtains the evaluation questions and the solution results of the evaluation questions output by the participating model, the evaluation tool can classify the evaluation questions, and this question type classification operation can be implemented by using the question type classifier. The evaluation tool can input the evaluation questions and their answer results into different evaluation models according to the question types corresponding to the evaluation questions. As Figure 3 shown, for any evaluation model, in the case of not inputting the reference answer corresponding to the evaluation question, the evaluation model can analyze the evaluation question. Specifically, the evaluation model can analyze the entities in the evaluation question and the relationships between the entities according to the evaluation question and its context information, and generate a reference answer according to the analysis result. The evaluation model can evaluate and score according to the generated reference answer and the input solution result. Among them, when any evaluation model evaluates and scores the reference answer and the solution result, it can adopt a rule-based scoring method, for example, use regular expressions to match the reference answer and the solution result, and output a score according to the matching result. Or, the evaluation model can adopt a deep learning-based method for scoring. For example, a pre-trained sub-model and a fine-tuned sub-model can be used to score according to the difference between the reference answer and the solution result.
[0072] In this implementation manner, by distinguishing the question types of the evaluation questions and calling different evaluation models to perform evaluation according to the question types, the evaluation can be completed at a faster speed and with lower token consumption while ensuring the evaluation accuracy.
[0073] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 101 to 104 can be device A; or, the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0074] In addition, in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0076] Figure 4 Schematically shows a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application, as Figure 4 shown, the electronic device includes: a memory 401 and a processor 402.
[0077] The memory 401 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device.
[0078] The processor 402 is coupled to the memory 401 and is used to execute the computer program in the memory 401 for: in response to a model evaluation operation, obtaining a response result output by the target model for the target question; obtaining the question type corresponding to the target question; calling a target evaluation model adapted to the question type among a plurality of evaluation models to evaluate the response result, and obtaining an evaluation result of the target model on the target question; among the plurality of evaluation models, different evaluation models correspond to different question types, and the number of parameters of different evaluation models is different.
[0079] Optionally, when the processor 402 calls a target evaluation model adapted to the question type among a plurality of evaluation models to evaluate the response result, it includes: if the question type is an objective question, calling a first evaluation model to evaluate the response result corresponding to the target question, and the first evaluation model is a lightweight model with the number of parameters less than a set first threshold.
[0080] Optionally, when the processor 402 calls the first evaluation model to evaluate the response result corresponding to the target question and obtains the evaluation result of the target model on the target question, it is specifically used for: calling the first evaluation model, and evaluating the response result corresponding to the target question according to the rule engine corresponding to the objective question to obtain the evaluation result of the target model on the target question; the rule engine is used to match the response result with the reference answer of the target question, and calculate the score of the response result according to the matching result and the predefined scoring rule.
[0081] Optionally, when the processor 402 calls a target evaluation model adapted to the type of the target question to evaluate the response result, it is specifically used for: if the question type is a subjective question, calling a second evaluation model to evaluate the response result corresponding to the target question, and the second evaluation model is a large language model with a parameter quantity greater than a set second threshold.
[0082] Optionally, when the processor 402 calls the second evaluation model to evaluate the response result corresponding to the target question, it is specifically used for: generating a prompt word adapted to the question type according to the question type, the target question, and the corresponding response result; the structures and / or contents of the prompt words corresponding to different question types are different; calling the second evaluation model to evaluate the response result corresponding to the target question according to the prompt word.
[0083] Optionally, generating a prompt word adapted to the question type according to the question type, the target question, and the corresponding response result includes: if the question type is a subjective question of the knowledge Q&A type, generating a prompt word adapted to the subjective question of the knowledge Q&A type according to the target question, the response result, the reference answer of the question, and the scoring rule; if the question type is a subjective question of the content creation type, generating a prompt word for the subjective question of the content creation type according to the creation question, the creation result output by the target model for the creation question, the creation requirements, and the scoring rule.
[0084] Optionally, when the processor 402 obtains the question type corresponding to the target question, it is specifically used for: obtaining the question type corresponding to the target question according to the question type setting information of the target question; or calling a question type classifier to identify the question type corresponding to the target question, and the question type classifier is trained on a question type training set using a machine learning algorithm.
[0085] Optionally, before the processor 402 invokes the target evaluation model adapted to the question type among multiple evaluation models to evaluate the response result, it is further configured to: obtain the matching degree between the response result corresponding to the target question and the target question; if the matching degree is greater than a set matching degree threshold, then invoke the target evaluation model adapted to the question type among the multiple evaluation models to evaluate the response result.
[0086] Further, as Figure 4 shown, the electronic device further includes: a communication component 403, a power supply component 404, a display component 405, an audio component 406, and other components. Figure 4 Only some components are schematically shown, which does not mean that the electronic device only includes Figure 4 the components shown. Figure 4 Among them, the components within the dashed box are optional components, rather than mandatory components, and can be determined according to the product form of the electronic device. The electronic device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include Figure 4 the components within the dashed box; if the electronic device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 4 the components within the dashed box.
[0087] Among them, the memory 401 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0088] Among them, the communication component 403 is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as Wi-Fi (wireless network communication technology), 2G (such as Global System for Mobile Communications (GSM), etc.), 3G (such as Wideband Code Division Multiple Access (WCDMA)), 4G (such as Long Term Evolution (LTE), etc.), 4G+ (such as LTE-Advanced (LTE-A), etc.) or 5G (5th Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0089] Among them, the power supply component 404 is used to provide power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0090] The display component includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from users. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0091] An audio component that can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in a memory or sent via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0092] In this embodiment, after obtaining the response result output by the target model for the target question, the question type corresponding to the target question can be obtained, and the target evaluation model adapted to this question type among multiple evaluation models can be called to evaluate the response result. In this implementation, among multiple evaluation models, different evaluation models correspond to different question types, and the number of parameters of different evaluation models is different. Based on this, a solution for collaboratively evaluating the response results output by the target model for different types of questions using evaluation models with different numbers of parameters is implemented. It is convenient to use an evaluation model with a lower order of magnitude of parameters to evaluate the response results of some questions with lower complexity, effectively reducing resource consumption, and it is convenient to use an evaluation model with a higher number of parameters to evaluate the response results of some questions with higher complexity, improving the evaluation accuracy, thereby achieving a balance of accuracy, cost, and time efficiency to a certain extent.
[0093] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by an electronic device in the above method embodiment.
[0094] An embodiment of the present application further provides a computer program product, including: computer program / instructions, and when the computer program / instructions are executed by a processor, they can implement the steps in the method provided by the embodiment of the present application.
[0095] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.
[0096] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing device create means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or blocks.
[0097] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or blocks.
[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or blocks.
[0099] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and memory.
[0100] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0101] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (Parallel Random Access Machine, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (Digital Video Disc, DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0102] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, product or device comprising the element.
[0103] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A model evaluation method, characterized in that: include: Respond to the model evaluation operation and obtain the answer result output by the target model for the target question; Obtaining the topic type corresponding to the target topic; Call a target evaluation model adapted to the question type from multiple evaluation models, evaluate the answer result, and obtain the evaluation result of the target model on the target question; among the multiple evaluation models, different evaluation models correspond to different question types, and different evaluation models have different parameter quantities; wherein, the different parameter quantities of the different evaluation models include: the orders of magnitude of the parameters of the different evaluation models are different, and the orders of magnitude are used to describe the scale and complexity of the model; the parameters of the evaluation model refer to the adjustable variables inside the evaluation model used to learn data patterns, and the more parameters a model has, the more resources the model requires during operation.
2. The method according to claim 1, characterized in that Calling a target evaluation model adapted to the question type from among multiple evaluation models to evaluate the answer result includes: If the question type is an objective question, a first evaluation model is called to evaluate the answer result corresponding to the target question, and the first evaluation model is a lightweight model with a parameter amount less than a set first threshold.
3. The method according to claim 2, characterized in that Calling the first evaluation model to evaluate the answer result corresponding to the target question to obtain the evaluation result of the target model on the target question, including: Call the first evaluation model, and evaluate the answer result corresponding to the target question according to the rule engine corresponding to the objective question to obtain the evaluation result of the target model on the target question; the rule engine is used to match the answer result with the reference answer of the target question, and calculate the score of the answer result based on the matching result and predefined scoring rules.
4. The method according to claim 1, characterized in that: Calling a target evaluation model adapted to the type of the target question to evaluate the answer result, including: If the question type is a subjective question, a second evaluation model is called to evaluate the answer result corresponding to the target question, and the second evaluation model is a large language model with a parameter quantity greater than a set second threshold.
5. The method according to claim 4, characterized in that Calling the second evaluation model to evaluate the answer result corresponding to the target question includes: According to the question type, the target question and the corresponding answer result, a prompt word adapted to the question type is generated; the structure and / or content of the prompt word corresponding to different question types are different; The second evaluation model is called to evaluate the answer result corresponding to the target question according to the prompt word.
6. The method according to claim 5, characterized in that According to the question type, the target question and the corresponding answer result, a prompt word adapted to the question type is generated, including: If the question type is a subjective question of the knowledge quiz type, then generating prompt words adapted to the subjective question of the knowledge quiz type according to the target question, the answer result, the reference answer to the question and the scoring rule; If the question type is a subjective question of content creation, a prompt word for the subjective question of content creation is generated according to the creation question, the creation result output by the target model for the creation question, the creation requirements and the scoring rules.
7. The method according to any one of claims 1 to 6, characterized in that: Get the topic type corresponding to the target topic, including: According to the question type setting information of the target question, obtaining the question type corresponding to the target question; or, A question type classifier is called to identify the question type corresponding to the target question, wherein the question type classifier is trained on a question type training set using a machine learning algorithm.
8. The method according to any one of claims 1 to 6, characterized in that: Calling a target evaluation model adapted to the question type from among multiple evaluation models, and before evaluating the answer result, further comprising: Obtaining the matching degree between the answer result corresponding to the target question and the target question; If the matching degree is greater than a set matching degree threshold, a target evaluation model adapted to the question type among the multiple evaluation models is called to evaluate the answer result.
9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute the one or more computer instructions to: perform the steps in the method according to any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps in the method described in any one of claims 1 to 8 can be implemented.
11. A computer program product, characterized in that include: A computer program / instruction, which, when executed by a processor, can implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Model evaluation method and device, electronic equipment, storage medium and program product
CN118012730A
Question solving model evaluation method and device
CN118504639A
Online paper marking method and system based on artificial intelligence
CN118780268A
Model evaluation method and device, equipment and medium
CN118886430A
Assessment method and device of large language model system and related equipment
CN119179631A