Evaluation for large language model
By constructing a test question bank based on the target application field, and using corpus samples to construct objective questions to evaluate large language models, the problem of insufficient evaluation methods for specific fields in the existing technology is solved, and the accuracy and reliability of the evaluation is improved.
Patent Information
- Application Number
- PCT/CN2024/119485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-09-18
- Publication Date
- 2025-05-22
AI Technical Summary
The lack of effective methods for evaluating large language models applied to specific fields leads to insufficient accuracy and reliability of model evaluation.
By constructing a test question bank based on the target application field, including objective questions constructed by analyzing corpus samples in the field, these questions are used to test and evaluate large language models.
Improve the accuracy and reliability of model evaluation, help verify the effectiveness of the model and improve the reliability of related projects.
Smart Images

Figure CN2024119485_22052025_PF_FP_ABST
Abstract
Description
Evaluation of Large Language Models Technical Field
[0001] The present application relates to the field of model evaluation technology, and in particular to a method and apparatus for evaluating a large language model, a storage medium, and a computer device. Background Art
[0002] Model evaluation is a crucial step in the entire lifecycle of large language models: training, evaluation, deployment, and maintenance. It directly influences model selection and provides a basis for model iteration. Without a systematic, rational, and comprehensive evaluation solution, it's impossible to assess model performance after training, select the best model from multiple candidate models, or verify the effectiveness of model improvements. In the past, when single models solved single tasks, model evaluation was often limited to a single task or category, ultimately limiting application to a single direction. This is no longer true in the era of large models. How to rationally and automatically evaluate the capabilities of large models will determine model selection, improvement, and iteration during the development process.
[0003] Current evaluation schemes for large models are primarily based on existing real-world exam questions. For example, these questions, such as those from academic exams and judicial examinations, assess the model's general knowledge or domain-specific capabilities. However, these exam questions are less applicable to the specific service areas of individual companies. For example, judicial examination questions make it difficult to accurately evaluate large models used in food delivery platforms.
[0004] Currently, there is a lack of effective evaluation methods for large models applied to specific fields.
[0005] Summary of the Invention
[0006] In view of this, the embodiments of the present application provide a large language model evaluation method and device, storage medium, and computer equipment, which construct objective questions for specific application fields so that the questions are more suitable for model testing in the application field, thereby improving the accuracy and reliability of model evaluation, and further, helping to verify the validity of the model and improve the reliability of model-related projects.
[0007] According to one aspect of the present application, a method for evaluating a large language model is provided, the method comprising:
[0008] Based on a target application domain in which a target large language model to be evaluated is applied, obtaining a test question bank corresponding to the target large language model, wherein the test question bank includes objective questions belonging to the target application domain, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application domain;
[0009] Testing the target large language model using the test question bank to obtain output results of the target large language model for the tested objective questions;
[0010] The target large language model is evaluated based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested.
[0011] Optionally,
[0012] The objective questions include multiple-choice questions, wherein the stem of the multiple-choice question is determined based on the masked corpus sample after keyword masking of the corpus sample, the correct option is determined based on the masked keyword, and the incorrect option is determined based on a target keyword different from the masked keyword; and / or,
[0013] The objective questions include judgment questions, wherein the stem of the judgment questions is determined by keyword masking of the corpus sample based on the masked corpus sample and the core words corresponding to the masked keywords, and the core words are determined based on the masked keywords or other keywords different from the masked keywords.
[0014] Optionally, the objective questions include multiple-choice questions; based on the target application domain to which the target large language model to be evaluated is applied, before obtaining a test question bank corresponding to the target large language model, the method further includes:
[0015] Obtaining at least one corpus sample of the target application field;
[0016] For any corpus sample, mask at least one keyword in the corpus sample, determine the stem of a multiple-choice question based on the masked corpus sample, determine the correct option based on the masked keyword, obtain a target keyword different from the masked keyword, determine the incorrect option based on the target keyword, and construct a multiple-choice question including the stem of the multiple-choice question, the correct option, and the incorrect option;
[0017] The test question bank is constructed based on the multiple-choice questions corresponding to the corpus samples.
[0018] Optionally, the objective questions include multiple-choice questions of multiple test question types; obtaining multiple corpus samples of the target application field includes:
[0019] Determining a test question type for the target large language model according to a function of the target large language model;
[0020] Based on the original data of the target application field, obtaining at least one corpus data group corresponding to each test question type, wherein any corpus data group includes at least one original data;
[0021] For any test question type, semantic analysis is performed on at least one corpus data group corresponding to the test question type to determine at least one corpus sample;
[0022] Accordingly, masking at least one keyword in the corpus sample includes:
[0023] At least one keyword in the corpus sample that matches the test question type corresponding to the corpus sample is masked.
[0024] Optionally, the target large language model is used to implement at least one of the following functions: dish category prediction, dish taste prediction, store category prediction, store brand prediction, store package content prediction, and recommended store prediction;
[0025] The test question types include at least one of dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction;
[0026] The original data includes at least one of dish library data, store data and user order data.
[0027] Optionally, obtaining a target keyword different from the masked keyword includes:
[0028] Determine the number of incorrect options based on the number of correct options and the total number of options, and determine the difficulty of the target question corresponding to the corpus sample;
[0029] According to the difficulty of the target question, a target keyword having the number of incorrect options and different from the masked keyword is obtained, wherein the target keyword is obtained by at least one of the following methods:
[0030] Determining the incorrect option phrase category corresponding to the corpus sample based on the target question difficulty and the masked category corresponding to the masked keyword, and performing negative sampling on the phrase samples corresponding to the incorrect option phrase category based on the masked keyword to obtain a target keyword different from the masked keyword; wherein the target question difficulty, the masked category, and the incorrect option phrase category are positively correlated;
[0031] Determining the edit distance of the wrong options corresponding to the corpus sample according to the difficulty of the target question, and obtaining a target keyword that meets the edit distance of the wrong options and is different from the masked keyword; wherein the difficulty of the target question is negatively correlated with the edit distance of the wrong options;
[0032] The target question difficulty is negatively correlated with the target question text similarity.
[0033] Optionally, after constructing the multiple-choice question including the multiple-choice question stem, the correct options, and the incorrect options, the method further includes at least one of the following:
[0034] According to the target question difficulty of the corpus sample, the difficulty of the multiple-choice questions corresponding to the corpus sample is marked; according to the test question type of the corpus sample, the question type of the multiple-choice questions corresponding to the corpus sample is marked.
[0035] Optionally, the objective questions include judgment questions; based on the target application field to which the target large language model to be evaluated belongs, before obtaining the test question bank corresponding to the target large language model, the method further includes:
[0036] Obtaining at least one corpus sample of the target application field;
[0037] For any corpus sample, at least one keyword in the corpus sample is masked, and the masked keyword or other keywords different from the masked keyword are determined as core words. A judgment question stem is constructed based on the masked corpus sample and the core words. The answer to the judgment question is determined based on whether the core word is the masked keyword, and a judgment question containing the judgment question stem and the judgment question answer is constructed.
[0038] Optionally, determining the masked keyword or other keywords different from the masked keyword as a core word includes:
[0039] Determine the difficulty level of the target questions corresponding to the corpus sample;
[0040] When the target question is easy, the masked keyword, other keywords belonging to a different main category from the masked keyword, other keywords having an edit distance from the masked keyword greater than a first edit distance, or other keywords having a text similarity with the masked keyword greater than a first similarity, or other keywords having a text similarity with the masked keyword less than the first similarity, are obtained as the core words;
[0041] When the target question is of medium difficulty, other keywords belonging to different subcategories under the same main category as the masked keyword, other keywords having an edit distance with the masked keyword less than or equal to a first edit distance and greater than a second edit distance, or other keywords having a text similarity with the masked keyword greater than or equal to a first similarity and less than a second similarity are obtained as the core words;
[0042] When the difficulty of the target question is difficult, other keywords belonging to the same subcategory as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the second similarity are obtained as the core words.
[0043] Optionally, the objective questions include multiple-choice questions, and the tested objective questions include tested multiple-choice questions; and evaluating the target large language model based on answers corresponding to the tested objective questions and output results of the target large language model for the tested objective questions includes:
[0044] Extracting the predicted options of the target large language model for the tested multiple-choice question based on the output result, marking whether the predicted options are correct based on the correct options in the answers corresponding to the tested multiple-choice question, and evaluating the target large language model based on the marked results of the tested multiple-choice question, wherein the predicted options are extracted by at least one of the following methods:
[0045] Determining the predicted option of the target large language model for the multiple-choice question based on the predicted probabilities of different options included in the output result;
[0046] Extracting, from the output results, the predicted options of the target large language model for the multiple-choice question being tested according to a preset answer extraction rule;
[0047] The output result is fuzzy matched according to the correct options of the tested multiple-choice question to determine the predicted options of the target large language model for the tested multiple-choice question.
[0048] Optionally, testing the target large language model using the multiple-choice questions includes:
[0049] Based on the multiple-choice question stem, correct options and incorrect options corresponding to the multiple-choice question being tested, a test sentence of the multiple-choice question being tested is synthesized, and the target large language model is tested using the test sentence.
[0050] Optionally, the objective questions are marked with corresponding question types and question difficulties; and testing the target large language model using the test question bank includes:
[0051] Acquire multiple objective questions of different difficulty levels corresponding to each question type from the test question bank as the objective questions to be tested, and test the target large language model using the objective questions to be tested;
[0052] Accordingly, based on the answer corresponding to the objective question and the output result of the target large language model for the objective question, evaluating the target large language model includes:
[0053] Based on the answers corresponding to the objective questions being tested, the output results are marked as correct. According to the marked results of each objective question being tested, the test accuracy corresponding to each question type is counted, and the prediction accuracy of the target large language model is determined based on the test accuracy of each question type.
[0054] Optionally, the test question bank also includes objective questions belonging to general application fields;
[0055] Testing the target large language model using the test question bank includes:
[0056] A plurality of objective questions in the target application field are obtained from the test question bank as first objective questions to be tested, and a plurality of objective questions in the general application field are obtained as second objective questions to be tested, wherein the objective questions to be tested include the first objective questions to be tested and the second objective questions to be tested.
[0057] According to another aspect of the present application, a large language model evaluation device is provided, the device comprising:
[0058] a question bank acquisition module, configured to acquire a test question bank corresponding to a target large language model to be evaluated based on a target application domain to which the target large language model is applied, wherein the test question bank includes objective questions belonging to the target application domain, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application domain;
[0059] A model testing module, configured to test the target large language model using the test question bank to obtain output results of the target large language model for the objective questions being tested;
[0060] The model evaluation module is used to evaluate the target large language model based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested.
[0061] Optionally, the objective questions include multiple-choice questions, wherein the stem of the multiple-choice questions is determined based on the masked corpus sample after keyword masking of the corpus sample, the correct options are determined based on the masked keywords, and the incorrect options are determined based on target keywords different from the masked keywords; and / or,
[0062] The objective questions include judgment questions, wherein the stem of the judgment questions is determined by keyword masking of the corpus sample based on the masked corpus sample and the core words corresponding to the masked keywords, and the core words are determined based on the masked keywords or other keywords different from the masked keywords.
[0063] Optionally, the objective questions include multiple-choice questions; the question bank acquisition module is further configured to:
[0064] Obtaining at least one corpus sample of the target application field;
[0065] For any corpus sample, mask at least one keyword in the corpus sample, determine the stem of a multiple-choice question based on the masked corpus sample, determine the correct option based on the masked keyword, obtain a target keyword different from the masked keyword, determine the incorrect option based on the target keyword, and construct a multiple-choice question including the stem of the multiple-choice question, the correct option, and the incorrect option;
[0066] The test question bank is constructed based on the multiple-choice questions corresponding to the corpus samples.
[0067] Optionally, the objective questions include multiple-choice questions of multiple test question types; the question bank acquisition module is further used to: determine the test question type of the target large language model according to the function of the target large language model; obtain at least one corpus data group corresponding to each test question type based on the original data of the target application field, wherein any corpus data group includes at least one original data; for any test question type, perform semantic analysis on the at least one corpus data group corresponding to the test question type to determine at least one corpus sample;
[0068] The question bank acquisition module is further used to: mask at least one keyword in the corpus sample that matches the test question type corresponding to the corpus sample.
[0069] Optionally, the target large language model is used to achieve at least one of the following functions: dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction; the test question types include at least one of dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction; the original data includes at least one of dish library data, store data, and user order data.
[0070] Optionally, the question bank acquisition module is further configured to: determine the number of incorrect options based on the number of correct options and the total number of options, and determine the difficulty of the target question corresponding to the corpus sample;
[0071] According to the difficulty of the target question, a target keyword having the number of incorrect options and different from the masked keyword is obtained, wherein the target keyword is obtained by at least one of the following methods:
[0072] According to the target question difficulty and the masked category corresponding to the masked keyword, the incorrect option phrase category corresponding to the corpus sample is determined, and negative sampling is performed on the phrase sample corresponding to the incorrect option phrase category based on the masked keyword to obtain a target keyword different from the masked keyword; wherein the target question difficulty, the correlation between the masked category and the incorrect option phrase category is positively correlated; the preset question difficulty includes at least one of simple question difficulty, medium question difficulty and difficult question difficulty, the incorrect option phrase category corresponding to the difficult question difficulty is the same as the target question type, the incorrect option phrase category corresponding to the medium question difficulty and the target question type are different subcategories under the same main category, and the incorrect option phrase category corresponding to the simple question difficulty and the target question type belong to different main categories;
[0073] Determining the edit distance of the wrong options corresponding to the corpus sample according to the difficulty of the target question, and obtaining a target keyword that meets the edit distance of the wrong options and is different from the masked keyword; wherein the difficulty of the target question is negatively correlated with the edit distance of the wrong options;
[0074] The target question difficulty is negatively correlated with the target question text similarity.
[0075] Optionally, the question bank acquisition module is further used to: mark the difficulty of the multiple-choice questions corresponding to the corpus sample based on the target question difficulty of the corpus sample; and / or mark the question type of the multiple-choice questions corresponding to the corpus sample based on the test question type of the corpus sample.
[0076] Optionally, the objective questions include judgment questions; the question bank acquisition module is also used to: obtain at least one corpus sample from the target application field; for any corpus sample, mask at least one keyword in the corpus sample, determine the masked keyword or other keywords different from the masked keyword as the core word, construct the judgment question stem based on the masked corpus sample and the core word, determine the judgment question answer based on whether the core word is the masked keyword, and construct a judgment question containing the judgment question stem and the judgment question answer, wherein when the core word is the masked keyword, the answer to the judgment question is yes, and when the core word is not the masked keyword, the answer to the judgment question is no.
[0077] Optionally, the question bank acquisition module is further used to:
[0078] Determine the difficulty level of the target questions corresponding to the corpus sample;
[0079] When the target question is easy, the masked keyword, other keywords belonging to a different main category from the masked keyword, other keywords having an edit distance from the masked keyword greater than a first edit distance, or other keywords having a text similarity with the masked keyword greater than a first similarity, or other keywords having a text similarity with the masked keyword less than the first similarity, are obtained as the core words;
[0080] When the target question is of medium difficulty, other keywords belonging to different subcategories under the same main category as the masked keyword, other keywords having an edit distance with the masked keyword less than or equal to a first edit distance and greater than a second edit distance, or other keywords having a text similarity with the masked keyword greater than or equal to a first similarity and less than a second similarity are obtained as the core words;
[0081] When the difficulty of the target question is difficult, other keywords belonging to the same subcategory as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the second similarity are obtained as the core words.
[0082] Optionally, the objective questions include multiple-choice questions, and the tested objective questions include tested multiple-choice questions; the model evaluation module is further used to:
[0083] Extracting the predicted options of the target large language model for the tested multiple-choice question based on the output result, marking whether the predicted options are correct based on the correct options in the answers corresponding to the tested multiple-choice question, and evaluating the target large language model based on the marked results of the tested multiple-choice question, wherein the predicted options are extracted by at least one of the following methods:
[0084] Determining the predicted option of the target large language model for the multiple-choice question based on the predicted probabilities of different options included in the output result;
[0085] Extracting, from the output results, the predicted options of the target large language model for the multiple-choice question being tested according to a preset answer extraction rule;
[0086] The output result is fuzzy matched according to the correct options of the tested multiple-choice question to determine the predicted options of the target large language model for the tested multiple-choice question.
[0087] Optionally, the model testing module is further used to: synthesize a test sentence of the multiple-choice question being tested based on the multiple-choice question stem, correct options and incorrect options corresponding to the multiple-choice question being tested, and use the test sentence to test the target large language model.
[0088] Optionally, the model testing module is further configured to: obtain a plurality of objective questions of different difficulty levels corresponding to each question type from the test question bank as the objective questions to be tested, and test the target large language model using the objective questions to be tested;
[0089] Correspondingly, the model evaluation module is also used to: mark whether the output result is correct based on the answer corresponding to the objective question being tested, count the test accuracy corresponding to each question type according to the marked results of each objective question being tested, and determine the prediction accuracy of the target large language model based on the test accuracy of each question type.
[0090] Optionally, the test question bank also includes objective questions belonging to general application fields;
[0091] The model testing module is also used to: obtain from the test question bank a plurality of objective questions of different difficulty levels corresponding to each question type in the target application field as first objective questions to be tested, and obtain a plurality of objective questions in the general application field as second objective questions to be tested, wherein the objective questions to be tested include the first objective questions to be tested and the second objective questions to be tested.
[0092] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the program implements the above-mentioned large language model evaluation method.
[0093] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned large language model evaluation method when executing the program.
[0094] By means of the above technical solution, the embodiment of the present application provides a method and device for evaluating a large language model, a storage medium, and a computer device, which pre-constructs a test question bank based on the target application field to which the target large language model is applied, wherein the test question bank contains objective questions constructed by analyzing corpus samples in the field, thereby testing the target large language model using objective questions belonging to the same application field as the model, so as to perform model evaluation. The embodiment of the present application pre-constructs objective questions using corpus samples belonging to the same application field as the target large language model, thereby evaluating the target large language model through objective questions. Compared with the prior art method of using subject examinations, judicial examinations and other examination questions to evaluate the model, the problem that the test questions are not suitable for the model being tested is solved. The objective questions constructed for a specific application field in the embodiment of the present application are more suitable for model testing in that application field, thereby improving the accuracy and reliability of the model evaluation, and facilitating the verification of the model validity and improving the reliability of model-related projects.
[0095] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, which can be implemented in accordance with the contents of this application, and to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0097] FIG1 shows a flow chart of a method for evaluating a large language model provided in an embodiment of the present application.
[0098] FIG2 shows a flow chart of another large language model evaluation method provided in an embodiment of the present application.
[0099] FIG3 shows a schematic structural diagram of a large language model evaluation device provided in an embodiment of the present application.
[0100] FIG4 shows a schematic diagram of the device structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0101] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0102] In this embodiment, a method for evaluating a large language model is provided. As shown in FIG1 , the method includes steps 101 to 103 .
[0103] Step 101: Based on the target application domain in which the target large language model to be evaluated is applied, a test question bank corresponding to the target large language model is obtained, wherein the test question bank contains objective questions belonging to the target application domain, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application domain.
[0104] In an embodiment of the present application, for the target large language model to be evaluated, a test question bank corresponding to the model is obtained according to the target application field to which the model is applied. The test question bank contains at least objective questions belonging to the target application field, and may also contain objective questions from other application fields or general fields, etc., wherein the objective questions of the target application field are constructed in advance by analyzing the corpus samples of the field. For example, the large model applied to the medical service platform has the corresponding corpus samples of the medical service field; for another example, the large model applied to the local life service platform has the corresponding corpus samples of the local life service field.
[0105] Specifically, objective questions can include multiple-choice questions, judgment questions, and other question types. Multiple-choice questions can specifically include single-choice questions, multiple-choice questions, open-ended multiple-choice questions, etc. Among them, in the model learning stage, by analyzing the corpus of the target application field, the corpus used for large language model learning can be obtained. After continuous learning, the model can implement corresponding functions based on the learned knowledge. In addition, in the question bank construction stage, by analyzing the corpus of the target application field, the correct information in the field can be obtained, so that objective questions can be constructed using the correct information. Taking multiple-choice questions as an example, multiple-choice questions include the question stem and the answers to the various options. The question stem and the correct answer to the option can be directly extracted from the correct data, while the wrong answer to the option can be obtained by obtaining content different from the correct answer to the option in the corpus.
[0106] In an embodiment of the present application, optionally, the objective questions include multiple-choice questions, wherein the stems of the multiple-choice questions are determined by keyword-masking the corpus sample based on the masked corpus sample, the correct options are determined based on the masked keywords, and the incorrect options are determined based on target keywords different from the masked keywords. In one example, generating the stems of the multiple-choice questions by keyword masking refers to replacing the keywords with brackets or underscores; and / or, the objective questions include true-or-false questions, wherein the stems of the true-or-false questions are determined by keyword-masking the corpus sample based on the masked corpus sample and the core words corresponding to the masked keywords, and the core words are determined based on the masked keywords or other keywords different from the masked keywords. In one example, generating the stems of the true-or-false questions by keyword masking refers to retaining the original corpus sample or replacing the keywords with core words.
[0107] In this embodiment, for objective questions of the multiple-choice type, keyword masking can be performed on the corpus samples of the target application field, and the masked corpus samples can be used to refine the multiple-choice question stems, and the masked keywords can be used to determine the content of the correct options. As for the content of the incorrect options, the masked keywords can be negatively sampled to obtain target keywords different from the masked keywords, and then determined using the target keywords. For objective questions of the judgment question type, based on the characteristics of the judgment question type, it can be seen that the judgment question stem is a question centered around a core word, and the answer can only be correct or wrong. Therefore, the corpus sample of the target application field can also be keyword masked first, and the masked corpus sample can be used to determine the part of the question in the judgment question stem other than the core word. The core word can be the masked keyword, or other keywords different from the masked keyword obtained by negative sampling of the masked keyword, so as to complete the question with the missing core word determined by using the masked corpus sample and obtain a complete judgment question stem. Among them, if the masked keyword is directly selected as the core word, then the answer to the judgment question is correct. If other keywords are selected as the core word, then the answer to the judgment question is wrong.
[0108] Optionally, the objective questions include multiple-choice questions; before step 101, it also includes: obtaining at least one corpus sample of the target application field; for any corpus sample, masking at least one keyword in the corpus sample, determining the stem of the multiple-choice question based on the masked corpus sample, determining the correct option based on the masked keyword, obtaining a target keyword different from the masked keyword, determining the incorrect option based on the target keyword, and constructing a multiple-choice question including the stem of the multiple-choice question, the correct option and the incorrect option; constructing the test question bank based on the multiple-choice questions corresponding to the corpus sample.
[0109] In this embodiment, corpus samples from the target application domain are obtained, and at least one multiple-choice question is generated from each corpus sample, so that a test question bank is constructed using the generated multiple-choice questions. When generating a multiple-choice question using any corpus sample, one or more keywords in the corpus sample are first masked, transforming the corpus sample from a complete corpus sample into a corpus sample with missing keywords. The masked missing corpus sample is then used to generate the stem of the multiple-choice question, and the correct answers are generated using the masked keywords. Finally, several target keywords different from the masked keywords are obtained to construct the incorrect answers for the multiple-choice question.
[0110] For example, consider the complete corpus sample "A user ordered [xx burger + Orleans chicken wings + medium cola] three times this week at ** Burger Shop." First, mask the "medium cola." Based on the masked corpus, the multiple-choice question stem is determined as follows: "At lunchtime, Xiao Ming ordered [xx burger*1, Orleans chicken wings*3] and a certain item at Wallace. Which of the following dishes is most likely the user ordered?" Based on the masked "medium cola," the correct answer is determined to be "medium cola." Next, target keywords distinct from "medium cola" are obtained, including "ice porridge," "porridge cake," and "soy milk." Further, incorrect answers corresponding to the target keywords are determined to include "Xiahe mung bean ice porridge," "porridge cake," and "fragrant soy milk." Finally, the sequence number is added to each answer to form the multiple-choice question options and the correct answer: A: Xiahe mung bean ice porridge; B: porridge cake; C: fragrant soy milk; D: medium cola. Answer: D. Medium cola. Finally, a complete multiple-choice question is formed.
[0111] Step 102: Test the target large language model using the test question bank to obtain output results of the target large language model for the tested objective questions.
[0112] In this embodiment of the present application, the target large language model is tested using objective questions in the test question bank to obtain the output of the target large language model for each objective question being tested. Taking the aforementioned multiple-choice questions as an example, the output of the target large language model may take various forms, such as directly outputting the predicted correct option, or outputting the predicted probability of each option, etc.
[0113] Step 103 : Evaluate the target large language model based on the answer corresponding to the objective question and the output result of the target large language model for the objective question.
[0114] In the embodiment of the present application, the output results of the target large language model for each objective question and the answers to the objective questions can be used to determine whether the prediction results of the target large language model for the objective questions are correct, thereby performing model evaluation.
[0115] By applying the technical solution of this embodiment, a test question bank is constructed in advance based on the target application field to which the target large language model is applied, wherein the test question bank contains objective questions constructed by analyzing corpus samples in the field, thereby testing the target large language model using objective questions belonging to the same application field as the model, so as to perform model evaluation. The embodiment of the present application constructs objective questions in advance using corpus samples belonging to the same application field as the target large language model, thereby evaluating the target large language model through objective questions. Compared with the prior art method of using subject examinations, judicial examinations and other examination questions to evaluate the model, the problem that the test questions are not suitable for the tested model is solved. The objective questions constructed for a specific application field in the embodiment of the present application are more suitable for model testing in that application field, thereby improving the accuracy and reliability of the model evaluation, and further, helping to verify the validity of the model and improve the reliability of model-related projects.
[0116] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, another large language model evaluation method is provided, as shown in Figure 2, which includes steps 201 to 208.
[0117] Step 201 : Determine the test question type of the target large language model according to the function of the target large language model.
[0118] In an embodiment of the present application, test question types can be determined based on the functions that the target large language model can implement, thereby generating objective questions of the corresponding test question type. Optionally, the target large language model is configured to implement at least one of the following functions: dish category prediction, dish flavor prediction, store category prediction, store brand prediction, store package content prediction, and store recommendation prediction; the test question types include at least one of dish category prediction, dish flavor prediction, store category prediction, store brand prediction, store package content prediction, and store recommendation prediction; store category prediction includes prediction of the store's main category and prediction of the store's auxiliary category. For example, in this application, categories can include staple foods, snacks, and dishes, and within each category, over 300 subcategories are hierarchically formed. For example, the category for "Kung Pao Chicken" is "Dishes." Category labels are the basic categorization information for gourmet food products, and the basic attributes of gourmet food products vary depending on the category. For example, the "Dishes" category has "Meat and Vegetable" and "Cuisine" categories, while the "Alcoholic Beverages" category does not have such attribute labels. This is just an example of a “category” given for ease of understanding and is not restrictive.
[0119] Step 202: Based on the original data of the target application field, at least one corpus data group corresponding to each test question type is obtained, wherein any corpus data group includes at least one piece of original data.
[0120] In an embodiment of the present application, for any test question type, at least one corpus data group can be extracted from the original data in the target application field, and a corpus data group contains one or more original data. The original data in a corpus data group can fully express an objective fact related to the corresponding test question type. For example, if the test question type is dish category prediction, the original data in the corpus data group may include the user's order data in a specific store (or chain store) and the dish library of the store (or chain store). The user order data can reflect the dish category ordered by the user in the store (or chain store) so as to predict the user's interested dish category in the store (or chain store). Optionally, the original data includes at least one of dish library data, store data and user order data. It should be noted that the dish library data and store data used in the embodiment of the present application are data obtained after authorization by the merchant, and the user order data is data obtained after authorization by the user.
[0121] Step 203: for any test question type, semantic analysis is performed on at least one corpus data group corresponding to the test question type to determine at least one corpus sample.
[0122] In the embodiment of the present application, after determining the corpus data group under each test question type, semantic analysis can be performed on the original data in each group to obtain the corresponding corpus sample. The corpus sample is a digital information of an objective fact related to the corresponding test question type.
[0123] Step 204 , constructing a multiple-choice question for any corpus sample, includes steps 204 - 1 to 204 - 3 .
[0124] Step 204-1 is to mask at least one keyword in the corpus sample that matches the test question type corresponding to the corpus sample, determine the multiple-choice question stem based on the masked corpus sample, and determine the correct option based on the masked keyword.
[0125] In an embodiment of the present application, one or more multiple-choice questions can be constructed using a corpus sample, and objective questions under different test question types can be constructed using the same corpus sample. First, according to the test question type corresponding to the corpus sample, at least one keyword in the corpus sample that matches the test question type is masked. For example, if the test question type is dish category prediction, the masked keyword is a keyword related to the dish category. Then, the masked corpus sample is used to generate the stem of the multiple-choice question, and the masked keyword is used to generate the correct options for the multiple-choice question, wherein the multiple-choice question can be a single-choice question or a multiple-choice question, and each correct option can be constructed based on one keyword or multiple keywords. For example, if the masked keywords are "medium cup of cola" and "Orleans chicken wings", two correct options for the multiple-choice question can be generated, namely "medium cup of cola" and "Orleans chicken wings", and one correct option for the single-choice question can be generated, namely "medium cup of cola, Orleans chicken wings".
[0126] It should be noted that the multiple-choice questions included in the test question bank in the embodiment of the present application can be single-choice questions, multiple-choice questions, or a combination of the two. When constructing a single-choice question, a correct option is generated based on the masked keyword; when constructing a multiple-choice question, the number of correct options can be randomly or specified first, and then the keyword masking of the corpus sample is performed according to the number of correct options (and the number of keywords corresponding to each correct option), thereby constructing a corresponding number of correct options; when constructing a multiple-choice question, the number of keywords that match the corresponding test question type in the corpus sample can also be identified first, and then the number of correct options (and the number of keywords corresponding to each correct option) is determined according to the number of keywords to perform keyword masking of the corpus sample, thereby constructing a corresponding number of correct options.
[0127] Step 204-2: Determine the number of incorrect options based on the number of correct options and the total number of options, and determine the difficulty level of the target question corresponding to the corpus sample.
[0128] In an embodiment of the present application, the number of incorrect options is calculated based on the number of correct options and the total number of preset options, and the difficulty of the target question corresponding to the corpus sample is obtained, so as to obtain a corresponding number of target keywords different from the masked keywords according to the difficulty of the target question, and construct a corresponding number of incorrect options based on the target keywords.
[0129] Step 204-3: acquiring target keywords that are different from the masked keywords and have the same number of incorrect options according to the difficulty of the target question, wherein the target keywords are acquired by at least one of the following methods:
[0130] Determining the incorrect option phrase category corresponding to the corpus sample based on the target question difficulty and the masked category corresponding to the masked keyword, and performing negative sampling on the phrase samples corresponding to the incorrect option phrase category based on the masked keyword to obtain a target keyword different from the masked keyword; wherein the target question difficulty, the masked category, and the incorrect option phrase category are positively correlated;
[0131] Determining the edit distance of the wrong options corresponding to the corpus sample according to the difficulty of the target question, and obtaining a target keyword that meets the edit distance of the wrong options and is different from the masked keyword; wherein the difficulty of the target question is negatively correlated with the edit distance of the wrong options;
[0132] The target question difficulty is positively correlated with the target question difficulty. The target question difficulty is positively correlated with the target question difficulty.
[0133] Step 204-4: determining incorrect options based on the target keyword.
[0134] Step 204-5: construct a multiple-choice question including the multiple-choice question stem, the correct options, and the incorrect options.
[0135] The present application examples list three methods for obtaining target keywords for generating incorrect options. Other methods can also be used to obtain target keywords, which are not limited here. After determining the target keywords, the incorrect options can be constructed, and a complete multiple-choice question containing the question stem, correct options, and incorrect options can be generated.
[0136] Method 1: Obtain target keywords by performing random negative sampling on masked keywords. Specifically, first determine the incorrect option phrase category based on the difficulty of the target question and the masked category corresponding to the masked keyword, so as to obtain target keywords different from the masked keyword through random negative sampling in the phrase samples under the incorrect option phrase category. For example, the phrase samples are the product names listed by merchants under the category in actual scenarios, such as cola, fragrant soy milk, etc. listed under the beverage category. Among them, the greater the difficulty of the target question, the greater the correlation between the corresponding incorrect option phrase category and the masked category, that is, the greater the difficulty of the target question, the more similar the target keyword is to the masked keyword. In a specific application scenario, the preset question difficulty includes at least one of the following: easy question difficulty, medium question difficulty, and difficult question difficulty. The wrong option phrase category corresponding to the difficult question difficulty belongs to the same subcategory as the masked category. The wrong option phrase category corresponding to the medium question difficulty and the masked category are different subcategories under the same main category. The wrong option phrase category corresponding to the easy question difficulty and the masked category belong to different main categories. For example, in the problem of predicting the content of a set meal, if the positive sample is "fried tofu", the options sampled under the difficult difficulty level may be [Japanese tofu, thousand-layer tofu, fried bean skin], while the options sampled under the easy difficulty level may be [cola, Korean rice cake, Chelsea chicken roll]. The difficulty between the two is obviously different. In some embodiments of the present application, the masked category refers to the category to which the masked keyword belongs, and the wrong option phrase category refers to the category to which the wrong option phrase belongs.
[0137] Method 2: Obtaining target keywords based on edit distance. Specifically, edit distance ranges can be set for different question difficulty levels, allowing target keywords of varying difficulty levels to be sampled based on these ranges. In practical applications, completely random negative sampling can easily sample options that are also correct for the current question, which is inappropriate for multiple-choice questions. Therefore, random negative sampling can be used to construct a candidate set of keywords that exceeds the number of incorrect options. By calculating the edit distance between the correct options and the candidate set, the probability of invalid selections can be reduced. Taking the set meal prediction problem as an example, after calculating the edit distance between dish names, a sampling interval can be set to control the similarity of sampled items in the same category. For example, if the positive example is "Teriyaki Chicken Drumstick Burger," with an edit distance between 2 and 4, the negative examples might be [Spicy Chicken Drumstick Burger, McChicken Burger, Spicy Chicken Wrap], etc.—similar but not identical, making the question more difficult. Specifically, edit distance is the edit distance between strings, also known as Levenshtein distance, which refers to the minimum number of operations required to transform string A into string B using character operations, including deletion, insertion, and modification. The wrong option edit distance is the edit distance between the masked keyword and the wrong option.
[0138] Method three: Obtain target keywords based on text similarity. Specifically, the text similarity range corresponding to different question difficulties can be set, so as to sample target keywords of different difficulty levels based on different text similarity ranges. In actual application scenarios, there are many ways to calculate the text similarity between two words. Relevant options can be recalled through text vector matching algorithms such as StructBERT and CoROM. In addition, since the number of stores or dish names may be very large, the computational cost required to complete random extraction at one time is high. A bucketing strategy can also be adopted to randomly generate a number to determine a bucket, and then extract negative samples in the bucket through similarity matching. In some embodiments of the present application, the text similarity of the wrong option is the text similarity between the masked keyword and the wrong option, and the difficulty of the target question is positively correlated with the text similarity of the wrong option. That is to say, the higher the text similarity between the masked keyword and the wrong option, the more difficult the question.
[0139] Step 205 : Mark the difficulty of the multiple-choice questions corresponding to the corpus sample according to the target question difficulty of the corpus sample; and / or mark the question type of the multiple-choice questions corresponding to the corpus sample according to the test question type of the corpus sample.
[0140] Step 206: construct the test question bank based on the multiple-choice questions corresponding to the corpus sample.
[0141] The various methods of constructing error options proposed in the embodiments of the present application can increase the difficulty of the problem or improve the passing rate of the problem annotation, thereby increasing the number of samples, verifying the generalization ability of the model, and preventing the distribution of test samples from being unbalanced, resulting in model overfitting.
[0142] Step 207: synthesize a test sentence for the multiple-choice question based on the stem, correct options, and incorrect options corresponding to the multiple-choice question, and use the test sentence to test the target large language model to obtain the output result of the target large language model for the objective question.
[0143] In the embodiment of the present application, when using multiple-choice questions to test the model, you can first use the multiple-choice question stem, correct options, and incorrect options to synthesize a test statement for asking questions, and then input the test statement into the model for testing. For example, the test statement can be:
[0144] The user asked, "What special dishes are very popular in the Tibet Autonomous Region?" Which of the following options meets the requirements? ( )
[0145] A. Sichuan pepper shrimp
[0146] B. Yellow Milk Maji Shao (Hot Skewer)
[0147] C. Shrimp and Egg Fried Rice
[0148] D. Tibetan tsampa
[0149] Step 208: Extract the predicted options of the target large language model for the multiple-choice question based on the output results, and mark whether the predicted options are correct based on the correct options of the multiple-choice question. Evaluate the target large language model based on the marked results of the multiple-choice question. The predicted options are extracted by at least one of the following methods:
[0150] Determining the predicted option of the target large language model for the multiple-choice question based on the predicted probabilities of different options included in the output result;
[0151] Extracting, from the output results, the predicted options of the target large language model for the multiple-choice question being tested according to a preset answer extraction rule;
[0152] The output result is fuzzy matched according to the correct options of the tested multiple-choice question to determine the predicted options of the target large language model for the tested multiple-choice question.
[0153] In an embodiment of the present application, in a scenario where the model is tested using the multiple-choice questions, since the output forms of the large language model may be varied, an embodiment of the present application proposes a variety of methods for extracting the model's prediction options from the output results of the model. 1. Calculate the Logits (logical value, a real number, refers to the confidence or score for the option, obtained by processing the input signal through a series of linear transformations and nonlinear activation functions, and used to represent the probability of the sample belonging to each category) corresponding to the option, and select the option with the largest Logits. The generated results of the model at each position are in the form of a probability distribution. The model is selected to generate probability scores for the four options A, B, C, and D at the first position, and the option with the highest score is selected as the final result (applicable to single-choice questions), or it is regarded as the correct answer when the probability score is greater than a certain threshold (applicable to multiple-choice questions). For example, the probability distribution generated by the model for each option is: A: [0.3, 0.1, 0.4, 0.2]; B: [0.2, 0.3, 0.2, 0.3]; C: [0.2, 0.2, 0.3, 0.3]; D: [0.4, 0.2, 0.3, 0.1], where the first 0.3 is P(A1), indicating that the probability of the model selecting A in the first position is 0.3, and the first 0.1 is P(A2), indicating that the probability of the model selecting A in the second position is 0.1. That is, at the first position, the probability scores for options A, B, C, and D are generated as P(A1) = 0.3, P(B1) = 0.2, P(C1) = 0.1, and P(D1) = 0.4. Therefore, for single-choice questions, option D, with the highest score, is selected as the final answer. For multiple-choice questions, assuming a preset threshold of 0.35, the probability of option D at the first position is greater than 0.35, and the probability of option A at the third position is 0.4, which is greater than 0.35, resulting in the final answers being D and A. 2. Rule-based extraction. The model output may not always generate the answer at the beginning. It may first analyze the incorrect options and then select the correct answer, for example, "Option B is incorrect due to.... Option C is incorrect due to.... Therefore, the final answer is: A." Therefore, rule extraction can also be performed on the model's answers. This involves using regular expressions to match phrases such as "the final answer is" or "the answer is:" to obtain the answer options. This method can be applied to both single-choice and multiple-choice questions. 3. Fuzzy matching methods. When outputting results, the model may not necessarily generate options A, B, C, or D. It may directly generate the content of the options and ignore the letters. For example, in the sentence "A chicken burger is not a drink, so the answer is Coke / You should choose Coke," the model actually generates the correct answer but does not generate the relevant options. Fuzzy matching is performed on the generated text using regular expressions. If the matching criteria are met, the model's choice is considered correct. This approach can be applied to both single-choice and multiple-choice questions.The embodiment of the present application can use the above methods one by one until the answer is extracted, or can use the above three methods in combination to extract the answer, and weight the answers obtained by each method to determine the final answer.
[0154] In an embodiment of the present application, optionally, the objective questions are marked with corresponding question types and question difficulties; and testing the target large language model using the test question bank includes: obtaining a plurality of objective questions of different question difficulties corresponding to each question type from the test question bank as tested objective questions, and testing the target large language model using the tested objective questions to obtain output results of the target large language model for the tested objective questions;
[0155] Accordingly, based on the answers corresponding to the objective questions being tested and the output results of the target large language model for the objective questions being tested, the target large language model is evaluated, including: marking whether the output results are correct based on the answers corresponding to the objective questions being tested, counting the test accuracy corresponding to each question type according to the marked results of each objective question being tested, and determining the prediction accuracy of the target large language model based on the test accuracy of each question type.
[0156] In the above embodiment, when constructing a test question bank, the objective questions in the question bank can also be marked with question types and question difficulties, so that during the test, the objective questions to be tested can be comprehensively selected according to the marking of question types and question difficulties. Specifically, multiple objective questions of different question difficulties can be selected for testing under each question type. Further, the predicted answer of each objective question to be tested is marked as correct according to the model. When evaluating the model, a macro-average method can be used, without distinguishing the difficulty of the questions, to calculate the test accuracy of each question category, and then the model can be evaluated based on the test accuracy of each question category. The specific evaluation method is only illustrated by way of example in the embodiment of this application. Those skilled in the art can also adopt other evaluation calculation methods, which are not limited here.
[0157] In an embodiment of the present application, optionally, a plurality of objective questions of different difficulty levels corresponding to each question type are obtained in the test question bank as the objective questions to be tested, including: obtaining a plurality of objective questions of different difficulty levels corresponding to each question type under the target application field in the test question bank as the first objective questions to be tested, and obtaining a plurality of objective questions under the general application field as the second objective questions to be tested, wherein the objective questions to be tested include the first objective questions to be tested and the second objective questions to be tested.
[0158] In the above embodiment, the large language model for the target application field can not only solve problems in the target application field, but also has the ability to solve problems in other fields or general fields. When evaluating the model, questions in general fields can also be used to test the model to improve the generalization of the model.
[0159] In an embodiment of the present application, optionally, the objective questions include judgment questions; based on the target application field to which the target large language model to be evaluated belongs, before obtaining the test question bank corresponding to the target large language model, the method also includes: obtaining at least one corpus sample of the target application field; for any corpus sample, masking at least one keyword in the corpus sample, determining the masked keyword or other keywords different from the masked keyword as the core word, constructing a judgment question stem based on the masked corpus sample and the core word, determining the judgment question answer based on whether the core word is the masked keyword, and constructing a judgment question containing the judgment question stem and the judgment question answer, wherein the answer to the judgment question is yes when the core word is the masked keyword, and the answer to the judgment question is no when the core word is not the masked keyword.
[0160] In the above embodiment, the test question bank may also include judgment questions, which are also constructed using corpus samples in the target application field. Specifically, when using any corpus sample to generate a judgment question, one or more keywords in the corpus sample are first masked to change the corpus sample from the original complete corpus sample to a missing corpus sample. Then, the masked missing corpus sample can be used to generate the stem of the judgment question. Furthermore, the masked keywords can be directly used to generate the core words of the judgment question, or the masked keywords can be negatively sampled to obtain the core words, thereby generating a judgment question using the stem and the core words. If the core words are directly generated using the masked keywords, the answer to the judgment question is correct, otherwise the answer is wrong. For example, if the corpus sample is "In the Tibet Autonomous Region, Tibetan tsampa is a popular specialty dish", the generated judgment question is "Is Tibetan tsampa a popular specialty dish in the Tibet Autonomous Region?" It can also be "Is shrimp fried rice a popular specialty dish in the Tibet Autonomous Region?"
[0161] In addition, when performing keyword masking, the keyword masking can also be performed according to the test question type corresponding to the corpus sample, thereby generating judgment questions of the corresponding question type.
[0162] In an embodiment of the present application, optionally, the masked keyword or other keywords different from the masked keyword are determined as core words, including: determining the difficulty of the target topic corresponding to the corpus sample; when the difficulty of the target topic is simple, obtaining the masked keyword, other keywords belonging to a different main category from the masked keyword, other keywords with an edit distance greater than a first edit distance from the masked keyword, or other keywords with a text similarity greater than a first similarity with the masked keyword or other keywords with a text similarity less than a first similarity with the masked keyword as the core words; when the difficulty of the target topic is medium, obtaining Take other keywords belonging to different subcategories under the same main category as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the first edit distance and greater than the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the first similarity and less than the second similarity as the core words; when the difficulty of the target question is difficult, obtain other keywords belonging to the same subcategory as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the second similarity as the core words.
[0163] In this embodiment, the core words of the judgment questions can also be determined based on the difficulty of the target questions corresponding to the corpus sample. Taking the difficulty level as simple, medium and difficult as an example, in the simple difficulty level, the masked keyword can be directly used as the core word, or other keywords that do not belong to the same main category as the masked keyword can be negatively sampled, other keywords with a large edit distance (greater than the first edit distance) between the negative sampling and the masked keyword can be negatively sampled, or other keywords with a small text similarity (less than the first similarity) between the negative sampling and the masked keyword can be negatively sampled; in the medium difficulty level, other keywords that belong to different subcategories under the same main category as the masked keyword can be negatively sampled, other keywords with a medium edit distance (less than or equal to the first edit distance and greater than the second edit distance) between the negative sampling and the masked keyword can be negatively sampled, or other keywords with a medium text similarity (greater than or equal to the first similarity and less than the second similarity) between the negative sampling and the masked keyword can be negatively sampled; in the difficult difficulty level, other keywords that belong to the same subcategory as the masked keyword can be negatively sampled, other keywords with a small edit distance (less than or equal to the second edit distance) between the negative sampling and the masked keyword can be negatively sampled, or other keywords with a large text similarity (text similarity greater than or equal to the second similarity) between the negative sampling and the masked keyword can be negatively sampled.
[0164] By applying the technical solution of this embodiment, for the target large language model that needs to be evaluated, the test question type is determined according to the model function, thereby obtaining a corpus data group under each test question type from the original data generated in the target application field, and determining the corpus sample by analyzing the original data in the corpus data group. For different types of questions, negative samples are obtained based on random negative sampling, edit distance negative sampling, and text similarity negative sampling, and objective questions of different difficulty levels and forms (including single-choice questions, multiple-choice questions, and judgment questions) are constructed to form a set of objective question banks for specific fields with diverse forms and stratified difficulty levels. Furthermore, in the model evaluation stage, not only can the model be tested using pre-constructed objective questions in the target application field, but also based on objective questions in general application fields. In addition, in view of the characteristics of the large language model answering questions in a generative form, a variety of answer extraction methods based on probability calculation, regular matching, and fuzzy matching are designed to improve the success rate of answer recognition. In addition, the final answer can be obtained based on the weighting of multiple extraction results, thereby comprehensively evaluating the output results of the model and achieving unified results.
[0165] Furthermore, as a specific implementation of the method of FIG1 , an embodiment of the present application provides an evaluation device 300 for a large language model. As shown in FIG3 , the device includes modules 301 to 303 .
[0166] The question bank acquisition module 301 is used to obtain a test question bank corresponding to the target large language model to be evaluated based on the target application field to which the target large language model is applied, wherein the test question bank contains objective questions belonging to the target application field, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application field.
[0167] The model testing module 302 is configured to test the target large language model using the test question bank to obtain output results of the target large language model for the objective questions being tested.
[0168] The model evaluation module 303 is configured to evaluate the target large language model based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested.
[0169] Optionally, the objective questions include multiple-choice questions, wherein the stems of the multiple-choice questions are determined by keyword-masking the corpus sample and then based on the masked corpus sample, the correct options are determined based on the masked keywords, and the incorrect options are determined based on target keywords different from the masked keywords; and / or, the objective questions include true-or-false questions, wherein the stems of the true-or-false questions are determined by keyword-masking the corpus sample and then based on the masked corpus sample and core words corresponding to the masked keywords, and the core words are determined based on the masked keywords or other keywords different from the masked keywords.
[0170] Optionally, the objective questions include multiple-choice questions; the question bank acquisition module 301 is further configured to:
[0171] Obtaining at least one corpus sample of the target application field;
[0172] For any corpus sample, mask at least one keyword in the corpus sample, determine the stem of a multiple-choice question based on the masked corpus sample, determine the correct option based on the masked keyword, obtain a target keyword different from the masked keyword, determine the incorrect option based on the target keyword, and construct a multiple-choice question including the stem of the multiple-choice question, the correct option, and the incorrect option;
[0173] The test question bank is constructed based on the multiple-choice questions corresponding to the corpus samples.
[0174] Optionally, the objective questions include multiple-choice questions of multiple test question types; the question bank acquisition module 301 is further used to: determine the test question type of the target large language model according to the function of the target large language model; obtain at least one corpus data group corresponding to each test question type based on the original data of the target application field, wherein any corpus data group includes at least one original data; for any test question type, perform semantic analysis on the at least one corpus data group corresponding to the test question type to determine at least one corpus sample.
[0175] The question bank acquisition module 301 is further configured to mask at least one keyword in the corpus sample that matches the test question type corresponding to the corpus sample.
[0176] Optionally, the target large language model is used to achieve at least one of the following functions: dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction; the test question types include at least one of dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction; the original data includes at least one of dish library data, store data, and user order data.
[0177] Optionally, the question bank acquisition module 301 is further configured to:
[0178] Determine the number of incorrect options based on the number of correct options and the total number of options, and determine the difficulty of the target question corresponding to the corpus sample;
[0179] According to the difficulty of the target question, a target keyword having the number of incorrect options and different from the masked keyword is obtained, wherein the target keyword is obtained by at least one of the following methods:
[0180] According to the target question difficulty and the masked category corresponding to the masked keyword, the incorrect option phrase category corresponding to the corpus sample is determined, and negative sampling is performed on the phrase sample corresponding to the incorrect option phrase category based on the masked keyword to obtain a target keyword different from the masked keyword; wherein the target question difficulty, the correlation between the masked category and the incorrect option phrase category is positively correlated; the preset question difficulty includes at least one of simple question difficulty, medium question difficulty and difficult question difficulty, the incorrect option phrase category corresponding to the difficult question difficulty is the same as the target question type, the incorrect option phrase category corresponding to the medium question difficulty and the target question type are different subcategories under the same main category, and the incorrect option phrase category corresponding to the simple question difficulty and the target question type belong to different main categories;
[0181] Determining the edit distance of the wrong options corresponding to the corpus sample according to the difficulty of the target question, and obtaining a target keyword that meets the edit distance of the wrong options and is different from the masked keyword; wherein the difficulty of the target question is negatively correlated with the edit distance of the wrong options;
[0182] The target question difficulty is negatively correlated with the target question text similarity.
[0183] Optionally, the question bank acquisition module 301 is further used to: mark the difficulty of the multiple-choice questions corresponding to the corpus sample according to the target question difficulty of the corpus sample; and / or mark the question type of the multiple-choice questions corresponding to the corpus sample according to the test question type of the corpus sample.
[0184] Optionally, the objective questions include judgment questions; the question bank acquisition module 301 is also used to: obtain at least one corpus sample in the target application field; for any corpus sample, mask at least one keyword in the corpus sample, determine the masked keyword or other keywords different from the masked keyword as the core word, construct the judgment question stem based on the masked corpus sample and the core word, determine the judgment question answer based on whether the core word is the masked keyword, and construct a judgment question containing the judgment question stem and the judgment question answer, wherein when the core word is the masked keyword, the answer to the judgment question is yes, and when the core word is not the masked keyword, the answer to the judgment question is no.
[0185] Optionally, the question bank acquisition module 301 is further configured to:
[0186] Determine the difficulty level of the target questions corresponding to the corpus sample;
[0187] When the target question is easy, the masked keyword, other keywords belonging to a different main category from the masked keyword, other keywords having an edit distance from the masked keyword greater than a first edit distance, or other keywords having a text similarity with the masked keyword greater than a first similarity, or other keywords having a text similarity with the masked keyword less than the first similarity, are obtained as the core words;
[0188] When the target question is of medium difficulty, other keywords belonging to different subcategories under the same main category as the masked keyword, other keywords having an edit distance with the masked keyword less than or equal to a first edit distance and greater than a second edit distance, or other keywords having a text similarity with the masked keyword greater than or equal to a first similarity and less than a second similarity are obtained as the core words;
[0189] When the difficulty of the target question is difficult, other keywords belonging to the same subcategory as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the second similarity are obtained as the core words.
[0190] Optionally, the objective questions include multiple-choice questions, and the tested objective questions include tested multiple-choice questions; the model evaluation module 303 is further configured to:
[0191] Extracting the predicted options of the target large language model for the tested multiple-choice question based on the output result, marking whether the predicted options are correct based on the correct options in the answers corresponding to the tested multiple-choice question, and evaluating the target large language model based on the marked results of the tested multiple-choice question, wherein the predicted options are extracted by at least one of the following methods:
[0192] Determining the predicted option of the target large language model for the multiple-choice question based on the predicted probabilities of different options included in the output result;
[0193] Extracting, from the output results, the predicted options of the target large language model for the multiple-choice question being tested according to a preset answer extraction rule;
[0194] The output result is fuzzy matched according to the correct options of the tested multiple-choice question to determine the predicted options of the target large language model for the tested multiple-choice question.
[0195] Optionally, the model testing module is further used to: synthesize a test sentence of the multiple-choice question being tested based on the multiple-choice question stem, correct options and incorrect options corresponding to the multiple-choice question being tested, and use the test sentence to test the target large language model.
[0196] Optionally, the model testing module 302 is further configured to: obtain a plurality of objective questions of different difficulty levels corresponding to each question type from the test question bank as the objective questions to be tested, and test the target large language model using the objective questions to be tested;
[0197] Correspondingly, the model evaluation module 303 is also used to: mark whether the output result is correct based on the answer to the objective question being tested, count the test accuracy corresponding to each question type according to the marked results of each objective question being tested, and determine the prediction accuracy of the target large language model based on the test accuracy of each question type.
[0198] Optionally, the test question bank also includes objective questions belonging to a general application field; the model testing module 302 is also used to: obtain from the test question bank a plurality of objective questions of different difficulty levels corresponding to each question type under the target application field as the first objective questions to be tested, and obtain a plurality of objective questions under the general application field as the second objective questions to be tested, wherein the objective questions to be tested include the first objective questions to be tested and the second objective questions to be tested.
[0199] It should be noted that for other corresponding descriptions of the functional units involved in the large language model evaluation device provided in the embodiment of the present application, reference can be made to the corresponding descriptions in the methods of Figures 1 to 2, and no further details will be given here.
[0200] The embodiment of the present application also provides a computer device, which can be specifically a personal computer, a server, a network device, etc. As shown in Figure 4, the computer device includes a bus, a processor, a memory and a communication interface, and may also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.
[0201] Those skilled in the art will understand that the structure shown in FIG4 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0202] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0203] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0204] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0205] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like.
[0206] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0207] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for evaluating a large language model, characterized in that: The method comprises: Based on the target application field to which the target large language model to be evaluated is applied, obtaining a test question bank corresponding to the target large language model, wherein the test question bank contains objective questions belonging to the target application field, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application field; Using the test question bank to test the target large language model, and obtaining an output result of the target large language model for the tested objective question; The target large language model is evaluated based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested.
2. The method according to claim 1, characterized in that The objective questions include multiple-choice questions, wherein the stem of the multiple-choice questions is determined based on the masked corpus sample after keyword masking of the corpus sample, the correct option is determined based on the masked keyword, and the wrong option is determined based on a target keyword different from the masked keyword; and / or, The objective questions include judgment questions, wherein the stem of the judgment questions is determined by keyword masking of the corpus sample, based on the masked corpus sample and the core words corresponding to the masked keywords, and the core words are determined based on the masked keywords or other keywords different from the masked keywords.
3. The method according to claim 1, characterized in that The objective questions include multiple-choice questions; before obtaining the test question bank corresponding to the target large language model based on the target application field to which the target large language model to be evaluated is applied, the method further includes: Acquire at least one corpus sample of the target application field; For any corpus sample, Mask at least one keyword in the corpus sample. Determine the stem of the multiple-choice questions based on the masked corpus samples. Determine the correct option based on the masked keyword. obtaining a target keyword different from the masked keyword, Determine the wrong options based on the target keyword, and Constructing a multiple-choice question including the multiple-choice question stem, the correct options, and the incorrect options; The test question bank is constructed based on the multiple-choice questions corresponding to the corpus sample.
4. The method according to claim 3, characterized in that The objective questions include multiple-choice questions of multiple test question types; the obtaining of at least one corpus sample of the target application field includes: Determining a test question type of the target large language model according to a function of the target large language model; Based on the original data of the target application field, obtaining at least one corpus data group corresponding to each test question type, wherein any corpus data group includes at least one original data; For any test question type, semantic analysis is performed on at least one corpus data group corresponding to the test question type to determine at least one corpus sample; Accordingly, at least one keyword in the corpus sample is masked, including: At least one keyword in the corpus sample that matches the test question type corresponding to the corpus sample is masked.
5. The method according to claim 4, characterized in that The target large language model is used to achieve at least one of the following functions: dish category prediction, dish taste prediction, store business category prediction, store brand prediction, store package content prediction, and recommended store prediction; The test topic types include dish category prediction, dish taste prediction, store business category prediction, store brand prediction, At least one of brand prediction, store package content prediction, and recommended store prediction; The original data includes at least one of dish library data, store data and user order data.
6. The method according to claim 4, characterized in that The obtaining of a target keyword different from the masked keyword includes: Determine the number of incorrect options according to the number of correct options and the preset total number of options, and determine the difficulty of the target question corresponding to the corpus sample; According to the difficulty of the target question, a target keyword different from the masked keyword and having the number of wrong options is obtained, wherein the target keyword is obtained by at least one of the following methods: According to the difficulty of the target question and the masked category corresponding to the masked keyword, determine the wrong option phrase category corresponding to the corpus sample, and perform negative sampling in the phrase sample corresponding to the wrong option phrase category based on the masked keyword to obtain a target keyword different from the masked keyword; wherein the difficulty of the target question, the correlation between the masked category and the wrong option phrase category is positively correlated; Determine the edit distance of the wrong option corresponding to the corpus sample according to the difficulty of the target question, and obtain a target keyword that meets the edit distance of the wrong option and is different from the masked keyword; wherein the difficulty of the target question is negatively correlated with the edit distance of the wrong option; The similarity of the wrong option text corresponding to the corpus sample is determined according to the difficulty of the target question, and a target keyword different from the masked keyword that meets the similarity of the wrong option text is obtained; wherein the difficulty of the target question is positively correlated with the similarity of the wrong option text.
7. The method according to claim 6, characterized in that After constructing the multiple-choice question including the multiple-choice question stem, the correct options, and the incorrect options, the method further includes at least one of the following: According to the target question difficulty of the corpus sample, mark the difficulty of the multiple-choice questions corresponding to the corpus sample; According to the test question type of the corpus sample, the question type of the multiple-choice questions corresponding to the corpus sample is marked.
8. The method according to claim 1, characterized in that The objective questions include judgment questions; before obtaining the test question bank corresponding to the target large language model based on the target application field to which the target large language model to be evaluated is applied, the method further includes: Acquire at least one corpus sample of the target application field; For any corpus sample, Mask at least one keyword in the corpus sample. Determine the masked keyword or other keywords different from the masked keyword as core words, Construct the judgment question stem based on the masked corpus sample and the core words, Determine the answer to the judgment question based on whether the core word is the masked keyword, and Construct a true or false question including the true or false question stem and the true or false question answer.
9. The method according to claim 8, characterized in that The step of determining the masked keyword or other keywords different from the masked keyword as core words includes: Determine the difficulty of the target questions corresponding to the corpus sample; When the target question is easy, the masked keyword, other keywords belonging to different main categories from the masked keyword, other keywords whose edit distance from the masked keyword is greater than a first edit distance, or other keywords whose text similarity with the masked keyword is greater than a first similarity, or other keywords whose text similarity with the masked keyword is less than the first similarity are obtained as the core words; When the target topic difficulty is medium difficulty, obtain other keywords belonging to different subcategories under the same main category as the masked keyword, and the edit distance between the other keywords and the masked keyword is less than or equal to the first edit distance. other keywords whose text similarity with the masked keyword is greater than or equal to the first similarity and less than the second similarity as the core word; When the difficulty of the target question is difficult, other keywords belonging to the same subcategory as the masked keyword, other keywords whose edit distance with the masked keyword is less than or equal to the second edit distance, or other keywords whose text similarity with the masked keyword is greater than or equal to the second similarity are obtained as the core words.
10. The method according to any one of claims 1 to 9, characterized in that The objective questions include multiple-choice questions, and the tested objective questions include tested multiple-choice questions; and the evaluating the target large language model based on the answers corresponding to the tested objective questions and the output results of the target large language model for the tested objective questions includes: Extracting the predicted options of the target large language model for the tested multiple-choice question based on the output result, marking whether the predicted options are correct based on the correct options in the answers corresponding to the tested multiple-choice question, and evaluating the target large language model according to the marked results of the tested multiple-choice question, wherein the predicted options are extracted by at least one of the following methods: Determining the predicted option of the target large language model for the multiple-choice question to be tested based on the predicted probabilities of different options included in the output result; Extracting, from the output result, the predicted options of the target large language model for the multiple-choice question being tested according to a preset answer extraction rule; The output result is fuzzy matched according to the correct options of the multiple-choice question to be tested, and the predicted options of the target large language model for the multiple-choice question to be tested are determined.
11. The method according to claim 10, characterized in that The testing of the target large language model using the tested multiple-choice questions includes: Based on the multiple-choice question stem, correct options and incorrect options corresponding to the multiple-choice question being tested, a test sentence of the multiple-choice question being tested is synthesized, and The target large language model is tested using the test sentence.
12. The method according to any one of claims 1 to 9, characterized in that The objective questions are marked with corresponding question types and question difficulties; The testing of the target large language model using the test question bank includes: Obtaining multiple objective questions of different difficulty levels corresponding to each question type in the test question bank as the objective questions to be tested, and Testing the target large language model using the objective questions to be tested; Accordingly, based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested, evaluating the target large language model includes: Mark whether the output result is correct based on the answer corresponding to the objective question being tested, According to the marking results of each objective question, the test accuracy rate corresponding to each question type is calculated, and The prediction accuracy of the target large language model is determined based on the test accuracy of each question type.
13. The method according to any one of claims 1 to 9, characterized in that The test question bank also includes objective questions belonging to a general application field; using the test question bank to test the target large language model includes: Acquire multiple objective questions in the target application field from the test question bank as first objective questions to be tested, and obtaining a plurality of objective questions in the general application field as second objective questions to be tested, The objective questions to be tested include the first objective questions to be tested and the second objective questions to be tested.
14. A large language model evaluation device, characterized in that: The device comprises: A question bank acquisition module is used to acquire a test question bank corresponding to the target large language model to be evaluated based on the target application field to which the target large language model is applied, wherein the test question bank contains objective questions belonging to the target application field, and the stems and answers of the objective questions are obtained by analyzing corpus samples of the target application field; A model testing module, used to test the target large language model using the test question bank to obtain an output result of the target large language model for the tested objective question; The model evaluation module is used to evaluate the target large language model based on the answer corresponding to the objective question being tested and the output result of the target large language model for the objective question being tested.
15. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
16. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Answer evaluation method, language model training method and related device
CN115146122A
Question automatic generation method and system
CN115934908A
Question and answer scoring method, question and answer scoring device, electronic equipment and storage medium
CN116561538A
Assessment method and device of large language model, storage medium and computer equipment
CN117291184A
Assessment method and device of large language model, storage medium and computer equipment
CN118627511A
Cited By
Intelligent cabin AI voice interaction test method, computer equipment and storage medium
CN120932630A
Intelligent cabin AI voice interaction test method, computer device, and storage medium
CN120932630B
Engineering investigation question bank generation method and system based on retrieval enhancement and rule engine
CN121255940A
Question and answer method and device based on large model, medium, equipment and program product
CN121303368A
Large language model comprehensive capability assessment method and related device
CN121581225A