Large model reply quality evaluation method, device, equipment and medium
By obtaining test data sets, configuring prompt words and scoring the reply text, the difficulty of reply quality evaluation in large language models is solved, and automated evaluation and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202411942717.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
There are difficulties in the evaluation of response quality of large language models in multiple fields and scenarios, and there is a lack of effective evaluation methods.
By obtaining the test dataset, configuring the prompt words, inputting the model to be tested to obtain the reply text, and comparing the reply text with the standard answers for a comprehensive score.
A comprehensive and automated evaluation of the reply content of large language models is realized, reducing manual intervention and subjective judgment, and improving evaluation efficiency.
Smart Images

Figure CN119939181A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a large model response quality assessment method, device, equipment and medium. Background Art
[0002] With the rapid development of artificial intelligence technology and the widespread application of deep learning in the field of natural language processing (NLP), semantic analysis technology based on large models has made significant progress. These large models, such as BERT and GPT series, have learned rich language knowledge and context understanding capabilities through pre-training on massive text data, and can accurately capture semantic information in texts, performing well in tasks such as semantic analysis and text generation.
[0003] With the widespread application of large language models, their comprehensiveness and multi-domain adaptability enable them to generate responses in multiple scenarios, showing great application potential. However, since current large language models can handle a variety of problems in different fields and scenarios, it is difficult to evaluate the quality of their responses. Therefore, how to comprehensively evaluate the responses of large language models is an urgent problem to be solved. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a large model response quality evaluation method, device, equipment and medium.
[0005] In a first aspect, the present disclosure provides a large model response quality evaluation method, comprising:
[0006] Get the test data set, which includes questions and standard answers;
[0007] Configure prompt words for the questions in the test data set. The prompt words are used to indicate the answer direction of the questions in the test data set.
[0008] Input the questions and corresponding prompt words in the test data set into the model under test to obtain the response text output by the model under test;
[0009] Compare the reply text with the standard answer, and give the reply text a comprehensive score based on the comparison results.
[0010] Optionally, configure prompt words for the questions in the test dataset, including:
[0011] For each question in the test data set, specify at least one of the output format, language style, and emotional color of the answer to the question.
[0012] Optionally, the questions and corresponding prompt words in the test data set are input into the model under test, including:
[0013] The application program interface of the model under test is called, and the questions and corresponding prompt words in the test data set are input into the model under test through the application program interface of the model under test.
[0014] Optionally, the test data set includes a subjective question set, the subjective question set includes subjective questions and corresponding subjective question answers, and correspondingly, the response text includes subjective question response text;
[0015] Compare the reply text with the standard answer and give a comprehensive score to the reply text based on the comparison results, including:
[0016] Calculate the text similarity between the subjective question response text and the subjective question answer;
[0017] The comprehensive score of the subjective question response text is determined based on the text similarity.
[0018] Optionally, calculating the text similarity between the subjective question reply text and the subjective question answer includes:
[0019] Calling the referee model to calculate the semantic similarity score and the lexical similarity score between the subjective question response text and the subjective question answer, and determining the partial match score, the order match score and the penalty factor of the subjective question response text relative to the subjective question answer;
[0020] The comprehensive score of the subjective question response text is determined based on the text similarity, including:
[0021] The semantic similarity score, lexical similarity score, partial match score, order match score and penalty factor are combined to determine the comprehensive score of the subjective question response text.
[0022] Optionally, the test data set further includes an objective question set, the objective question set includes objective questions and corresponding objective question answers, and correspondingly, the response text includes objective question response text;
[0023] Compare the reply text with the standard answer and give a comprehensive score to the reply text based on the comparison results, including:
[0024] The accuracy of the objective question response text is determined based on whether the objective question response text is consistent with the objective question answer;
[0025] The comprehensive score of the objective question response text is determined based on the accuracy of the objective question response text.
[0026] Optionally, before comparing the reply text with the standard answer, the following is also included:
[0027] Post-process the reply text to obtain a post-processed reply text;
[0028] Among them, post-processing of the text of the subjective question responses includes cleaning irrelevant information, text sentence segmentation and filtering, abstract extraction, text reorganization and formatting, grammar correction, entity standardization, and removal of semantic noise;
[0029] Post-processing of the objective question response text includes using regular expressions to extract the corresponding options in the objective question response text;
[0030] Compare the response text with the model answer, including:
[0031] Compare the post-processed response text with the standard answer.
[0032] In a second aspect, the present disclosure provides a large model response quality evaluation device, comprising:
[0033] A data collection module is used to obtain a test data set, which includes questions and standard answers;
[0034] A prompt word configuration module is used to configure prompt words for questions in the test data set, and the prompt words are used to indicate the answer direction of the questions in the test data set;
[0035] An execution module is used to input the questions and corresponding prompt words in the test data set into the model under test to obtain the reply text output by the model under test;
[0036] The evaluation module is used to compare the reply text with the standard answer and give a comprehensive score to the reply text based on the comparison results.
[0037] Optionally, when configuring prompt words for questions in the test data set, the prompt word configuration module is specifically used to:
[0038] For each question in the test data set, specify at least one of the output format, language style, and emotional color of the answer to the question.
[0039] Optionally, when the execution module inputs the questions and corresponding prompt words in the test data set into the model under test, it is specifically used to:
[0040] The application program interface of the model under test is called, and the questions and corresponding prompt words in the test data set are input into the model under test through the application program interface of the model under test.
[0041] Optionally, the test data set includes a subjective question set, the subjective question set includes subjective questions and corresponding subjective question answers, and correspondingly, the response text includes subjective question response text;
[0042] When the evaluation module compares the reply text with the standard answer and gives a comprehensive score to the reply text based on the comparison result, it is specifically used to:
[0043] Calculate the text similarity between the subjective question response text and the subjective question answer;
[0044] The comprehensive score of the subjective question response text is determined based on the text similarity.
[0045] Optionally, when calculating the text similarity between the subjective question reply text and the subjective question answer, the evaluation module is specifically used to:
[0046] Calling the referee model to calculate the semantic similarity score and the lexical similarity score between the subjective question response text and the subjective question answer, and determining the partial match score, the order match score and the penalty factor of the subjective question response text relative to the subjective question answer;
[0047] When the evaluation module determines the comprehensive score of the subjective question response text based on the text similarity, it is specifically used to:
[0048] The semantic similarity score, lexical similarity score, partial match score, order match score and penalty factor are combined to determine the comprehensive score of the subjective question response text.
[0049] Optionally, the test data set further includes an objective question set, the objective question set includes objective questions and corresponding objective question answers, and correspondingly, the response text includes objective question response text;
[0050] When the evaluation module compares the reply text with the standard answer and gives a comprehensive score to the reply text based on the comparison result, it is specifically used to:
[0051] The accuracy of the objective question response text is determined based on whether the objective question response text is consistent with the objective question answer;
[0052] The comprehensive score of the objective question response text is determined based on the accuracy of the objective question response text.
[0053] Optionally, the large model reply quality assessment device further includes a post-processing module, and before comparing the reply text with the standard answer, the post-processing module is used to:
[0054] Post-process the reply text to obtain a post-processed reply text;
[0055] Among them, post-processing of the text of the subjective question responses includes cleaning irrelevant information, text sentence segmentation and filtering, abstract extraction, text reorganization and formatting, grammar correction, entity standardization, and removal of semantic noise;
[0056] Post-processing of the objective question response text includes using regular expressions to extract the corresponding options in the objective question response text;
[0057] When comparing the response text with the standard answer, the evaluation module is specifically used to:
[0058] Compare the post-processed response text with the standard answer.
[0059] In a third aspect, the present disclosure provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method as described in any one of the first aspects is implemented.
[0060] In a fourth aspect, the present disclosure provides a computer-readable storage medium having program instructions stored thereon, which implement any method of the first aspect when the program instructions are executed.
[0061] In a fifth aspect, the present disclosure provides a computer program product, which is stored in a storage medium. When the program product is executed, the method of the first aspect can be implemented.
[0062] Compared with the prior art, the technical solution provided by the present invention has the following advantages:
[0063] The large model response quality evaluation method, device, equipment and medium provided by the present disclosure obtains a test data set, the test data set includes questions and standard answers, and then configures prompt words for the questions in the test data set, the prompt words are used to indicate the answer direction of the questions in the test data set, and then the questions in the test data set and the corresponding prompt words are input into the tested model to obtain the response text output by the tested model, and finally the response text is compared with the standard answer, and the response text is comprehensively scored according to the comparison result. The present disclosure can perform comprehensive and automated evaluation on the response content of the model, greatly reducing manual intervention and subjective judgment, and improving the efficiency of the evaluation work. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0066] Figure 1 A flow chart of a large model response quality evaluation method provided in an embodiment of the present disclosure;
[0067] Figure 2 A schematic diagram of the structure of a large model response quality evaluation device provided in an embodiment of the present disclosure;
[0068] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0069] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0070] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0071] Figure 1 The present invention provides a flow chart of a large model response quality evaluation method, which can be performed by a large model response quality evaluation device, which can be implemented in software and / or hardware, and can be configured in an electronic device, including a server or a terminal. Figure 1 As shown, the large model response quality evaluation method includes the following steps:
[0072] S101. Obtain a test data set, where the test data set includes questions and standard answers.
[0073] Exemplarily, test questions of different types and fields can be collected from preset channels to form a test data set; wherein the preset channels include at least one of the following: academic websites, corporate websites, open source community websites, industry websites, data sets generated by various tools, and user feedback data.
[0074] Test datasets can be public datasets from academia, enterprises or open source communities, professional datasets from specific industries, or generated by crawling web pages, synthesizing text or using specific tools, and the content of the dataset can be continuously optimized through user testing and feedback. Manual classification can be performed based on the diversity, representativeness, consistency and accuracy of the dataset, and the classified dataset can also be uploaded to a third-party public evaluation tool for evaluation, and the quality of the test dataset can be distinguished based on the evaluation results. The evaluation dataset can also be divided into different levels of difficulty through manual review. For example, the following difficulty levels are usually set: elementary, intermediate and advanced.
[0075] The questions in the elementary-level test data set are relatively direct and clear. The knowledge or information involved in the questions are relatively basic and common. Answering these questions does not involve complex reasoning or logical operations, and the model may only need to perform simple matching or classification based on the input questions to get the answers to the questions.
[0076] Questions in the intermediate-level test data set require a certain degree of reasoning and logical inference. The knowledge or information involved in the questions is relatively broad and may span multiple fields or topics. The questions may require the model to find implicit information in the text or make appropriate inferences to get the answer. In addition, the model may need to perform multiple steps or comprehensively consider multiple factors to get an accurate answer or result.
[0077] Questions in the advanced test data sets require complex reasoning and logical operations, and may involve multiple steps or complex conditional judgments. In addition, the knowledge or information involved in the questions is relatively in-depth and professional, and the model may need to have a high level of domain knowledge and comprehension. The questions may require the model to integrate and understand multimodal information, such as images, voice, etc. In addition, the model may be required to make decisions or judgments on the questions in actual application scenarios.
[0078] S102: configuring prompt words for the questions in the test data set, where the prompt words are used to indicate the answering direction of the questions in the test data set.
[0079] Configuring effective prompts based on the questions in the test dataset can speed up the model's response generation, because the prompts provide clear directions and restrictions for the model under test to answer these questions. This reduces the irrelevant exploration that the model may conduct during the generation process, making resources more efficiently utilized. In addition, by using prompts, users can have more detailed control over the response text generated by the model under test, such as unifying the format of the response text to facilitate subsequent scoring of the response text.
[0080] In some embodiments, configuring prompt words for questions in the test data set includes: for each question in the test data set, specifying at least one of an output format, a language style, and an emotional color for replying to the question.
[0081] By configuring prompt words to specify the format, language style, and emotional color of the text output by the tested model to the question, you can ensure that the generated content is highly consistent with user expectations. The following takes a common model tax data prompt word as an example to explain how to create a configuration prompt word:
[0082] Please evaluate the quality of the AI assistant's answers to user questions as an impartial judge. Your evaluation should consider the usefulness, accuracy, and comprehensiveness of the response. When conducting the evaluation, you need to follow the following process: 1. Compare the AI assistant's answers with the reference answers, identify and point out the strengths or weaknesses of the AI assistant's answers. Be as objective as possible. 2. Evaluate the AI assistant's answers from three dimensions: usefulness, accuracy, and comprehensiveness. Briefly describe the performance of each dimension and give a score ranging from 1 to 5 points (1 point represents the lowest and 5 points represent the highest). 3. Calculate the overall score based on the scores of these dimensions. When scoring, please strictly follow the following JSON format output and ensure that the scores are all integers. Scoring format example: {"relevance": score 1, "accuracy": score 2, "comprehensiveness": score 3, "overall score": score 4}.
[0083] S103: Input the questions and corresponding prompt words in the test data set into the model under test to obtain the reply text output by the model under test.
[0084] The questions and corresponding prompt words in the test data set are input into the large language model under test, and the model under test outputs the reply text that answers the question according to the prompt words.
[0085] In some embodiments, inputting the questions and corresponding prompt words in the test data set into the model under test includes: calling the application program interface of the model under test, and inputting the questions and corresponding prompt words in the test data set into the model under test through the application program interface of the model under test.
[0086] In actual application scenarios, the Application Programming Interface (API) is currently one of the most common ways to open up large-model AI capabilities. Requesting a large model by calling the API usually involves several key steps, including preparing request data, sending HTTP (HyperText Transfer Protocol) requests, processing responses, and parsing returned data.
[0087] S104, comparing the reply text with the standard answer, and comprehensively scoring the reply text based on the comparison result.
[0088] Before the comparison, the reply text is preprocessed to ensure the consistency of the format. The reply text generated by the tested model is compared with the standard answer. The similarity calculation method can be used to evaluate the similarity between the reply text and the standard answer, so as to judge its accuracy and relevance. On this basis, the reply text is comprehensively scored in combination with multiple evaluation dimensions to form an evaluation result reflecting the quality of the reply text.
[0089] The disclosed embodiment obtains a test data set, which includes questions and standard answers, and then configures prompt words for the questions in the test data set, which are used to indicate the answer direction of the questions in the test data set. Then, the questions in the test data set and the corresponding prompt words are input into the tested model to obtain the reply text output by the tested model, and finally the reply text is compared with the standard answer, and the reply text is comprehensively scored according to the comparison result. The disclosed embodiment can perform comprehensive and automatic evaluation on the reply content of the model, greatly reducing manual intervention and subjective judgment, and improving the efficiency of the evaluation work.
[0090] In some embodiments, the test data set includes a subjective question set, the subjective question set includes subjective questions and corresponding subjective question answers, and the response text includes subjective question response text. The response text is compared with the standard answer, and the response text is comprehensively scored according to the comparison result, including: calculating the text similarity between the subjective question response text and the subjective question answer; and determining the comprehensive score of the subjective question response text according to the text similarity.
[0091] Exemplarily, the subjective question may be a question in a text summary dataset, such as translating the content of a text summary, or performing reading comprehension on the text summary and summarizing the content of the text summary. The answer to the subjective question may be a reference translation text or a reference summary.
[0092] A variety of text similarity calculation methods are used to compare the reply text with the standard answer. Common similarity calculation methods include: Based on exact match: Check whether the reply text is exactly the same as the standard answer. Lexical overlap: Calculate the number of common words between the reply text and the standard answer to determine their similarity. Semantic similarity: Use embedding models (such as Word2Vec, BERT, etc.) to convert the text into a vector and calculate the cosine similarity to measure the semantic similarity. Then determine the comprehensive score of the subjective question reply text based on the similarity between the reply text and the standard answer.
[0093] In some embodiments, before comparing the reply text with the standard answer, the method further includes: post-processing the reply text to obtain a post-processed reply text; wherein the post-processing of the reply text for the subjective question includes cleaning irrelevant information, text sentence segmentation and filtering, summary extraction, text reorganization and formatting, grammar correction, entity standardization, and semantic noise removal; the post-processing of the reply text for the objective question includes using regular expressions to extract corresponding options in the reply text for the objective question. Accordingly, comparing the reply text with the standard answer includes: comparing the post-processed reply text with the standard answer.
[0094] The subjective question response text generated by the tested model is a piece of text containing a summary. The summary information can be accurately extracted from the text generated by the model through corresponding post-processing steps, including cleaning up irrelevant information, text sentence segmentation and filtering, summary extraction, text reorganization and formatting, grammar correction, entity standardization, removal of semantic noise, and output verification.
[0095] The objective question response text output by the tested model for objective questions containing questions and options usually contains the options selected by the tested model as answers. The objective question response text is post-processed to extract the options that the model believes to be correct. Regular expressions are used to match text such as "the answer is A" and "choose B" to extract the letter options (A, B, C, or D). Ensure that the extracted answers are standardized, for example: capital letters (A, B, C, D), and remove irrelevant content (such as the "most likely option is" part of "the most likely option is A"). Post-processing the objective question response text can ensure that the objective question response is a clear option (A, B, C, or D).
[0096] The disclosed embodiment performs post-processing on the reply text to remove information in the reply text that is irrelevant to the answer to the question, thereby reducing errors in comparing subsequent reply texts with standard answers and making the comprehensive score of the reply text more accurate.
[0097] In some embodiments, calculating the text similarity between the subjective question reply text and the subjective question answer includes: calling the referee model, calculating the semantic similarity score and the lexical similarity score between the subjective question reply text and the subjective question answer, and determining the partial match score, order match score and penalty factor of the subjective question reply text relative to the subjective question answer. Accordingly, determining the comprehensive score of the subjective question reply text based on the text similarity includes: combining the semantic similarity score, the lexical similarity score, the partial match score, the order match score and the penalty factor to determine the comprehensive score of the subjective question reply text.
[0098] In the disclosed embodiment, the referee model can comprehensively evaluate the performance of the large model under test through a series of preset standards or evaluation indicators. These evaluation indicators may include multiple dimensions such as model accuracy, consistency, logical reasoning ability, language generation quality, etc. The referee model can objectively quantify these indicators and provide reliable data support for the evaluation. The type of referee model is usually different from the type of the model under test, and the referee model can also be used by calling an API.
[0099] The lexical similarity score can be measured using the Rouge-N evaluation index, which evaluates the similarity between texts by comparing the overlap of consecutive n words (n-grams). The semantic similarity can be measured using the cosine similarity evaluation index, which measures the semantic similarity between texts by representing the texts as vectors and calculating the cosine value of the angle between them. The Rouge-N value and / or cosine similarity value between the response text and the answer to each subjective question are calculated to measure the similarity between the response text and the answer to the subjective question generated by the tested model.
[0100] You can also use the BLEU metric or the Meteor metric to measure the similarity between texts. BLEU calculates the number of n-grams that overlap between the generated text and the reference text. For example, when n=1, unigram is calculated, when n=2, bigram is calculated, and so on. The accuracy of the overlap between the two sides is calculated and multiplied by the length penalty to get the final score. Meteor calculates the number of phrases that completely match the translation result and the reference answer, and combines partial matches, order changes, and penalty factors to calculate the final score.
[0101] The precision score is the match between the generated reply text and the reference answer text at the vocabulary and phrase level, including word form restoration and synonym matching. The order matching score is a measure of whether the matching words or phrases in the reference answer text appear in the generated reply text in order. The order score can be calculated by penalizing misplaced matches. The penalty factor is introduced based on repetition, text length, grammatical errors, etc. Among them, α, β, and γ are weight factors that can be adjusted according to task requirements.
[0102] Finally, the final comprehensive score of the subjective question response text is calculated by combining the scores of these evaluation criteria: the influence of each factor on the final score can be further adjusted. For example, partial matching may be more important, so its weight in the final score can be increased; if the order change is very important to the task, the penalty for sequential matching can be increased. By combining partial matching, order change, and penalty factors, the similarity between the generated text and the reference text can be more comprehensively evaluated. This multi-dimensional evaluation method is more flexible and accurate than simple n-gram matching (such as BLEU), especially when dealing with complex text generation tasks, it can more fully consider multiple factors such as semantics, structure, and fluency.
[0103] In addition, the quality of the code generated by the model under test can be evaluated using the Pass@k metric. The Pass@k metric is used to measure the proportion of the code generated by the model under test that passes specific test conditions in a specific test set, which helps to evaluate the effectiveness and correctness of the code.
[0104] For the subjective question answering text output by the tested model for the subjective question set, an overall evaluation is performed. The average values of the Rouge-N, BLEU, and Meteor indicators of all the subjective question answering texts can be calculated to assess the overall quality of the subjective question answering results output by the tested evaluation model.
[0105] In some embodiments, the test data set further includes an objective question set, the objective question set includes objective questions and corresponding objective question answers, and correspondingly, the response text includes objective question response text;
[0106] The reply text is compared with the standard answer, and the reply text is comprehensively scored based on the comparison result, including: determining the accuracy of the objective question reply text based on whether the objective question reply text is consistent with the objective question answer; determining the comprehensive score of the objective question reply text based on the accuracy of the objective question reply text.
[0107] Objective questions can be multiple-choice questions that include questions and options. After obtaining the post-processed objective question response text, it is compared with the standard answers to the objective questions in the data set. If the options of the model's objective question response text are exactly the same as the options of the standard answer, then it can be confirmed that the model under test has given the correct answer to this objective question. In this process, the output of the model under test for each objective question is compared with the standard answer of the objective question one by one to determine the overall correct rate of the model under test's response to the objective question set. Based on this correct rate, the comprehensive score of the objective question response text can be determined. For example, if the entire objective question set includes 100 objective questions and the correct rate is 90%, the comprehensive score of the objective question response text can be determined as 90 points. If the length of the comprehensive score of the objective question response text is limited to a full score of 5 points, the comprehensive score of the objective question response text can be obtained as 4.5 points by scaling the score.
[0108] The disclosed embodiment sets subjective questions and objective questions in a test data set, determines the comprehensive score of the subjective question answers by calculating the similarity between the answers to the subjective questions of the tested model and the standard answers, and determines the comprehensive score of the objective question answers by calculating the accuracy of the answers to the objective questions of the tested model relative to the standard answers, thereby making the evaluation of the responses of the tested model more comprehensive and accurate.
[0109] Figure 2The schematic diagram of the structure of the large model reply quality evaluation device provided by the embodiment of the present disclosure. The large model reply quality evaluation device provided by the embodiment of the present disclosure can execute the processing flow provided by the large model reply quality evaluation method embodiment, such as Figure 2 As shown, the large model response quality evaluation device 200 includes:
[0110] The data collection module 201 is used to obtain a test data set, which includes questions and standard answers;
[0111] A prompt word configuration module 202, configured to configure prompt words for the questions in the test data set, the prompt words are used to indicate the answer direction of the questions in the test data set;
[0112] An execution module 203 is used to input the questions and corresponding prompt words in the test data set into the model under test to obtain the reply text output by the model under test;
[0113] The evaluation module 204 is used to compare the reply text with the standard answer and give a comprehensive score to the reply text according to the comparison result.
[0114] In some embodiments, when configuring prompt words for questions in the test data set, the prompt word configuration module 202 is specifically used to: for each question in the test data set, specify at least one of the output format, language style and emotional color of the reply to the question.
[0115] In some embodiments, when the execution module 203 inputs the questions and corresponding prompt words in the test data set into the model under test, it is specifically used to: call the application program interface of the model under test, and input the questions and corresponding prompt words in the test data set into the model under test through the application program interface of the model under test.
[0116] In some embodiments, the test data set includes a subjective question set, which includes subjective questions and corresponding answers to the subjective questions. Correspondingly, the reply text includes the subjective question reply text; when the evaluation module 204 compares the reply text with the standard answer and gives a comprehensive score to the reply text based on the comparison result, it is specifically used to: calculate the text similarity between the subjective question reply text and the subjective question answer; determine the comprehensive score of the subjective question reply text based on the text similarity.
[0117] In some embodiments, when calculating the text similarity between the subjective question response text and the subjective question answer, the evaluation module 204 is specifically used to: call the referee model, calculate the semantic similarity score and the lexical similarity score between the subjective question response text and the subjective question answer, and determine the partial match score, order match score and penalty factor of the subjective question response text relative to the subjective question answer; when determining the comprehensive score of the subjective question response text based on the text similarity, the evaluation module 204 is specifically used to: determine the comprehensive score of the subjective question response text by combining the semantic similarity score, lexical similarity score, partial match score, order match score and penalty factor.
[0118] In some embodiments, the test data set also includes an objective question set, which includes objective questions and corresponding objective question answers, and correspondingly, the reply text includes objective question reply text; when the evaluation module 204 compares the reply text with the standard answer and performs a comprehensive score on the reply text based on the comparison result, it is specifically used to: determine the accuracy of the objective question reply text based on whether the objective question reply text is consistent with the objective question answer; determine the comprehensive score of the objective question reply text based on the accuracy of the objective question reply text.
[0119] In some embodiments, the large model reply quality assessment device further includes a post-processing module 205. Before comparing the reply text with the standard answer, the post-processing module is used to: post-process the reply text to obtain a post-processed reply text; wherein the post-processing of the reply text for the subjective question includes cleaning irrelevant information, text sentence segmentation and filtering, summary extraction, text reorganization and formatting, grammar correction, entity standardization, and semantic noise removal; the post-processing of the reply text for the objective question includes extracting corresponding options in the reply text for the objective question using regular expressions;
[0120] When comparing the reply text with the standard answer, the evaluation module 204 is specifically used to compare the post-processed reply text with the standard answer.
[0121] Figure 2 The large model response quality evaluation device of the illustrated embodiment can be used to implement the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effect are similar and will not be repeated here.
[0122] Figure 3 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 3 , which shows a structural schematic diagram of an electronic device 300 suitable for implementing the embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0123] like Figure 3As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303 to implement the large model reply quality evaluation method of the embodiment described in the present disclosure. In RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0124] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0125] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart, thereby implementing the large model response quality evaluation method as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0126] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0127] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0128] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0129] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0130] Get the test data set, which includes questions and standard answers;
[0131] Configure prompt words for the questions in the test data set. The prompt words are used to indicate the answer direction of the questions in the test data set.
[0132] Input the questions and corresponding prompt words in the test data set into the model under test to obtain the response text output by the model under test;
[0133] Compare the reply text with the standard answer, and give the reply text a comprehensive score based on the comparison results.
[0134] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.
[0135] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0136] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0137] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0138] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0139] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] The embodiments of the present disclosure also provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any of the above embodiments can be implemented. The execution method and beneficial effects are similar and will not be repeated here.
[0141] The embodiments of the present disclosure also provide a computer program product, which is stored in a storage medium. When the program product is run, the method of any of the above embodiments can be implemented. The execution method and beneficial effects are similar and will not be repeated here.
[0142] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0143] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0144] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
[0145] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0146] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large model response quality evaluation method, characterized in that: include: Obtain a test data set, wherein the test data set includes questions and standard answers; Configuring prompt words for the questions in the test data set, wherein the prompt words are used to indicate the answer direction of the questions in the test data set; Input the questions and corresponding prompt words in the test data set into the model under test to obtain the reply text output by the model under test; The reply text is compared with the standard answer, and a comprehensive score is given to the reply text based on the comparison result.
2. The large model response quality evaluation method according to claim 1, characterized in that: The configuring prompt words for the questions in the test data set includes: For each question in the test data set, at least one of an output format, a language style, and an emotional color for responding to the question is specified.
3. The large model response quality evaluation method according to claim 1, characterized in that: The step of inputting the questions and corresponding prompt words in the test data set into the model under test comprises: The application program interface of the model under test is called, and the questions and corresponding prompt words in the test data set are input into the model under test through the application program interface of the model under test.
4. The large model response quality evaluation method according to claim 1, characterized in that: The test data set includes a subjective question set, the subjective question set includes subjective questions and corresponding subjective question answers, and correspondingly, the response text includes subjective question response text; The step of comparing the reply text with the standard answer and comprehensively scoring the reply text according to the comparison result includes: Calculating the text similarity between the subjective question reply text and the subjective question answer; The comprehensive score of the subjective question reply text is determined according to the text similarity.
5. The large model response quality evaluation method according to claim 4, characterized in that: The calculating the text similarity between the subjective question reply text and the subjective question answer includes: Calling the referee model to calculate the semantic similarity score and the lexical similarity score between the subjective question reply text and the subjective question answer, and determining the partial match score, the order match score and the penalty factor of the subjective question reply text relative to the subjective question answer; Determining the comprehensive score of the subjective question reply text according to the text similarity includes: The comprehensive score of the subjective question reply text is determined by combining the semantic similarity score, the lexical similarity score, the partial match score, the order match score and the penalty factor.
6. The large model response quality evaluation method according to claim 1, characterized in that: The test data set also includes an objective question set, the objective question set includes objective questions and corresponding objective question answers, and correspondingly, the response text includes objective question response text; The step of comparing the reply text with the standard answer and comprehensively scoring the reply text according to the comparison result includes: Determining the accuracy of the objective question response text according to whether the objective question response text is consistent with the objective question answer; The comprehensive score of the objective question response text is determined according to the accuracy rate of the objective question response text.
7. The large model response quality evaluation method according to claim 4 or 6, characterized in that: Before comparing the reply text with the standard answer, the method further includes: Post-processing the reply text to obtain a post-processed reply text; The post-processing of the subjective question reply text includes cleaning irrelevant information, text sentence segmentation and filtering, abstract extraction, text reorganization and formatting, grammar correction, entity standardization, and semantic noise removal; Post-processing the objective question response text includes extracting corresponding options in the objective question response text using regular expressions; The comparing the reply text with the standard answer includes: The post-processed reply text is compared with the standard answer.
8. A large model response quality evaluation device, characterized in that: include: A data collection module, used to obtain a test data set, wherein the test data set includes questions and standard answers; A prompt word configuration module, used to configure prompt words for the questions in the test data set, wherein the prompt words are used to indicate the answer direction of the questions in the test data set; An execution module, used for inputting the questions and corresponding prompt words in the test data set into the model under test to obtain a reply text output by the model under test; The evaluation module is used to compare the reply text with the standard answer and give a comprehensive score to the reply text according to the comparison result.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: Program instructions are stored thereon, and when the program instructions are executed, the method according to any one of claims 1 to 7 is implemented.