Question and answer model evaluation method and device

By constructing an evaluation dataset and a large model to generate standard answers and step sequences, the evaluation problem of the question-answering model in multi-step, process-based instruction scenarios is solved, and the effective evaluation of the answers output by the question-answering model and the enhancement of its conversational capabilities are achieved.

CN120706581AActive Publication Date: 2025-09-26TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511225781.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-09-26
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively evaluate the accuracy and completeness of the answers output by question-answering models under the question-answering requirements of multi-step and process-based instructions.

Method used

Construct an evaluation dataset, including historical question-answer pairs, multiple instruction documents, and user query question-answer pairs. Generate standard answers and step sequences through a large model. Use prompt words to guide the question-answering model to be tested to output predicted answers and step sequences, and evaluate by comparing the generated answers and step sequences.

Benefits of technology

It achieves effective evaluation of the answers output by the question-answering model in multi-step, process-based instruction scenarios, enhances the assessment of dialogue capabilities, and supports multi-round contextual reasoning and continuous contextual interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706581A_ABST
    Figure CN120706581A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a question and answer model evaluation method and device, and the method comprises the steps: inputting a historical question and answer pair in an evaluation data set, a plurality of instruction documents, a question in a user query question and answer pair, and a first prompt word into a to-be-tested question and answer model, obtaining a predicted answer and a predicted step sequence of a question in the user query question-answer pair, which are reasoned and output by the question-answer model to be tested according to the first cue word; wherein the multiple instruction documents are obtained based on user query text retrieval, questions in the historical question and answer pairs and the user query question and answer pairs are generated by the first large model based on the instruction documents, and the standard answers and the standard step sequences are generated by the first large model based on the questions in the historical question and answer pairs, the multiple instruction documents and the user query question and answer pairs; comparing the standard answer with the prediction answer, and comparing the standard step sequence with the prediction step sequence to obtain an evaluation result. According to the method, the to-be-tested question and answer model is effectively evaluated under the scene of a multi-step and process instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a question-answering model evaluation method and device. Background Art

[0002] With the rapid development of large-scale language models, they have been widely used in question-answering, text generation, and human-computer dialogue. Evaluating the performance of question-answering models (such as accuracy and completeness) is a primary means of evaluating their quality. However, current mainstream question-answering model evaluation methods still focus on factual questions and answers based on single texts. However, they are unable to effectively evaluate the output of question-answering models for real-world, everyday question-answering scenarios involving multi-step, process-based instructions, particularly when faced with instruction documents from multiple heterogeneous sources. Summary of the Invention

[0003] The present invention provides a question-answering model evaluation method and device to solve the problem in the prior art that when users have question-answering requirements for multi-step, process-based instructions, the output answers of the question-answering model cannot be effectively evaluated.

[0004] The present invention provides a question-answering model evaluation method, comprising the following steps: The historical question-answer pairs, multiple instruction documents, questions in the user query question-answer pairs, and the first prompt word in the evaluation data set are input into the question-answer model to be tested, and the predicted answers and predicted step sequences of the questions in the user query question-answer pairs are obtained by the question-answer model to be tested according to the reasoning output of the first prompt word. The first prompt word is used to instruct the question-answer model to be tested to analyze the multiple instruction documents and the historical question-answer pairs to reason and output the predicted answers and predicted step sequences of the questions in the user query question-answer pairs. The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query texts, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs. Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained.

[0005] According to a question-answering model evaluation method provided by the present invention, before inputting historical question-answer pairs, multiple instruction documents, questions in user query question-answer pairs, and a first prompt word in an evaluation dataset into a question-answering model to be tested, and obtaining a predicted answer and a sequence of prediction steps for the question-answering model to be tested based on the first prompt word, the method further includes: Obtaining the user query text; Acquire a plurality of instruction documents retrieved according to the user query text; Inputting the plurality of instruction documents, user query texts, and a second prompt word into the first large model, obtaining the historical question-answer pairs and the question in the user query question-answer pairs output by the first large model after analyzing the plurality of instruction documents and user query texts according to the instruction in the second prompt word; Inputting the plurality of instruction documents, the historical question-answer pairs, the questions in the user query question-answer pairs, and the first prompt word into the first large model, obtaining the standard answer and the standard step sequence of the question in the user query question-answer pair, which are inferred and output by the first large model after analyzing the plurality of instruction documents and the historical question-answer pairs according to the instruction in the first prompt word; The evaluation dataset is constructed based on the historical question-answer pairs, user query question-answer pairs, standard step sequences, and multiple instruction documents.

[0006] According to a question-answering model evaluation method provided by the present invention, obtaining the user query text includes: Get all the query contents entered by users in the user query pool; Inputting all the query contents and the third prompt word into the second large model, obtaining the user query text filtered from all the query contents after the second large model analyzes all the query contents according to the instruction of the third prompt word; The third prompt word is used to instruct the second largest model to filter out query content containing the user's task completion requirements from all query content as the user query text.

[0007] According to a question-answering model evaluation method provided by the present invention, after constructing the evaluation dataset based on the historical question-answer pairs, user query question-answer pairs, standard step sequences, and multiple instruction documents, the method further includes: Revisions of historical question-answer pairs, standard answers in user query-answer pairs, and / or standard step sequences are received from external input.

[0008] According to a question-answering model evaluation method provided by the present invention, based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained, including: Calculating a recall-oriented summary evaluation metric based on the ground truth answer and the predicted answer; Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, the relevance score and task completion rate representing the answer of the question to be tested by the question-answering model are calculated; wherein, the relevance score represents the degree of relevance between the predicted answer and the question in the user query question-answer pair, and the task completion rate represents the ratio of the number of steps that appear in both the predicted step sequence and the standard step sequence to the total number of steps in the standard step sequence.

[0009] According to a question-answering model evaluation method provided by the present invention, based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, a relevance score and a task completion rate representing the answer of the question to be tested by the question-answering model are calculated, including: The historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence, the predicted step sequence and the fourth prompt word are input into the third largest model, and the third largest model outputs a correlation score and a task completion rate after analyzing the historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence and the predicted step sequence according to the instructions in the fourth prompt word.

[0010] The present invention also provides a question-answering model evaluation device, comprising the following modules: An answer output unit is used to input historical question-answer pairs, multiple instruction documents, questions in user query question-answer pairs, and a first prompt word in an evaluation data set into a question-answer model to be tested, and obtain a predicted answer and a predicted step sequence for the question in the user query question-answer pair inferred and output by the question-answer model to be tested according to the first prompt word, wherein the first prompt word is used to instruct the question-answer model to be tested to analyze multiple instruction documents and the historical question-answer pairs to infer and output the predicted answer and the predicted step sequence for the question in the user query question-answer pair; the evaluation data set includes: historical question-answer pairs, user query question-answer pairs, a standard step sequence corresponding to the standard answer in the user query question-answer pair, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query text, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answer and the standard step sequence are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs; The answer comparison unit is used to obtain an evaluation result based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence.

[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, it implements the question-answering model evaluation method as described in any one of the above.

[0012] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the question-answering model evaluation method as described in any one of the above.

[0013] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the question-answering model evaluation methods described above.

[0014] In the question-answering model evaluation method and device provided by the present invention, the evaluation data set includes multiple instruction documents corresponding to the user query text, and the first large model is used to infer the questions in the historical question-answer pairs and the user query question-answer pairs based on the multiple instruction documents, simulating a multi-step (or multi-round) question-answering scenario. The question-answering model to be tested, under the instruction of the first prompt word, infers the predicted answer and predicted step sequence of the question in the user query question-answering pair based on the analysis of the multi-step (or multi-round) question-answering scenario and multiple instruction documents, and then compares the predicted answer and the predicted step sequence with the standard answer and the standard step sequence respectively, thereby achieving effective evaluation of the output answer of the question-answering model to be tested in the scenario of user question-answering requirements for multi-step, process-based instructions. Moreover, the question-answering format of the multi-step (or multi-round) question-answering scenario and multiple instruction documents is closer to the user's usage habits, supports multi-round contextual reasoning and continuous context interaction, and enhances the evaluation of the conversational ability of the question-answering model to be tested. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 It is a flowchart of the question-answering model evaluation method provided by the present invention.

[0017] Figure 2 It is a schematic diagram of the evaluation data in the question-answering model evaluation method provided by the present invention.

[0018] Figure 3 This is a schematic diagram of the first prompt word in the question-answering model evaluation method provided by the present invention.

[0019] Figure 4 This is a schematic diagram of the second prompt word in the question-answering model evaluation method provided by the present invention.

[0020] Figure 5This is a schematic diagram of the application field distribution of the evaluation data set in the question-answering model evaluation method provided by the present invention.

[0021] Figure 6 This is a schematic diagram of the third prompt word in the question-answering model evaluation method provided by the present invention.

[0022] Figure 7 This is a schematic diagram of the fourth prompt word in the question-answering model evaluation method provided by the present invention.

[0023] Figure 8 It is a structural diagram of the question-answering model evaluation device provided by the present invention.

[0024] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] The question-answering model evaluation method of the embodiment of the present invention is as follows: Figure 1 As shown, it includes: step S110 and step S120.

[0027] Step S110: Input the historical question-and-answer pairs, multiple instruction documents, questions in the user query question-and-answer pairs, and the first prompt word in the evaluation data set into the question-and-answer model to be tested, and obtain the predicted answers and prediction step sequences of the questions in the user query question-and-answer pairs inferred and output by the question-and-answer model to be tested according to the first prompt word. The first prompt word is used to instruct the question-and-answer model to be tested to analyze the multiple instruction documents and the historical question-and-answer pairs to infer and output the predicted answers and prediction step sequences of the questions in the user query question-and-answer pairs.

[0028] The evaluation dataset is a pre-constructed sample set, specifically including: historical question-and-answer pairs, user query-and-answer pairs, standard step sequences corresponding to standard answers in the user query-and-answer pairs, and multiple instruction documents. The multiple instruction documents are retrieved based on user query text, the questions in the historical question-and-answer pairs and the user query-and-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-and-answer pairs, the multiple instruction documents, and the questions in the user query-and-answer pairs.

[0029] It's understandable that a user query text is a question posed by a user. After a user asks a question online, other users post answers based on the question. These posts serve as instruction documents, instructing the user on how to proceed. Different users may provide different answers, so a single user query text typically corresponds to multiple instruction documents.

[0030] The questions in historical question-and-answer pairs and user query question-and-answer pairs are generated by the first largest model based on multiple instruction documents, that is, the first largest model is used to analyze the contents of multiple instruction documents to generate multiple historical question-and-answer pairs and user query question-and-answer pairs. The historical question-and-answer pairs are question-and-answer pairs inferred by the first largest model based on multiple instruction documents, which are before the user query question-and-answer pairs and have a logical relationship of continuous questions and answers with the user query question-and-answer pairs. In this way, the multi-step or multi-round question-and-answer scenarios before user query questions and answers in daily real task scenarios are simulated through historical question-and-answer pairs.

[0031] The first model generates standard answers and standard step sequences based on the historical Q&A pairs, multiple instruction documents, and the questions in the user's query. Specifically, the first model analyzes the historical Q&A pairs and multiple instruction documents, and answers the questions in the user's query to obtain standard answers. The first model further refines the standard answers into clear, structured, process-based instructions, generating a standard step sequence with no sublists, clear logical order, and complete semantics. The user can solve their question by following this standard step sequence.

[0032] For example: Figure 2 As shown, the user query text corresponds to four instruction documents. Question Q3 in the user query question-answer pair is the last question. Q1A1 and Q2A2 are both historical question-answer pairs output by the first largest model based on the reasoning of the four instruction documents. A3 is the standard answer to Q3 output by the first largest model based on the reasoning of the four instruction documents and Q1A1 and Q2A2. 1-4 below A3 is the standard step sequence obtained by further refining the standard answer by the first largest model.

[0033] In this step, when the question-answering model to be tested is evaluated, the historical question-answer pairs, multiple instruction documents, questions in the user query question-answer pairs, and the first prompt word in the evaluation data set are input into the question-answering model to be tested. For example, the first prompt word is Figure 3 The first prompt word is used to instruct the question-answering model to be tested to understand and analyze the multiple instruction documents and historical question-answer pairs, and to infer and output the predicted answer and predicted step sequence of the question in the user query question-answer pair.

[0034] Step S120: Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained. Specifically, the closer the predicted answer is to the standard answer, and the closer the predicted step sequence is to the standard step sequence, the better the answer provided by the tested question-answering model, i.e., the better the evaluation result. For example, the evaluation result can be evaluated using the traditional recall-oriented summary evaluation metric ROUGE.

[0035] It should be noted that: since the first model generates historical question-answer pairs (such as Q1A1~Q2A2), questions in user query question-answer pairs, and standard answers and standard step sequences for questions in user query question-answer pairs based on multiple instruction documents, the question-answer model to be tested generates predicted answers and predicted step sequences based on multiple instruction documents and historical question-answer pairs, that is, it refers to the historical question-answer pairs generated by the first model to generate predicted answers and predicted step sequences. Therefore, the question-answer model to be tested can be evaluated by comparing the generated predicted answers and predicted step sequences with the standard answers and standard step sequences respectively.

[0036] In the question-answering model evaluation method of this embodiment, the evaluation data set includes multiple instruction documents corresponding to the user query text, and the first large model is used to infer the questions in the historical question-answer pairs and the user query question-answer pairs based on the multiple instruction documents, simulating a multi-step (or multi-round) question-answering scenario. The question-answering model to be tested, under the instruction of the first prompt word, infers the predicted answer and predicted step sequence of the question in the user query question-answering pair based on the analysis of the multi-step (or multi-round) question-answering scenario and multiple instruction documents, and then compares the predicted answer and the predicted step sequence with the standard answer and the standard step sequence respectively, thereby achieving effective evaluation of the output answer of the question-answering model to be tested in the scenario where the user requires question-answering of multi-step, process-based instructions. Moreover, the question-answering format of the multi-step (or multi-round) question-answering scenario and multiple instruction documents is closer to the user's usage habits, supports multi-round contextual reasoning and continuous context interaction, and enhances the evaluation of the conversational ability of the question-answering model to be tested.

[0037] In some embodiments, before step S110, a step of constructing an evaluation dataset is further included, specifically including: Step 1: Obtain the user query text. The user query text is a natural language query entered by the user when asking a question on the Internet, and can be obtained by querying the user query pool of a certain question-and-answer website or question-and-answer app.

[0038] Step 2: Acquire multiple command documents retrieved based on the user query text, that is, retrieve multiple command documents replied by other users based on the user query text.

[0039] Step 3: Input the plurality of instruction documents, user query texts, and the second prompt word into the first large model, and obtain the output of the historical question-answer pairs and the question in the user query question-answer pair after the first large model analyzes the plurality of instruction documents and user query texts according to the instructions in the second prompt word. The question in the user query question-answer pair is the last question, for example Figure 2 Question Q3 in . For example, the second prompt word is Figure 4 As shown, the first large model, based on the second prompt word, understands and analyzes multiple instruction documents and the user query text, inferring the questions in multiple historical Q&A pairs and the user query Q&A pair. Because the question in the user query Q&A pair corresponds to the user query text and represents the current question, the first large model outputs the question in the user query Q&A pair at the end of the historical Q&A pairs.

[0040] Step 4: Input the multiple instruction documents, the historical question-answer pairs, the questions in the user query question-answer pairs and the first prompt word into the first large model, and obtain the standard answer and the standard step sequence of the question in the user query question-answer pair after the first large model analyzes the multiple instruction documents and the historical question-answer pairs according to the instructions in the first prompt word. The first prompt word is as follows: Figure 3 As shown, the predicted answer and prediction step sequence output by the question-answering model to be tested are the same, thus avoiding the impact of different output results due to different prompt words on the accuracy of the evaluation results.

[0041] Step 5: Construct the evaluation data set based on the historical question-answer pairs, user query question-answer pairs, standard step sequences and multiple instruction documents. For each user query text in step 1, there is a corresponding evaluation data, for example Figure 2 The evaluation data shown.

[0042] In this embodiment, the evaluation data set is constructed through the above steps. Figure 5 As shown, to evaluate question-answering models across multiple application domains, the evaluation dataset covers 13 real-world application domains and contains 13,959 sample question-answering evaluation data. On average, each question in a user-query question-answer pair corresponds to 3.11 rounds of questions (both historical and user-query question-answer pairs), references 4.09 instruction documents, and ultimately generates 6.5 structured process instructions, or standard step sequences. This demonstrates the high diversity, practicality, and process-oriented nature of this evaluation dataset, enabling effective evaluation of the output of question-answering models in scenarios where users require multi-step, process-based instructions.

[0043] In some embodiments, obtaining the user query text specifically includes: Get all the query content entered by users in the user query pool. For some question-and-answer websites or question-and-answer apps, the query content entered by each user will be stored in the user query pool. Therefore, you can get all the query content entered by users by accessing the user query pool.

[0044] All the query contents and the third prompt word are input into the second model, and the second model analyzes all the query contents according to the instruction of the third prompt word, and the user query text is filtered out from all the query contents, wherein the third prompt word is used to instruct the second model to filter out the query contents containing the user's task completion requirements from all the query contents as the user query text. That is, the second model is used to filter all the query contents to filter out the query contents containing the user's task completion requirements. For example, the third prompt word is as follows: Figure 6 As shown, for the query content containing the user's task completion requirements, the answer is yes, thereby filtering out the query content containing the user's task completion requirements.

[0045] In this embodiment, filtering query content that contains the user's task completion requirements can be understood as removing content that is highly time-dependent or subjective, retaining query content with clear task objectives, and ensuring the quality and directiveness of the constructed dataset.

[0046] In some embodiments, after constructing the evaluation data set based on the historical question-answer pairs, user query question-answer pairs, standard step sequences, and multiple instruction documents, the process further includes: receiving external input for revisions to the standard answers and / or standard step sequences in the historical question-answer pairs and user query question-answer pairs. Exemplarily, three manual annotators are introduced to review the samples, eliminating evaluation data with problems, such as question-answer swaps, mismatched question-answer content, marketing content, low-quality language generation, politically sensitive information, and incorrect step order in the standard step sequence. Furthermore, 10% of the evaluation data in the evaluation data set is manually reviewed, ultimately achieving an accuracy rate of over 95%.

[0047] In this embodiment, after constructing the evaluation data set based on the historical question and answer pairs, user query question and answer pairs, standard step sequences and multiple instruction documents, the standard answers and / or standard step sequences in the historical question and answer pairs and user query question and answer pairs in the evaluation data set are manually revised, further ensuring the matching degree of each question and answer pair in the evaluation data set and the correctness of the step order of the standard step sequence, thereby achieving more accurate and effective evaluation of the question and answer model to be tested.

[0048] In some embodiments, step S120 specifically includes: The recall-oriented summary evaluation index is calculated based on the standard answer and the predicted answer. The recall-oriented summary ROUGE is a traditional evaluation index. In this embodiment, the ROUGE score can be represented by the literal and semantic similarity between the standard answer and the predicted answer. , ROUGE's score The specific calculation formula is as follows: ; ; .

[0049] in, X Indicates the standard answer. m express X The character length of Y represents the predicted answer, n express Y The character length of β is a hyperparameter, Recall, represents the accuracy (Precision), Indicates solving two sentences X and Y The longer the longest common subsequence, the higher the similarity between the two.

[0050] Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, the relevance score and task completion rate representing the answer of the question to be tested by the question-answering model are calculated; wherein, the relevance score represents the degree of relevance between the predicted answer and the question in the user query question-answer pair, and the task completion rate represents the ratio of the number of steps that appear in both the predicted step sequence and the standard step sequence to the total number of steps in the standard step sequence.

[0051] For example, the predicted answer output by the question-answering model to be tested is compared with the standard answer. For example, by comparing the similarity between the two, the relevance between the answer output by the question-answering model to be tested and the question in the user query question-answer pair can be evaluated. The greater the similarity between the two, the greater the relevance, and the higher the accuracy of the predicted answer.

[0052] Task completion rate Task completion rate It can be calculated as follows: .

[0053] in, M c represents the number of steps that appear in both the predicted step sequence and the standard step sequence, K Indicates the total number of steps in a standard step sequence.

[0054] It should be noted that the number of steps in the predicted step sequence and the standard step sequence may be inconsistent. The number of steps that appear in both the predicted step sequence and the standard step sequence is based on the steps in the standard step sequence to ensure the task completion rate. Task completion rate The value of is not greater than 1. For example, if the total number of steps in the prediction step sequence is five, and the total number of steps in the standard step sequence is four, and the contents of two steps in the prediction step sequence are represented by one step in the standard step sequence, then the two steps in the prediction step sequence are counted as one step.

[0055] This example effectively evaluates the ROUGE metric, total rating, and task completion rate of the question-answering model under test by comparing standard answers with predicted answers, as well as standard step sequences with predicted step sequences, in a multi-step, process-based question-answering scenario. For each evaluation dimension, a higher score indicates better question-answering performance.

[0056] In some embodiments, based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, calculating the relevance score and task completion rate representing the answer of the question to be tested by the question-answering model includes: The historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence, the predicted step sequence and the fourth prompt word are input into the third largest model, and the third largest model outputs a correlation score and a task completion rate after analyzing the historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence and the predicted step sequence according to the instructions in the fourth prompt word.

[0057] For example, the fourth prompt word is as follows: Figure 7 As shown, in this embodiment, the third model is used to analyze and predict whether the answer solves the problem in the user's query question and answer based on the fourth prompt word combined with the historical question and answer pair, and the predicted answer is compared and analyzed with the standard answer, and finally a score of 1-10 is output. The third model also outputs a relevance score of 1-10 based on the fourth prompt word representing the completion rate of the above task. Task completion rate The calculation formula automatically calculates the task completion rate.

[0058] It should be noted that each evaluation data in the evaluation data set will have a relevance score and task completion rate. The final evaluation result of the question-answering model to be tested is the average of the relevance scores and the average of the task completion rates obtained for each evaluation data.

[0059] In this embodiment, a large model (the third largest model) is introduced to assist in the evaluation mechanism, so that the evaluation process is standardized and automated, which significantly improves the evaluation efficiency and the consistency of the evaluation results.

[0060] It can be understood that the first largest model, the second largest model and the third largest model in the above embodiments can all be implemented using currently mature large language models (such as DeepSeek R1 or Qwen3, etc.).

[0061] The question-answering model evaluation device provided by the present invention is described below. The question-answering model evaluation device described below and the question-answering model evaluation method described above can be referenced to each other.

[0062] The question-answering model evaluation device of this embodiment is as follows: Figure 8 Shown, including: The answer output unit 810 is used to input historical question and answer pairs, multiple instruction documents, questions in user query question and answer pairs, and a first prompt word in the evaluation data set into the question and answer model to be tested, and obtain the predicted answers and predicted step sequences of the questions in the user query question and answer pairs inferred and output by the question and answer model to be tested according to the first prompt word. The first prompt word is used to instruct the question and answer model to be tested to analyze the multiple instruction documents and the historical question and answer pairs to infer and output the predicted answers and predicted step sequences of the questions in the user query question and answer pairs; the evaluation data set includes: historical question and answer pairs, user query question and answer pairs, standard step sequences corresponding to standard answers in user query question and answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query text, the questions in the historical question and answer pairs and user query question and answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question and answer pairs, multiple instruction documents, and questions in the user query question and answer pairs.

[0063] The answer comparison unit 820 is used to obtain an evaluation result based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence.

[0064] In some embodiments, the question-answering model evaluation apparatus further includes: The user query text acquiring unit is configured to acquire the user query text.

[0065] The instruction document acquisition unit is used to acquire a plurality of the instruction documents retrieved according to the user query text.

[0066] The historical question-answer pair output unit is used to input multiple instruction documents, user query texts and a second prompt word into the first large model, and obtain the historical question-answer pairs and the questions in the user query question-answer pairs output by the first large model after analyzing the multiple instruction documents and user query texts according to the instructions in the second prompt word.

[0067] The standard answer output unit is used to input multiple instruction documents, the historical question and answer pairs, the questions in the user query question and answer pairs, and the first prompt word into the first large model, and obtain the standard answers and the standard step sequence of the questions in the user query question and answer pairs after the first large model analyzes the multiple instruction documents and the historical question and answer pairs according to the instructions in the first prompt word.

[0068] A data set construction unit is used to construct the evaluation data set based on the historical question-answer pairs, user query question-answer pairs, standard step sequences and multiple instruction documents.

[0069] In some embodiments, the user query text acquisition unit specifically includes: The query content acquisition unit is used to acquire all query contents input by the user in the user query pool.

[0070] The query content screening unit is used to input all the query contents and the third prompt word into the second large model, and obtain the user query text screened from all the query contents after the second large model analyzes all the query contents according to the instruction of the third prompt word.

[0071] The third prompt word is used to instruct the second largest model to filter out query content containing the user's task completion requirements from all query content as the user query text.

[0072] In some embodiments, the question-answer model evaluation device also includes: a revision receiving unit for receiving external input of revisions to the standard answers and / or standard step sequences in the historical question-answer pairs, user query question-answer pairs, after constructing the evaluation data set based on the historical question-answer pairs, user query question-answer pairs, standard step sequences and multiple instruction documents.

[0073] In some embodiments, the answer comparison unit 820 is specifically configured to: Calculating a recall-oriented summary evaluation metric based on the ground truth answer and the predicted answer; Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, the relevance score and task completion rate representing the answer of the question to be tested by the question-answering model are calculated; wherein, the relevance score represents the degree of relevance between the predicted answer and the question in the user query question-answer pair, and the task completion rate represents the ratio of the number of steps that appear in both the predicted step sequence and the standard step sequence to the total number of steps in the standard step sequence.

[0074] In some embodiments, the answer comparison unit 820 is specifically used to: input the historical question and answer pairs, the questions in the user query question and answer pairs, the standard answers, the predicted answers, the standard step sequences, the predicted step sequences and the fourth prompt word into a third model, and obtain the correlation score and task completion rate output by the third model after analyzing the historical question and answer pairs, the questions in the user query question and answer pairs, the standard answers, the predicted answers, the standard step sequences and the predicted step sequences according to the instructions in the fourth prompt word.

[0075] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute the question-answering model evaluation method, which includes: The historical question-answer pairs, multiple instruction documents, questions in the user query question-answer pairs, and the first prompt word in the evaluation data set are input into the question-answer model to be tested, and the predicted answers and predicted step sequences of the questions in the user query question-answer pairs are obtained by the question-answer model to be tested according to the inference output of the first prompt word. The first prompt word is used to instruct the question-answer model to be tested to analyze the multiple instruction documents and the historical question-answer pairs to infer and output the predicted answers and predicted step sequences of the questions in the user query question-answer pairs. The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query texts, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs.

[0076] Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained.

[0077] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0078] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the question-answering model evaluation method provided by the above methods, which includes: The historical question-answer pairs, multiple instruction documents, questions in the user query question-answer pairs, and the first prompt word in the evaluation data set are input into the question-answer model to be tested, and the predicted answers and predicted step sequences of the questions in the user query question-answer pairs are obtained by the question-answer model to be tested according to the inference output of the first prompt word. The first prompt word is used to instruct the question-answer model to be tested to analyze the multiple instruction documents and the historical question-answer pairs to infer and output the predicted answers and predicted step sequences of the questions in the user query question-answer pairs. The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query texts, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs.

[0079] Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained.

[0080] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the question-answering model evaluation method provided by the above methods, the method comprising: The historical question-answer pairs, multiple instruction documents, questions in the user query question-answer pairs, and the first prompt word in the evaluation data set are input into the question-answer model to be tested, and the predicted answers and predicted step sequences of the questions in the user query question-answer pairs are obtained by the question-answer model to be tested according to the inference output of the first prompt word. The first prompt word is used to instruct the question-answer model to be tested to analyze the multiple instruction documents and the historical question-answer pairs to infer and output the predicted answers and predicted step sequences of the questions in the user query question-answer pairs. The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query texts, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs.

[0081] Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained.

[0082] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0083] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A question-answering model evaluation method, characterized in that: include: Inputting historical question-and-answer pairs, multiple instruction documents, questions in user query question-and-answer pairs, and a first prompt word in the evaluation dataset into the question-and-answer model to be tested, and obtaining a predicted answer and a sequence of predicted steps for the question in the user query question-and-answer pair output by the question-and-answer model to be tested based on the first prompt word, wherein the first prompt word is used to instruct the question-and-answer model to analyze the multiple instruction documents and the historical question-and-answer pairs to infer and output a predicted answer and a sequence of predicted steps for the question in the user query question-and-answer pair; The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query text, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs; Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained.

2. The question-answering model evaluation method according to claim 1, characterized in that: Before inputting historical question-answer pairs, multiple instruction documents, questions in user query question-answer pairs, and a first prompt word in the evaluation dataset into the question-answer model to be tested, and obtaining a predicted answer and a sequence of prediction steps output by the question-answer model to be tested based on the first prompt word, the method further includes: Obtaining the user query text; Acquire a plurality of instruction documents retrieved according to the user query text; Inputting the plurality of instruction documents, user query texts, and a second prompt word into the first large model, obtaining the historical question-answer pairs and the question in the user query question-answer pairs output by the first large model after analyzing the plurality of instruction documents and user query texts according to the instruction in the second prompt word; Inputting the plurality of instruction documents, the historical question-answer pairs, the questions in the user query question-answer pairs, and the first prompt word into the first large model, obtaining the standard answer and the standard step sequence of the question in the user query question-answer pair, which are inferred and output by the first large model after analyzing the plurality of instruction documents and the historical question-answer pairs according to the instruction in the first prompt word; The evaluation dataset is constructed based on the historical question-answer pairs, user query question-answer pairs, standard step sequences, and multiple instruction documents.

3. The question-answering model evaluation method according to claim 2, characterized in that: Obtaining the user query text includes: Get all the query contents entered by users in the user query pool; Inputting all the query contents and the third prompt word into the second large model, obtaining the user query text filtered from all the query contents after the second large model analyzes all the query contents according to the instruction of the third prompt word; The third prompt word is used to instruct the second largest model to filter out query content containing the user's task completion requirements from all query content as the user query text.

4. The question-answering model evaluation method according to claim 2, characterized in that: After constructing the evaluation dataset based on the historical question-answer pairs, the user query question-answer pairs, the standard step sequence, and the plurality of instruction documents, the further step includes: Revisions of historical question-answer pairs, standard answers in user query-answer pairs, and / or standard step sequences are received from external input.

5. The question-answering model evaluation method according to claim 1, characterized in that: Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, an evaluation result is obtained, including: Calculating a recall-oriented summary evaluation metric based on the ground truth answer and the predicted answer; Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, the relevance score and task completion rate representing the answer of the question to be tested by the question-answering model are calculated; wherein, the relevance score represents the degree of relevance between the predicted answer and the question in the user query question-answer pair, and the task completion rate represents the ratio of the number of steps that appear in both the predicted step sequence and the standard step sequence to the total number of steps in the standard step sequence.

6. The question-answering model evaluation method according to claim 5, characterized in that: Based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence, a relevance score and a task completion rate representing the answer of the question to be tested by the question-answering model are calculated, including: The historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence, the predicted step sequence and the fourth prompt word are input into the third largest model, and the third largest model outputs a correlation score and a task completion rate after analyzing the historical question-and-answer pairs, the question in the user query question-and-answer pairs, the standard answer, the predicted answer, the standard step sequence and the predicted step sequence according to the instructions in the fourth prompt word.

7. A question-answering model evaluation device, characterized in that: include: An answer output unit is configured to input historical question-and-answer pairs, multiple instruction documents, questions in user query question-and-answer pairs, and a first prompt word into the question-and-answer model to be tested, and obtain a predicted answer and a sequence of predicted steps for the question in the user query question-and-answer pair as inferred and output by the question-and-answer model to be tested based on the first prompt word. The first prompt word is configured to instruct the question-and-answer model to analyze the multiple instruction documents and the historical question-and-answer pairs to infer and output a predicted answer and a sequence of predicted steps for the question in the user query question-and-answer pair. The evaluation data set includes: historical question-answer pairs, user query question-answer pairs, standard step sequences corresponding to standard answers in the user query question-answer pairs, and multiple instruction documents, wherein the multiple instruction documents are retrieved based on user query text, the questions in the historical question-answer pairs and the user query question-answer pairs are generated by the first large model based on the multiple instruction documents, and the standard answers and standard step sequences are generated by the first large model based on the historical question-answer pairs, the multiple instruction documents, and the questions in the user query question-answer pairs; The answer comparison unit is used to obtain an evaluation result based on the comparison between the standard answer and the predicted answer, and the comparison between the standard step sequence and the predicted step sequence.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the question-answering model evaluation method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the question-answering model evaluation method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the question-answering model evaluation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Large model evaluation method, device, equipment, system and program product

    CN120106210A

  • Apparatus and methods for the generation and improvement of efficiency data

    US20250225426A1