RAG answer quality evaluation method and system based on large language model

Through the RAG answer quality evaluation method based on the large language model, using weak length alignment and strong alignment techniques, combined with the results of multiple experiments, the bias problem of answer quality evaluation in the RAG system is solved, achieving a more scientific and accurate answer quality evaluation.

CN120523705APending Publication Date: 2025-08-22WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510491675.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing technology lacks stable answer quality evaluation criteria in search-enhanced generation (RAG) systems, resulting in model iteration direction deviation and decision-making errors in user interaction. In addition, there are length bias, position bias and experimental instability bias in large language models, making it difficult to accurately evaluate answer quality.

Method used

The answer length is adjusted by weak and strong alignment methods, and combined with preset evaluation indicators and multiple trial results, the RAG answer quality evaluation method based on large language model is adopted to reduce the impact of bias and increase the scientificity and accuracy of the evaluation.

Benefits of technology

It effectively reduces the biased impact caused by defects in the large language model itself, improves the scientificity, objectivity and accuracy of answer quality evaluation, solves the problems that are difficult to distinguish when the quality of traditional methods is close to the answer, and reduces the impact of instability through multiple experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523705A_ABST
    Figure CN120523705A_ABST
Patent Text Reader

Abstract

The invention provides an RAG answer quality evaluation method based on a large language model. The method comprises the following steps: inputting questions into two to-be-compared RAG systems to generate an answer pair corresponding to each question; comparing the lengths of the answer pairs, and if the length difference exceeds a preset threshold value, enabling the RAG system to which the short answers belong to generate answers again through an answer length weak alignment method; if the length difference of the regenerated answers still does not meet the preset threshold value, aligning the lengths of the shorter answers with the lengths of the longer answers through an answer length strong alignment method; scoring the answer pairs with the aligned lengths respectively; exchanging an answer sequence for re-scoring, and taking the sum of the two scores as a final score; and 5, comparing the answer scores of the two RAG systems to be compared, and drawing a box plot based on multiple test results. According to the method, the problems of length prejudice, position prejudice and test instability in the original scheme are solved by adopting the modes of answer alignment, position exchange, multiple tests and the like, and the evaluation problem of the RAG system capability is more accurately solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of answer quality assessment, and in particular to a RAG answer quality assessment method and system based on a large language model. Background Art

[0002] Answer quality evaluation (AQE) is dedicated to systematically evaluating the comprehensiveness, relevance, and reliability of generated answers to assess the quality of the answers and provide quantifiable feedback for optimizing large language models. In systems based on retrieval-augmented generation (RAG), there is a lack of stable answer quality evaluation standards for questions without correct answers, making it difficult to correctly evaluate the actual effectiveness of RAG systems. This situation can lead to multiple problems. For example, when optimizing model generation strategies, if factual errors and logical flaws in answers cannot be accurately distinguished, the direction of model iteration will deviate, making it difficult to effectively improve knowledge accuracy and reasoning capabilities. For example, in user interaction scenarios, if the system fails to filter out low-quality answers that contain false information or do not answer the question, it may lead to decision-making errors in key areas (such as medical diagnosis and financial analysis), seriously damaging user trust.

[0003] Existing methods often rely solely on the capabilities of the large language model itself, and directly use the large model to evaluate the answer quality from multiple angles. Such methods ignore the defects of the large language model itself, and its own defects will directly lead to length bias, position bias, and experimental instability bias in the process of evaluating answers. Specifically, length bias will cause the large language model to tend to believe that longer answers have better quality, position bias will cause the large language model to tend to believe that answers received earlier have better quality, and experimental instability bias will cause the evaluation results output by the large language model each time to be inconsistent. These biases will directly lead to errors in the evaluation results. Therefore, there is an urgent need for an answer quality assessment method that can effectively solve the above-mentioned biases and more accurately solve the problem of evaluating the capabilities of the RAG system. Summary of the Invention

[0004] The purpose of the present invention is to address the shortcomings of the existing technology and provide a RAG answer quality assessment method and system based on a large language model to evaluate the quality of answers generated by the RAG system, thereby scientifically revealing the true performance level of the RAG system in terms of answer generation quality.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A RAG answer quality assessment method based on a large language model includes the following steps:

[0007] Step 1: Input the questions in the test problem set into the two RAG systems to be compared, generate the answer pairs corresponding to each question, and record the length of each answer pair;

[0008] Step 2: Compare the lengths of the answer pairs for the same question. If the length difference exceeds a preset threshold, the RAG system to which the shorter answer belongs is regenerated using the weak alignment method. If the length difference of the regenerated answer meets the preset threshold, proceed to step 4; otherwise, proceed to step 3.

[0009] Step 3: If the length difference of the answers regenerated in step 2 still does not meet the preset threshold, the shorter answer is aligned with the longer answer through the strong answer length alignment method to form a length-aligned answer pair; if it still does not meet the threshold, the answer pair is discarded;

[0010] Step 4: Based on the preset evaluation metrics, the large language model scores the length-aligned answer pairs separately. The order of the answers is swapped and re-scored, and the sum of the two scores is used as the final score.

[0011] Step 5: Compare the answer scores of the two RAG systems to be compared, and count the number of wins and draws of the two RAG systems respectively;

[0012] Step 6: Repeat steps 4 and 5 multiple times and draw a box plot based on the results of multiple trials.

[0013] Furthermore, in step 1, the answers generated by the two RAG systems to be compared and their corresponding lengths are stored in the data structure form of: [RAG method; answer; answer length].

[0014] Furthermore, in step 2, the weak length alignment method includes: modifying the built-in prompt of the RAG system to which the shorter answer belongs, setting the expected answer length in the prompt word of the final generated answer, and setting the expected answer length to the answer length generated by the RAG model to which the longer answer belongs under the same question.

[0015] Furthermore, in step 2, the preset threshold is that the length difference between the two answers is ≤ 10 words.

[0016] Furthermore, the strong length alignment method in step 3 includes: using a large language model to add words without changing the meaning of the shorter answer, aligning the answer length again, recalculating the answer length, and checking whether its length meets the standard.

[0017] Furthermore, in step 3, the number of words added is the length difference between the answer after weak length alignment and the longer answer.

[0018] Furthermore, in step 4, the preset evaluation indicators are set as: comprehensiveness, relevance, empowerment and directness.

[0019] Furthermore, in step 4, the comprehensiveness is defined as the extent to which the answer covers all aspects and details of the question, the relevance is defined as the degree to which the answer is related to the question, the empowerment is defined as the extent to which the answer helps the questioner solve the problem, and the directness is defined as the clarity and readability of the answer.

[0020] Furthermore, in step 4, the scoring is specifically set as follows: each evaluation angle will be scored in the range of 0-5 points, where 0 points represents the worst performance and 5 points represents the best performance.

[0021] On the other hand, the present invention provides a RAG answer quality assessment system based on a large language model, comprising:

[0022] Initial answer generation module: It is used to input the questions in the test problem set into the two RAG systems to be compared, generate the answer pair corresponding to each question, and record the length of each answer pair;

[0023] Weak alignment module: It is used to compare the lengths of answer pairs for the same question. If the length difference exceeds a preset threshold, the RAG system to which the shorter answer belongs is regenerated through weak alignment of the answer length.

[0024] Strong alignment module: If the length difference of the answers regenerated in the weak alignment module still does not meet the preset threshold, the shorter answer is aligned with the longer answer through the strong alignment method to form a length-aligned answer pair; if it still does not meet the threshold, the answer pair is discarded;

[0025] Scoring module: This module is used to score length-aligned answer pairs using a large language model based on preset evaluation metrics. It also swaps the order of the answers and re-scores them, taking the sum of the two scores as the final score.

[0026] Quality assessment module: It is used to compare the answer scores of the two RAG systems to be compared, and count the number of wins and draws of the two RAG systems respectively;

[0027] Visualization module: It is used to repeat the steps in the above module multiple times and draw a box plot based on the results of multiple experiments.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] This invention utilizes a large language model-based RAG answer quality assessment method. This method reduces the impact of three types of biases on the assessment results due to inherent limitations of the large language model during the evaluation of the actual capabilities of different RAG methods, thereby enhancing the scientificity, objectivity, and accuracy of the assessment. The weak answer alignment step aims to minimize the assessment error caused by varying answer lengths while ensuring that the answers reflect the capabilities of the evaluated RAG model. The strong alignment step aims to address situations where the length difference is reduced after weak alignment but still falls short of the target. In such cases, since the distance from the target is small, directly using the large language model to add text allows for the best possible alignment without changing the meaning of the answer, thereby preventing the large language model from showing a preference for longer answers, which could affect the assessment accuracy. The scoring step in the answer quality assessment step makes the large language model assessment process more quantitative, resolving the difficulty of traditional methods in distinguishing between two answers of similar quality. It also makes the assessment process easier to understand and more scientific. The position swapping step aims to address the positional bias of the large language model, preventing the evaluation results from being affected by the placement of the answer. Repeating the test steps multiple times solves the instability problem in the large language model generation phase. Replacing the results of a single test with the statistical distribution of multiple tests avoids the inaccurate evaluation in certain test cases caused by unstable output. Therefore, the present invention increases the scientificity, objectivity, and accuracy of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1 This is a flow chart of a RAG answer quality assessment method based on a large language model proposed in the present invention.

[0032] Figure 2 It is the specific information of the answer evaluation indicator in step 4.

[0033] Figure 3 This is an example of how to grade the comprehensiveness of the answer evaluation criteria in step 4.

[0034] Figure 4 This is a specific flowchart of step 4.

[0035] Figure 5 This is an example plot of the boxplot in step 6. DETAILED DESCRIPTION

[0036] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0037] Example 1

[0038] See also Figure 1 The RAG answer quality assessment method based on a large language model proposed in the present invention includes the following steps.

[0039] Step 1: Enter the questions in the test set into the two RAG systems to be compared (System A and System B), generating two sets of answers for each question. Next, calculate the length of each set of answers and record this length information. Next, check the length difference between the two sets of answers for each question. If the length difference meets the requirements, proceed directly to Step 4; if not, proceed to Step 2.

[0040] The test set questions in this embodiment and the answers generated by system A and system B are shown in Table 1:

[0041] Table 1

[0042]

[0043]

[0044] Step 2: Weakly align answer lengths. For two answers to the same question generated by two different RAG methods, if one answer is shorter, regenerate the answer by adjusting the RAG method's built-in prompt to bring its length closer to the longer answer. After regenerating the answer, calculate its length again and check whether it meets the length requirements. If the length meets the requirements, proceed to Step 4; if it still does not meet the requirements, proceed to Step 3.

[0045] The answers generated by system A and system B after weak alignment in this embodiment are shown in Table 2:

[0046] Table 2

[0047]

[0048]

[0049] Step 3: Strongly align answer lengths. If, after step 2, the answer lengths of the questions are still not aligned, the large language model can be used directly. Without changing the original meaning of the shorter answer, the answer length is aligned with the longer answer by adding a few words. Then, the answer length is recalculated and checked to see if it meets the criteria. If so, proceed to step 4; if not, the answer is simply discarded.

[0050] The answers generated by System A and System B after strong alignment in this embodiment are shown in Table 3:

[0051] Table 3

[0052]

[0053]

[0054] Step 4: Generate answer quality scores. After Steps 2 and 3, answer pairs whose answer length difference meets the criteria are evaluated for answer quality. During this evaluation, the large language model is used to score the answers based on pre-defined answer evaluation metrics. After the order of the answers is evaluated, the order is swapped and the scoring is repeated. The scoring results for one question are shown in Table 4. Finally, the sum of the scores under the two orderings is used as the final score for the answer pair.

[0055] Table 4

[0056]

[0057] Step 5: Answer quality evaluation. Compare and analyze the generated answer scores. The specific rules are as follows:

[0058] If the answer score generated by method A is higher than the answer score generated by method B, the counter indicating method A's win is incremented by 1. If the answer score generated by method A is equal to the answer score generated by method B, the counter indicating the two RAG methods have tied is incremented by 1. If the answer score generated by method A is lower than the answer score generated by method B, the counter indicating method B's win is incremented by 1. After all answers to the entire test set have been judged, the number of wins by method A, the number of wins by method B, and the number of ties are recorded.

[0059] Step 6: Repeat the experiment and draw a boxplot. To minimize the impact of the large language model's inherent output instability on the results, repeat steps 4 and 5 multiple times. Then, based on the results of these repeated experiments, draw a boxplot, which serves as the basis for the final evaluation.

[0060] See also Figure 2The answer evaluation indicators in step 4 specifically include comprehensiveness, relevance, empowerment and directness.

[0061] Specifically, comprehensiveness measures how much detail the answer provides, covering all aspects and details of the question. Relevance measures whether the answer is relevant to the question. Empowerment measures the extent to which the answer helps readers understand and make informed decisions about the topic. Directness measures how specifically and clearly the answer addresses the question.

[0062] See also Figure 3 The comprehensiveness rating in step 4 is as follows:

[0063] A score of 0 indicates that the answer is very limited in scope and does not cover any important aspects or details of the question. A score of 1 indicates that the answer covers only a few aspects of the question, provides little detail, and omits important elements. A score of 2 indicates that the answer touches on some aspects of the question but lacks depth, missing several key details and broader context. A score of 3 indicates that the answer is moderately comprehensive, covering most aspects of the question in reasonable detail, although some areas could be expanded. A score of 4 indicates that the answer is very comprehensive, providing comprehensive coverage of the question and providing detailed explanations and context for most aspects. A score of 5 indicates that the answer is extremely comprehensive, thoroughly covering all aspects and details of the question, providing a complete and well-rounded understanding.

[0064] See also Figure 4 The specific process of the order exchange in step 4 is: first, score the answer pairs with the input order of [A, B]. After completing the scoring, exchange the order of A and B in the prompt words to [B, A], and score again in this order. After the two scorings, the sum of the two scores is taken as the actual score of the RAG method.

[0065] See also Figure 5 , the box plot drawn by the present invention after multiple repetitions is shown in the figure, where the vertical axis is the winning rate and the horizontal axis is A winning, drawing and B winning. The winning rate is defined as the proportion of the number of problems that a RAG model wins in the test problem set to the total number of problems. In the box plot, the upper and lower edges of the box part represent the first quartile (Q1, 25% quantile) and the third quartile (Q3, 75% quantile), respectively, and the length of the box is the interquartile range (IQR), which reflects the middle 50% distribution range of the data. The orange horizontal line in the middle of the box represents the median (50% quantile), which is the center position of the data and can intuitively reflect the central trend of the data. The whiskers extending above and below the box extend to the minimum and maximum values ​​within the range of no more than 1.5 times the IQR, respectively. Data points outside this range are regarded as outliers and are represented by separate dots.

[0066] Example 2

[0067] A specific embodiment of the present invention provides a RAG answer quality assessment system based on a large language model, comprising:

[0068] Initial answer generation module: It is used to input the questions in the test problem set into the two RAG systems to be compared, generate the answer pair corresponding to each question, and record the length of each answer pair;

[0069] Weak alignment module: It is used to compare the lengths of answer pairs for the same question. If the length difference exceeds a preset threshold, the RAG system to which the shorter answer belongs is regenerated through weak alignment of the answer length.

[0070] Strong alignment module: If the length difference of the answers regenerated in the weak alignment module still does not meet the preset threshold, the shorter answer is aligned with the longer answer through the strong alignment method to form a length-aligned answer pair; if it still does not meet the threshold, the answer pair is discarded;

[0071] Scoring module: This module is used to score length-aligned answer pairs using a large language model based on preset evaluation metrics. It also swaps the order of the answers and re-scores them, taking the sum of the two scores as the final score.

[0072] Quality assessment module: It is used to compare the answer scores of the two RAG systems to be compared, and count the number of wins and draws of the two RAG systems respectively;

[0073] Visualization module: It is used to repeat the steps in the above module multiple times and draw a box plot based on the results of multiple experiments.

[0074] The above is only a preferred specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be covered by the scope of protection of the present application.

[0075] It should be understood that any portion not elaborated in detail in this specification belongs to the prior art. It should be understood that the above description of the preferred embodiments is relatively detailed, but it should not be considered as limiting the scope of protection of the present invention. A person of ordinary skill in the art, guided by the present invention, may make substitutions or modifications without departing from the scope of protection of the claims of the present invention, and all such modifications shall fall within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the appended claims.

Claims

1. A RAG answer quality assessment method based on a large language model, characterized by: The following steps are involved: Step 1: Input the questions in the test problem set into the two RAG systems to be compared, generate the answer pairs corresponding to each question, and record the length of each answer pair; Step 2: Compare the lengths of the answer pairs for the same question. If the length difference exceeds a preset threshold, the RAG system to which the shorter answer belongs is regenerated using the weak alignment method. If the length difference of the regenerated answer meets the preset threshold, proceed to step 4; otherwise, proceed to step 3. Step 3: If the length difference of the answers regenerated in step 2 still does not meet the preset threshold, the shorter answer is aligned with the longer answer through the strong answer length alignment method to form a length-aligned answer pair; if it still does not meet the threshold, the answer pair is discarded; Step 4: Based on the preset evaluation metrics, the large language model scores the length-aligned answer pairs separately. The order of the answers is swapped and re-scored, and the sum of the two scores is used as the final score. Step 5: Compare the answer scores of the two RAG systems to be compared, and count the number of wins and draws of the two RAG systems respectively; Step 6: Repeat steps 4 and 5 multiple times and draw a box plot based on the results of multiple trials.

2. The RAG answer quality assessment method based on a large language model according to claim 1, characterized in that: In step 1, the answers generated by the two RAG systems to be compared and their corresponding lengths are stored in the data structure form of: [RAG method; answer; answer length].

3. The RAG answer quality assessment method based on a large language model according to claim 1, characterized in that: In step 2, the weak length alignment method includes: modifying the built-in prompt of the RAG system to which the shorter answer belongs, setting the expected answer length in the prompt word of the final generated answer, and setting the expected answer length to the answer length generated by the RAG model to which the longer answer belongs under the same question.

4. The RAG answer quality assessment method based on a large language model according to claim 3 is characterized by: In step 2, the preset threshold is that the length difference between the two answers is ≤ 10 words.

5. The RAG answer quality assessment method based on a large language model according to claim 1, characterized in that: The strong length alignment method in step 3 includes: using a large language model to add words without changing the meaning of the shorter answer, aligning the answer length again, recalculating the answer length, and checking whether its length meets the standard.

6. The RAG answer quality assessment method based on a large language model according to claim 5, characterized in that: In step 3, the number of words added is the length difference between the answer after weak length alignment and the longer answer.

7. The RAG answer quality assessment method based on a large language model according to claim 1, characterized in that: In step 4, the preset evaluation indicators are set as: comprehensiveness, relevance, empowerment and directness.

8. The RAG answer quality assessment method based on a large language model according to claim 7, characterized in that: In step 4, comprehensiveness is defined as the extent to which the answer covers all aspects and details of the question, relevance is defined as the degree to which the answer is related to the question, empowerment is defined as the extent to which the answer helps the questioner solve the problem, and directness is defined as the clarity and readability of the answer.

9. The RAG answer quality assessment method based on a large language model according to claim 1, characterized in that: In step 4, the scoring is specifically set as follows: each evaluation angle will be scored in the range of 0-5 points, where 0 points is the worst performance and 5 points is the best performance.

10. A RAG answer quality assessment system based on a large language model, characterized by: include: Initial answer generation module: It is used to input the questions in the test problem set into the two RAG systems to be compared, generate answer pairs corresponding to each question, and record the length of each answer pair; Weak alignment module: It is used to compare the lengths of answer pairs for the same question. If the length difference exceeds a preset threshold, the RAG system to which the shorter answer belongs is regenerated through weak alignment of the answer length. Strong alignment module: If the length difference of the answers regenerated in the weak alignment module still does not meet the preset threshold, the shorter answer is aligned with the longer answer through the strong alignment method to form a length-aligned answer pair; if it still does not meet the threshold, the answer pair is discarded; Scoring module: This module is used to score length-aligned answer pairs using a large language model based on preset evaluation metrics. It also swaps the order of the answers and re-scores them, taking the sum of the two scores as the final score. Quality assessment module: It is used to compare the answer scores of the two RAG systems to be compared, and count the number of wins and draws of the two RAG systems respectively; Visualization module: It is used to repeat the steps in the above module multiple times and draw a box plot based on the results of multiple experiments; The RAG answer quality assessment system based on a large language model is used to perform the steps in the RAG answer quality assessment method based on a large language model according to any one of claims 1 to 9.