Quantitative evaluation system and method for context retention capability of multi-round dialogues of large language model
By building a dynamic scene generation and interference injection module, combined with memory accuracy, association depth and anti-interference evaluation, the subjectivity and high cost problems in the multi-round dialogue evaluation of large language models are solved, and the quantitative evaluation and optimization of context retention capabilities are achieved.
Patent Information
- Application Number
- CN202510859942.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for evaluating the multi-round dialogue capabilities of large language models suffer from strong subjectivity, single dimensions, and high costs, and are unable to effectively measure the model's ability to maintain context in long conversations.
A quantitative evaluation system for the context-preserving ability of large language models in multi-round conversations is constructed, including a dynamic scenario generation module, an interference injection module, a three-layer evaluation module, and a decay rate analysis module. By inserting controllable interference and using quantitative indicators to evaluate the model's memory decay, topic relevance, and anti-interference robustness, an interactive evaluation report is generated.
It enables objective evaluation of the contextual understanding capabilities of large language models, significantly reduces evaluation costs, provides precise optimization directions, and improves evaluation accuracy and reproducibility.
Smart Images

Figure CN120803920A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing evaluation, in particular to a system and method for quantitatively evaluating the context retention capability of a large language model in multi-turn dialogue. BACKGROUND
[0002] There are three major pain points in the current context evaluation of large language models: first, it is highly subjective and relies on manually designed test cases and subjective scoring; second, it is single-dimensional, only testing basic QA memory capabilities, ignoring topic relevance and anti-interference; third, it is costly, requiring professional personnel to construct complex dialogue scenarios.
[0003] Existing tools have obvious limitations.
[0004] For example, LangChain Bench is a tool designed to help benchmark various large language model (LLM) related tasks. Its test tasks are organized according to end-to-end use cases and highly dependent on the Langsmith platform. However, it only supports single-turn dialogue testing and cannot evaluate the context retention capability of large language models in long dialogue. In actual application scenarios, large language models often need to handle multi-turn, complex dialogue interactions, such as the scenario of intelligent customer service communicating with users for multiple turns to solve problems. LangChain Bench cannot fully measure the performance of large language models in such scenarios due to its lack of evaluation of long dialogue context retention capability.
[0005] MT-Bench, on the other hand, is a framework for evaluating the performance of large language models in multi-turn dialogue. It uses a structured approach to test the performance of large language models in multi-turn dialogue scenarios, designs realistic multi-domain multi-turn questions to evaluate the model's ability in context understanding, answer accuracy, and logical reasoning, and uses an artificial intelligence-based scoring assistant to evaluate the model's output. However, although MT-Bench supports multi-turn dialogue testing, it does not introduce a quantitative decay indicator to measure the persistence of large language model context memory. For example, in multi-turn dialogue, as the number of turns increases, the large language model's memory of previous information may gradually fade, and MT-Bench cannot accurately evaluate the impact of this memory weakening on model performance due to the lack of a quantitative decay indicator. SUMMARY
[0006] The present application provides a system and method for quantitatively evaluating the context retention capability of a large language model in multi-turn dialogue to address the lack of standards and subjective dependence in evaluating the multi-turn dialogue capability of large language models.
[0007] In the first aspect, the technical solution adopted by the system for quantitatively evaluating the context retention capability of a large language model in multi-turn dialogue to solve the above technical problems is as follows:
[0008] A quantitative evaluation system for the multi-turn dialogue context maintenance capability of a large language model, comprising:
[0009] A dynamic scenario generation module responsible for generating a multi-turn dialogue flow containing marked information points based on rule templates and a large language model (LLM) collaboration, providing standardized test materials for subsequent interference injection;
[0010] An interference injection module responsible for inserting controllable interference into the dialogue flow generated by the dynamic scenario generation component, and outputting the dialogue flow with interference as input data for the three-layer evaluation module;
[0011] A three-layer evaluation module including a memory accuracy layer, a correlation depth layer, and an interference resistance layer, which quantifies the basic memory decay, topic logical correlation, and interference resistance robustness of the large language model (LLM), respectively, and provides quantitative data support for the decay rate analysis module by calculating the information decay rate, decay index, and interference resistance coefficient;
[0012] An attenuation rate analysis module based on the information decay rate, decay index, and interference resistance coefficient output by the three-layer evaluation module, corresponding to generate three types of curves: memory decay, correlation decay, and anti-interference decay, to intuitively present the performance trend of the large language model (LLM) in different dimensions with the increase of dialogue turns or the introduction of interference, providing intuitive data analysis support for large language model (LLM) optimization;
[0013] A visualization report module that integrates the curves generated by the decay rate analysis module, the quantitative results output by the three-layer evaluation module, and the original dialogue flow generated by the dynamic scenario generation module, outputs an interactive evaluation report, and clearly shows the overall performance of the large language model (LLM) in the context evaluation.
[0014] Optionally, the interference injection module involved based on the marked information points in the dialogue flow generated by the dynamic scenario generation module, injects three types of interference: topic jump, irrelevant QA pairs, and long text noise at specified turns or nodes, simulating complex scenarios in real interactions, and the output dialogue flow with interference is directly used as input data for the three-layer evaluation module.
[0015] Further optionally, the memory accuracy layer, correlation depth layer, and interference resistance layer of the three-layer evaluation module take the dialogue flow with interference output by the interference injection module as input, wherein:
[0016] The memory accuracy layer focuses on key information points in the dialogue flow, detects the memory decay of the large language model (LLM) on historical information through verification questions, calculates the information decay rate δ, δ = (1 - correct recall information points / total information points) x 100%;
[0017] The association depth layer focuses on the logical dependency relationship of the dialogue flow, calculates the semantic similarity between the model output and the standard answer through a semantic similarity algorithm, and introduces a decay factor a combined with the dialogue flow round to calculate the semantic similarity decay index Score; Score = cos(E (standard answer), E (large language model output)) * a, a = 1 - current round / total round;
[0018] The interference resistance layer directly tests the robustness of the model under three types of interference: topic jump, irrelevant QA pairs, and long text noise. The anti-interference ability is quantified by comparing the accuracy before and after interference, and the anti-interference coefficient κ is calculated, κ = accuracy after interference / accuracy before interference.
[0019] Further optionally, the decay rate analysis module generates a memory decay curve based on the information decay rate output by the three-layer evaluation module, the specific process being:
[0020] First, obtain the dialogue information point data output by the memory accuracy layer in the three-layer evaluation module, that is, insert N marked information points in each round of dialogue through the dynamic scene generation module, and insert verification questions after every K rounds to count the number of correctly recalled information points M by the large language model;
[0021] Then calculate the information decay rate δ of each round according to the formula "δ = (1 - correctly recalled information points / total information points) × 100%";
[0022] Finally, draw a memory decay curve with rounds as the horizontal axis and information decay rate δ as the vertical axis to visually present the trend of the large language model's ability to maintain historical key information points decreasing with rounds in multi-round dialogue.
[0023] Further optionally, the decay rate analysis module generates an association decay curve based on the decay index output by the three-layer evaluation module, the specific process being:
[0024] First, obtain the dialogue logical dependency data output by the association depth layer in the three-layer evaluation module, that is, construct a dialogue chain with logical dependency and insert questions that need to be associated with the previous context in subsequent dialogue;
[0025] Then use a semantic similarity calculation model to calculate the semantic similarity between the large language model output and the standard answer;
[0026] Then calculate the decay index Score through the formula "Score = cos(E (standard answer), E (large language model output)) * a, a = 1 - current round / total round";
[0027] Finally, draw an association decay curve with rounds as the horizontal axis and semantic similarity decay index Score as the vertical axis to visually present the trend of the large language model's ability to maintain logical association and the decline of semantic understanding ability in the process of topic evolution.
[0028] Further optionally, the involved decay rate analysis module generates an anti-interference decay curve based on the anti-interference coefficient output by the three-layer evaluation module, and the specific process is:
[0029] First, obtain the dialogue data after the interference injection module inserts the three types of interference, namely topic jump, irrelevant QA pair, and long text noise, in the dialogue flow;
[0030] Then record the correct rate of the large language model before and after the interference, calculate the anti-interference coefficient K through K = correct rate after interference / correct rate before interference, and further obtain the anti-interference decay rate η according to η = 1-K;
[0031] Finally, take the round as the horizontal axis and the anti-interference decay rate η as the vertical axis to draw the anti-interference decay curve, so as to present the change trend of the context retention ability and robustness of the large language model under controllable interference.
[0032] In the second aspect, the technical solution adopted to solve the above technical problems is as follows:
[0033] A quantitative evaluation method for the context retention ability of a large language model in multi-round dialogue, which comprises the following steps:
[0034] S1, dialogue flow generation stage: based on the cooperation of rule templates and large language model LLM, generate multi-round dialogue flow containing marked information points, and provide standardized test materials for subsequent interference injection;
[0035] S2, interference injection stage: insert controllable interference in the generated dialogue flow, and output the dialogue flow with interference as the input data of the quantitative evaluation stage;
[0036] S3, quantitative evaluation stage: through the memory accuracy layer, the correlation depth layer and the interference resistance layer, the basic memory decay, the topic logical correlation degree and the anti-interference robustness of the large language model LLM are quantified respectively, and the information decay rate, decay index and anti-interference coefficient are calculated to provide quantitative data support for the decay rate analysis stage;
[0037] S4, decay rate analysis stage: based on the calculated information decay rate, decay index and anti-interference coefficient, three types of curves of memory decay, correlation decay and anti-interference decay are generated correspondingly, so as to intuitively present the performance change trend of the large language model LLM in different dimensions with the increase of dialogue rounds or the introduction of interference, and provide intuitive data analysis support for the optimization of the large language model LLM;
[0038] S5, evaluation report generation stage: integrate the curve generated in the decay rate analysis stage, the quantitative results output in the quantitative evaluation stage, and the original dialogue flow generated in the dialogue flow generation stage, output an interactive evaluation report to clearly show the overall performance of the large language model LLM in the context evaluation.
[0039] Optionally, in the interference injection stage, based on the marked information points in the dialogue flow generated in the dialogue flow generation stage, three types of interference, topic jump, irrelevant QA pairs, and long text noise, are injected at specified turns or nodes to simulate complex scenarios in real interactions. The output dialogue flow with interference is directly used as input data for the quantitative evaluation stage.
[0040] Further optionally, in the quantitative evaluation stage, the memory accuracy layer, the correlation depth layer, and the interference resistance layer take the dialogue flow with interference output by the interference injection module as input, wherein:
[0041] The memory accuracy layer focuses on key information points in the dialogue flow, detects the memory decay of the large language model LLM through verification questions, calculates the information decay rate δ, δ = (1 - correct recall information points / total information points) x 100%;
[0042] The correlation depth layer pays attention to the logical dependency relationship of the dialogue flow, calculates the semantic matching degree of the model output and the standard answer through a semantic similarity algorithm, and introduces a decay factor α combined with the dialogue flow turns to calculate the semantic similarity decay index Score; Score = cos(E(standard answer), E(large language model output)) * α, α = 1 - current turn / total turns;
[0043] The interference resistance layer directly tests the robustness of the model under three types of interference: topic jump, irrelevant QA pairs, and long text noise, quantifies the anti-interference ability by comparing the accuracy before and after interference, and calculates the anti-interference coefficient κ, κ = accuracy after interference / accuracy before interference.
[0044] Further optionally, in the decay rate analysis stage, based on the information decay rate output by the quantitative evaluation stage, a memory decay curve is generated, the specific process is as follows:
[0045] (4.1.1) First, obtain the dialogue information point data output by the memory accuracy layer in the three-layer evaluation module, that is, insert N marked information points in each round of dialogue through the dynamic scene generation module, and insert verification questions after K rounds to count the number of information points M correctly recalled by the large language model;
[0046] (4.1.2) Then calculate the information decay rate δ of each round according to the formula "δ = (1 - correct recall information points / total information points) x 100%";
[0047] (4.1.3) Finally, we plot a memory decay curve with rounds as the horizontal axis and the information decay rate δ as the vertical axis to visually demonstrate how the large language model's ability to retain historical key information points decreases with each round of conversation.
[0048] In the decay rate analysis phase, the associated decay curve is generated based on the decay index output in the quantitative evaluation phase. The specific process is as follows:
[0049] (4.2.1) First, obtain the dialogue logic dependency data output by the associated depth layer in the three-layer evaluation module. This means constructing a dialogue chain with logical dependencies and inserting questions that require the previous context in the subsequent dialogue.
[0050] (4.2.2) Then, use the semantic similarity calculation model to calculate the semantic similarity between the large language model output and the standard answer;
[0051] (4.2.3) Calculate the decay index Score using the formula "Score = cos(E(standard answer), E(large language model output)) * α, α = 1 - current round / total rounds";
[0052] (4.2.4) Finally, a correlation decay curve is plotted with rounds as the horizontal axis and the semantic similarity decay index Score as the vertical axis. This curve visually demonstrates the decline in the large language model's ability to maintain logical relevance and semantic understanding during topic evolution.
[0053] In the attenuation rate analysis phase, the anti-interference attenuation curve is generated based on the anti-interference coefficient output in the quantitative evaluation phase. The specific process is as follows:
[0054] (4.3.1) First, obtain the conversation data after the interference injection module inserts three types of interference into the conversation flow: topic jump, irrelevant QA pairs, and long text noise;
[0055] (4.3.2) Then, record the accuracy of the large language model before and after the interference, calculate the anti-interference coefficient κ using "κ = accuracy after interference / accuracy before interference", and further calculate the anti-interference attenuation rate η using "η = 1-κ";
[0056] (4.3.3) Finally, an anti-interference attenuation curve is generated with rounds as the horizontal axis and the anti-interference attenuation rate η as the vertical axis to show the changing trend of the context retention ability and robustness of the large language model under controllable interference.
[0057] The system and method for quantitatively evaluating the context-preserving capability of a large language model in multiple rounds of dialogues of the present invention have the following advantages compared with the prior art:
[0058] 1. This paper constructs a three-tiered evaluation system based on memory accuracy, association depth, and interference resistance, and combines quantitative indicators such as information decay rate, semantic similarity decay index, and interference resistance coefficient to achieve an objective evaluation of the model's contextual understanding ability. By providing an automated quantitative evaluation process, it addresses the issues of missing standards and subjective dependence in multi-round dialogue ability evaluation, providing data support for the optimization of large language models.
[0059] 2. This invention establishes resistance testing based on three types of controllable interference: topic jumping, irrelevant QA, and long text noise. This can systematically expose the context loss defects of large language models in real scenarios and provide precise guidance for large language model optimization. It also adopts a hybrid architecture based on the collaboration of rule templates and large language models (LLMs), ensuring the diversity of test cases while maintaining the reproducibility of evaluation results, significantly reducing evaluation costs.
[0060] 3. This invention can be applied to large language model selection, parameter tuning, and dialogue system optimization, significantly reducing evaluation costs and improving evaluation accuracy, solving the problems of traditional evaluation methods that are highly subjective, single-dimensional, and high-cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Attachment Figure 1 This is a block diagram of system module connections in accordance with the first embodiment of the present invention;
[0062] Attachment Figure 2 This is a flow chart of the method of embodiment 2 of the present invention. DETAILED DESCRIPTION
[0063] In order to make the technical solution, the technical problems solved and the technical effects of the present invention more clear, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.
[0064] Example 1:
[0065] Reference Attachment Figure 1 This embodiment proposes a quantitative evaluation system for the multi-round dialogue context preservation capability of a large language model, which includes:
[0066] The dynamic scene generation module is responsible for generating multi-round dialogue flows containing marked information points based on the collaboration of rule templates and the large language model (LLM), providing standardized test materials for subsequent interference injection;
[0067] The interference injection module is responsible for inserting controllable interference into the dialogue flow generated by the dynamic scene generation component. The output dialogue flow with interference is directly used as input data for the three-layer evaluation module;
[0068] The three-layer evaluation module includes a memory accuracy layer, a correlation depth layer, and an interference resistance layer, respectively quantifying the basic memory decay, topic logical correlation, and interference resistance robustness of the large language model LLM, and providing quantitative data support for the decay rate analysis module by calculating the information decay rate, decay index, and interference resistance coefficient;
[0069] The decay rate analysis module generates three types of curves corresponding to memory decay, correlation decay, and anti-interference decay based on the information decay rate, decay index, and interference resistance coefficient output by the three-layer evaluation module, to intuitively present the performance change trend of the large language model LLM in different dimensions after the dialogue round increases or interference is introduced, and to provide intuitive data analysis support for the optimization of the large language model LLM.
[0070] The visualization report module integrates the curves generated by the decay rate analysis module, the quantitative results output by the three-layer evaluation module, and the original dialogue flow generated by the dynamic scene generation module, and outputs an interactive evaluation report (in HTML format, etc.), which clearly shows the overall performance of the large language model LLM in the context evaluation.
[0071] In this embodiment, the interference injection module involved based on the dialogue flow generated by the dynamic scene generation module, injects three types of interference, topic jump, irrelevant QA pair, and long text noise at specified rounds or nodes, simulates complex scenes in real interactions, and the output dialogue flow with interference is directly used as input data for the three-layer evaluation module.
[0072] In this embodiment, the memory accuracy layer, the correlation depth layer, and the interference resistance layer of the three-layer evaluation module take the dialogue flow with interference output by the interference injection module as input, wherein:
[0073] The memory accuracy layer focuses on key information points in the dialogue flow, detects the memory decay of the large language model LLM on historical information through verification questions, calculates the information decay rate δ, δ = (1 - correct recall information point / total information point) x 100%;
[0074] The correlation depth layer pays attention to the logical dependency relationship of the dialogue flow, calculates the semantic matching degree of the model output and the standard answer through a semantic similarity algorithm, and introduces a decay factor α combined with the dialogue flow round, to calculate the semantic similarity decay index Score; Score = cos(E(standard answer), E(large language model output)) * α, α = 1 - current round / total round;
[0075] The interference resistance layer directly tests the robustness of the model under three types of interference, topic jump, irrelevant QA pair, and long text noise, quantifies the anti-interference ability by comparing the accuracy before and after the interference, and calculates the anti-interference coefficient κ, κ = post-interference accuracy / pre-interference accuracy.
[0076] In this embodiment, the decay rate analysis module generates a memory decay curve based on the information decay rate output by the three-layer evaluation module. The specific process is as follows:
[0077] First, obtain the dialogue information point data output by the memory accuracy layer in the three-layer evaluation module, that is, insert N marked information points in each round of dialogue through the dynamic scene generation module, and insert verification questions after every K rounds to count the number of correctly recalled information points M by the large language model;
[0078] Then, calculate the information decay rate δ of each round according to the formula "δ = (1 - correctly recalled information points / total information points) x 100%";
[0079] Finally, draw a memory decay curve with the round as the horizontal axis and the information decay rate δ as the vertical axis to intuitively present the trend of the large language model's ability to maintain historical key information points decreasing with the round.
[0080] Taking the customer service scenario as an example, the user discusses the order status query with the customer service large language model (specifically ChatGPT).
[0081] Information point 1: The user's order number is "SP20250606001" (3rd round of dialogue).
[0082] Information point 2: The user's address is "XX City, XX District, 518, XX Road" (5th round of dialogue).
[0083] Information point 3: The order is expected to be delivered on "June 10, 2025" (7th round of dialogue).
[0084] Test process:
[0085] a) Insert verification questions at the 3rd, 5th, and 7th rounds (e.g. "What is my order number?" "Please confirm my delivery address" "When will the order be delivered?").
[0086] b) The large language model needs to correctly recall the above information points when the user asks questions related to the order in subsequent rounds (e.g. 10th, 15th).
[0087] c) Count the number of correctly recalled information points (M) and the total number of information points (N = 3) by the large language model ChatGPT in each round, and calculate the information decay rate.
[0088] Test results:
[0089] Round 10: M = 2 (missing order number), which means that the large language model ChatGPT only correctly recalled the user's address and the estimated delivery time of the order in the 10th round of conversation, forgetting the order number information, and the information decay rate = (1-2 / 3) x 100% = 33.3%.
[0090] Round 15: M = 1 (missing address and delivery time), that is, the large language model ChatGPT only remembers one information point in the 15th round of conversation, and the information decay rate = (1-1 / 3) x 100% = 66.7%.
[0091] Perform the above process in different scenarios to generate different memory decay curves. By observing the characteristics of all memory decay curves, it can be concluded that the memory capacity of large language models represented by ChatGPT decreases significantly with the increase of conversation rounds. It can be considered to optimize the context storage mechanism to change this defect of large language models represented by ChatGPT.
[0092] In this embodiment, the decay rate analysis module generates an associated decay curve based on the decay index output by the three-layer evaluation module. The specific process is as follows:
[0093] First, obtain the dialogue logic dependency data output by the associated depth layer in the three-layer evaluation module, that is, construct a dialogue chain with logical dependency and insert questions that need to be associated with the previous context in subsequent conversations;
[0094] Next, use a semantic similarity calculation model to calculate the semantic similarity between the large language model output and the standard answer;
[0095] Then calculate the decay index Score by the formula "Score = cos(E(standard answer), E(large language model output)) * a, a = 1-current round / total rounds";
[0096] Finally, take the round as the horizontal axis and the semantic similarity decay index Score as the vertical axis to generate an associated decay curve, which directly shows the ability of large language models to maintain logical association and the downward trend of semantic understanding ability in the process of topic evolution.
[0097] Taking the medical consultation scenario as an example, the user discusses symptom analysis and treatment suggestions with a medical large language model (specifically taking MedGPT as an example). The dynamic scenario generation module constructs the following logically dependent dialogue chain:
[0098] Symptom description (2nd round): "The patient has been continuously fever for 3 days, accompanied by cough."
[0099] Cause analysis (5th round): "It may be a viral cold or bacterial infection."
[0100] Solution (Round 8): "Suggest a blood test, and if the white blood cell count is high, use antibiotics."
[0101] Assume a total of 15 rounds of dialogue.
[0102] Test procedure:
[0103] a) Insert a verification question at Round 12: "Based on the previous symptoms and analysis, how should it be handled?"
[0104] b) Large language model MedGPT needs to generate semantically coherent answers in combination with the previous information, and use a semantic similarity algorithm to calculate the similarity between the large language model MedGPT output and the standard answer.
[0105] Test results:
[0106] Large language model output: "Suggest a blood test, and decide whether to use antibiotics based on the results."
[0107] Standard answer: "Suggest a blood test, and if the white blood cell count is high, use antibiotics."
[0108] Semantic similarity = 0.82, semantic similarity decay index = 0.82 x (1-12 / 15) = 0.82 x 0.2 = 0.164.
[0109] In different scenarios, the above procedure is performed to generate different correlation decay curves. By observing the characteristics of all correlation decay curves, it can be concluded that the long-distance semantic correlation ability of large language models represented by MedGPT decreases with the increase of dialogue rounds. It can be considered to enhance the memory mechanism of context logical reasoning chain to improve the correlation ability of large language models represented by MedGPT for logically dependent information in multi-round dialogue.
[0110] In this embodiment, the interference attenuation rate analysis module generates an anti-interference decay curve based on the anti-interference coefficient output by the three-layer evaluation module. The specific process is as follows:
[0111] First, obtain the dialogue data after the interference injection module inserts three types of interference in the dialogue flow: topic jump, irrelevant QA pairs, and long text noise.
[0112] Then record the accuracy of the large language model before and after the interference, calculate the anti-interference coefficient κ by "κ = accuracy after interference / accuracy before interference", and further obtain the anti-interference decay rate η by "η = 1-κ".
[0113] Finally, take the round as the horizontal axis and the anti-interference decay rate η as the vertical axis to draw the anti-interference decay curve, to present the change trend of the context retention ability and robustness of the large language model under controllable interference.
[0114] Take the government Q&A scenario as an example, the user discusses policy consultation with the government Q&A large language model (specifically, take Huawei Cloud Panggu government large model as an example). The interference injection module inserts the following interference in the dialogue flow:
[0115] Topic jump (round 6): "What is the weather like in XX city today?"
[0116] Irrelevant QA pair (round 8): "How to handle a passport?" → "Identity card and photo are required."
[0117] Long text noise (round 10): "Here is a lengthy policy interpretation about tax reform..."
[0118] Test procedure:
[0119] a) Before interference (round 5): The user asks "What is the personal income tax threshold?" Huawei Cloud Panggu government large model correctly answers "5000 yuan / month". This round judges that the large language model answers correctly, and the correct rate before interference is 100%.
[0120] b) After interference (round 12): The user asks again "What is the personal income tax threshold?" Huawei Cloud Panggu government large model needs to exclude interference information and give the correct response.
[0121] c) Calculate the anti-interference coefficient = correct rate after interference / correct rate before interference.
[0122] Test results:
[0123] Correct rate before interference = 100% (large language model correctly answers).
[0124] Correct rate after interference = 60% (large language model incorrectly answers "3500 yuan / month").
[0125] Anti-interference coefficient = 60% / 100% = 0.6, anti-interference attenuation rate = 1-0.6 = 40%.
[0126] Perform the above procedure in different scenarios, such as setting different types of topic jumps, irrelevant QA pairs, and long text noise interference, and conduct multiple rounds of testing for different policy types such as social security policy consultation and public accumulation fund withdrawal policy consultation. Draw different anti-interference attenuation curves. By observing the characteristics of all anti-interference attenuation curves, we can conclude that large language models represented by Huawei Cloud Panggu government large model lack stability in interference scenarios. We can consider optimizing the context filtering mechanism, such as using more accurate semantic recognition technology to filter out interference information in advance, to avoid its impact on the judgment of the large language model, thereby improving the reliability of the large language model in complex dialogue environments.
[0127] Example two:
[0128] Reference Attachment Figure 2 This embodiment proposes a quantitative evaluation method for the multi-round dialogue context preservation capability of a large language model, which includes the following steps:
[0129] S1. Dialogue flow generation stage: Based on the collaboration of rule templates and large language models (LLM), multi-round dialogue flows containing marked information points are generated to provide standardized test materials for subsequent interference injection.
[0130] S2, Interference Injection Phase: Based on the generated marked information points in the dialogue flow in the dialogue flow generation phase, three types of interference are injected into the specified rounds or nodes: topic jumps, irrelevant QA pairs, and long text noise. This simulates the complex scenarios in real interactions. The output of the interference dialogue flow is directly used as input data for the quantitative evaluation phase.
[0131] S3. Quantitative evaluation stage: Through the memory accuracy layer, association depth layer and interference resistance layer, the basic memory decay, topic logical relevance and anti-interference robustness of the large language model (LLM) are quantified respectively. By calculating the information decay rate, decay index and anti-interference coefficient, quantitative data support is provided for the decay rate analysis stage.
[0132] Specifically, the memory accuracy layer focuses on key information points in the conversation flow, detects the memory decay of the large language model (LLM) on historical information through verification questions, and calculates the information decay rate δ, δ = (1-correctly recalled information points / total information points) × 100%;
[0133] The association depth layer focuses on the logical dependencies of the dialogue flow. It uses a semantic similarity algorithm to calculate the semantic match between the model output and the standard answer. It also introduces a decay factor α based on the dialogue flow turns to calculate the semantic similarity decay index Score; Score = cos(E(standard answer), E(large language model output)) * α, where α = 1-current turn / total turns.
[0134] The interference resistance layer directly tests the robustness of the model under three types of interference: topic jumping, irrelevant QA pairs, and long text noise. The anti-interference ability is quantified by comparing the accuracy before and after interference, and the anti-interference coefficient κ is calculated, κ = accuracy after interference / accuracy before interference.
[0135] S4, decay rate analysis stage: Based on the calculated information decay rate, decay index, and anti-interference coefficient, three types of curves are generated: memory decay, association decay, and anti-interference decay. This visually presents the performance change trends of the large language model (LLM) in different dimensions as the number of conversation turns increases or interference is introduced, providing intuitive data analysis support for large language model (LLM) optimization.
[0136] (4.1) In the decay rate analysis stage, the memory decay curve is generated based on the information decay rate corresponding to the output of the quantitative evaluation stage. The specific process is as follows:
[0137] (4.1.1) First, obtain the dialogue information point data output by the memory precision layer in the three-layer evaluation module, that is, insert N marked information points in each round of dialogue through the dynamic scene generation module, and insert verification questions after K rounds to count the number of information points M correctly recalled by the large language model;
[0138] (4.1.2) Then calculate the information decay rate δ of each round according to the formula "δ = (1 - correct recall information point / total information point) x 100%";
[0139] (4.1.3) Finally, draw a memory decay curve with round as horizontal axis and information decay rate δ as vertical axis to intuitively present the trend of the large language model's ability to maintain historical key information points decreasing with rounds in multi-round dialogue.
[0140] Taking the customer service scenario as an example, the user discusses the order status query with the customer service large language model (specifically ChatGPT).
[0141] Information point 1: User order number "SP20250606001" (3rd round of dialogue).
[0142] Information point 2: User address "XX City, XX District, 518, XX Road" (5th round of dialogue).
[0143] Information point 3: Order estimated delivery time "June 10, 2025" (7th round of dialogue).
[0144] Test process:
[0145] a) Insert verification questions at the 3rd, 5th, and 7th rounds (e.g. "What is my order number?" "Please confirm my delivery address" "When will the order be delivered?").
[0146] b) The large language model needs to correctly recall the above information points when the user asks questions related to the order in subsequent rounds (e.g. 10th, 15th).
[0147] c) Count the number of correct recalls (M) and total information points (N = 3) for each round of the large language model ChatGPT, and calculate the information decay rate.
[0148] Test results:
[0149] Round 10: M = 2 (missing order number), which means that the large language model ChatGPT only correctly recalled the user's address and the estimated delivery time of the order in the 10th round of conversation, forgetting the order number information, and the information decay rate = (1-2 / 3) x 100% = 33.3%.
[0150] Round 15: M = 1 (missing address and delivery time), that is, the large language model ChatGPT only remembers one information point in the 15th round of conversation, and the information decay rate = (1-1 / 3) x 100% = 66.7%.
[0151] Perform the above process in different scenarios to generate different memory decay curves. By observing the characteristics of all memory decay curves, it can be concluded that the memory capacity of large language models represented by ChatGPT decreases significantly with the increase of conversation rounds. It can be considered to optimize the context storage mechanism to change this defect of large language models represented by ChatGPT.
[0152] (4.2) In the decay rate analysis stage, the decay index output by the quantitative evaluation stage is used to generate the correlation decay curve. The specific process is as follows:
[0153] (4.2.1) First, obtain the dialogue logic dependency data output by the correlation depth layer in the three-layer evaluation module, that is, construct a dialogue chain with logical dependency and insert questions that need to be associated with the previous context in subsequent conversations;
[0154] (4.2.2) Then use a semantic similarity calculation model to calculate the semantic similarity between the large language model output and the standard answer;
[0155] (4.2.3) Then calculate the decay index Score through the formula "Score = cos(E(standard answer), E(large language model output)) * a, a = 1-current round / total rounds";
[0156] (4.2.4) Finally, take the round as the horizontal axis and the semantic similarity decay index Score as the vertical axis to draw the correlation decay curve, which directly shows the ability of large language models to maintain logical correlation and the decline trend of semantic understanding ability in the process of topic evolution.
[0157] Taking the medical consultation scenario as an example, the user discusses symptom analysis and treatment recommendations with a medical large language model (specifically taking MedGPT as an example). The dynamic scenario generation module constructs the following logical dependency dialogue chain:
[0158] Symptom description (2nd round): "The patient has been low fever for 3 days, accompanied by cough."
[0159] Cause analysis (5th round): "It may be viral or bacterial infection."
[0160] Solution (Round 8): "Suggest a blood routine test, and if the white blood cell count is elevated, use antibiotics."
[0161] Assume a total of 15 rounds of dialogue.
[0162] Test procedure:
[0163] a) Insert a verification question at Round 12: "Based on the previous symptoms and analysis, how should it be handled?"
[0164] b) Large language model MedGPT needs to combine the previous information to generate a semantically coherent answer, and use a semantic similarity algorithm to calculate the similarity between the large language model MedGPT output and the standard answer.
[0165] Test results:
[0166] Large language model output: "Suggest checking blood routine, decide whether to use antibiotics according to the results."
[0167] Standard answer: "Suggest a blood routine test, and if the white blood cell count is elevated, use antibiotics."
[0168] Semantic similarity = 0.82, semantic similarity decay index = 0.82 x (1-12 / 15) = 0.82 x 0.2 = 0.164.
[0169] In different scenarios, execute the above procedure to generate different correlation decay curves, observe the characteristics of all correlation decay curves, and conclude that the long-distance semantic correlation ability of large language models represented by MedGPT decreases with the increase of dialogue rounds. It can be considered to enhance the memory mechanism of context logical reasoning chain to improve the correlation ability of large language models represented by MedGPT for logical dependent information in multi-round dialogue.
[0170] (4.3) In the decay rate analysis stage, generate an anti-interference decay curve based on the anti-interference coefficient output in the quantitative evaluation stage. The specific process is as follows:
[0171] (4.3.1) First, obtain the dialogue data after the interference injection module inserts three types of interference in the dialogue flow: topic jump, irrelevant QA pairs, and long text noise;
[0172] (4.3.2) Then record the accuracy of the large language model before and after the interference, calculate the anti-interference coefficient κ by "κ = accuracy after interference / accuracy before interference", and further obtain the anti-interference decay rate η by "η = 1-κ";
[0173] (4.3.3)Finally, the anti-interference attenuation curve is drawn with the round as the horizontal axis and the anti-interference attenuation rate η as the vertical axis, to present the change trend of the contextual retention ability and robustness of the large language model under controllable interference.
[0174] Taking the government Q&A scene as an example, the user discusses policy consultation with a government Q&A large language model (specifically, taking Huawei Cloud Pan Gu government large model as an example). The interference injection module inserts the following interference in the dialogue flow:
[0175] Topic jump (Round 6): “What is the weather like in XX city today?”
[0176] Irrelevant QA pair (Round 8): “How to apply for a passport?” → “Identity card and photo are required.”
[0177] Long text noise (Round 10): “The following is a lengthy policy interpretation of tax reform…”
[0178] Test procedure:
[0179] a) Before interference (Round 5): The user asks “What is the personal income tax threshold?” The Huawei Cloud Pan Gu government large model correctly answers “5000 yuan / month”. This round determines that the large language model answers correctly, and the correct rate before interference is 100%.
[0180] b) After interference (Round 12): The user asks again “What is the personal income tax threshold?” The Huawei Cloud Pan Gu government large model needs to exclude interference information and give the correct response.
[0181] c) Calculate the anti-interference coefficient = correct rate after interference / correct rate before interference.
[0182] Test results:
[0183] Correct rate before interference = 100% (large language model correctly answers).
[0184] Correct rate after interference = 60% (large language model incorrectly answers “3500 yuan / month”).
[0185] Anti-interference coefficient = 60% / 100% = 0.6, anti-interference attenuation rate = 1-0.6 = 40%.
[0186] The above process is performed in different scenarios, such as setting different types of topic jumps, irrelevant QA pairs, and long text noise interference, and multiple rounds of testing are performed for different policy types such as social security policy consultation and public accumulation withdrawal policy consultation, to draw different anti-interference attenuation curves. By observing the characteristics of all anti-interference attenuation curves, it can be concluded that large language models represented by Huawei Yunpogu government affairs large models have insufficient stability in interference scenarios, and the context filtering mechanism can be optimized, such as through more accurate semantic recognition technology, to filter out interference information in advance and avoid its impact on the judgment of large language models, thereby improving the reliability of large language models in complex conversation environments.
[0187] S5, evaluation report generation stage: integrating the curves generated in the attenuation rate analysis stage, the quantitative results output in the quantitative evaluation stage, and the original dialogue flow generated in the dialogue flow generation stage, outputting an interactive evaluation report (in HTML format, etc.), clearly showing the overall performance of the large language model LLM in the context evaluation.
[0188] As can be seen from the above, the large language model multi-round dialogue context retention capability quantitative evaluation system and method of the present application can realize objective evaluation of the model context understanding capability, solve the standard deficiency and subjective dependence problem in multi-round dialogue capability evaluation, and provide data support for large language model optimization.
[0189] The above application specific examples have described the principles and implementation modes of the present application in detail, and these examples are only used to help understand the core technical content of the present application. Based on the above specific embodiments of the present application, any improvement and modification of the present application made by those skilled in the art without departing from the principles of the present application shall fall within the scope of the patent protection of the present application.
Claims
1. A quantitative evaluation system for the multi-turn conversation context preservation capability of a large language model, characterized by: It includes: The dynamic scene generation module is responsible for generating multi-round dialogue flows containing marked information points based on the collaboration of rule templates and the large language model (LLM), providing standardized test materials for subsequent interference injection; The interference injection module is responsible for inserting controllable interference into the dialogue flow generated by the dynamic scene generation component. The output dialogue flow with interference is directly used as input data for the three-layer evaluation module; The three-layer evaluation module includes a memory accuracy layer, an association depth layer, and an interference resistance layer. These layers quantify the basic memory decay, topic logical relevance, and interference robustness of the large language model (LLM). By calculating the information decay rate, decay index, and interference resistance coefficient, they provide quantitative data support for the decay rate analysis module. The decay rate analysis module generates three types of curves: memory decay, association decay, and anti-interference decay, based on the information decay rate, decay index, and anti-interference coefficient output by the three-layer evaluation module. This module visually demonstrates the performance trends of the large language model (LLM) in different dimensions as the number of conversation turns increases or interference is introduced, providing intuitive data analysis support for large language model (LLM) optimization. The visualization report module integrates the curves generated by the decay rate analysis module, the quantitative results output by the three-layer evaluation module, and the original dialogue flow generated by the dynamic scene generation module to output an interactive evaluation report that clearly demonstrates the overall performance of the large language model (LLM) in contextual evaluation.
2. A quantitative evaluation system for the multi-round dialogue context preservation capability of a large language model according to claim 1, characterized in that: The interference injection module generates marked information points in the dialogue flow based on the dynamic scenario generation module, and injects three types of interference at specified turns or nodes: topic jumps, irrelevant Q&A pairs, and long text noise, to simulate complex scenarios in real interactions. The output of the interference-containing dialogue flow is directly used as input data for the three-layer evaluation module.
3. The quantitative evaluation system for multi-round dialogue context preservation capability of a large language model according to claim 2 is characterized in that: The memory accuracy layer, association depth layer, and interference resistance layer of the three-layer evaluation module take the interference-containing dialogue flow output by the interference injection module as input, where: The memory accuracy layer focuses on key information points in the conversation flow, detects the memory decay of the large language model (LLM) on historical information through verification questions, and calculates the information decay rate δ, where δ = (1-correctly recalled information points / total information points) × 100%; The association depth layer focuses on the logical dependencies of the dialogue flow. It uses a semantic similarity algorithm to calculate the semantic match between the model output and the standard answer. It also introduces a decay factor α based on the dialogue flow turns to calculate the semantic similarity decay index Score; Score = cos(E(standard answer), E(large language model output)) * α, where α = 1-current turn / total turns. The interference resistance layer directly tests the robustness of the model under three types of interference: topic jumping, irrelevant QA pairs, and long text noise. The anti-interference ability is quantified by comparing the accuracy before and after interference, and the anti-interference coefficient κ is calculated, κ = accuracy after interference / accuracy before interference.
4. The quantitative evaluation system for multi-round dialogue context preservation capability of a large language model according to claim 3 is characterized in that: The decay rate analysis module generates a memory decay curve based on the information decay rate output by the three-layer evaluation module. The specific process is as follows: First, we obtain the dialogue information point data output by the memory accuracy layer in the three-layer evaluation module. Specifically, we insert N labeled information points into each round of dialogue through the dynamic scene generation module. Then, we insert verification questions after K rounds to count the number of information points M correctly recalled by the large language model. Then, the information decay rate δ of each round is calculated according to the formula "δ = (1-correctly recalled information points / total information points) × 100%"; Finally, a memory decay curve is plotted with rounds as the horizontal axis and the information decay rate δ as the vertical axis to visually demonstrate the decreasing trend in the large language model's ability to retain historical key information points in multi-round conversations.
5. The quantitative evaluation system for multi-round dialogue context preservation capability of a large language model according to claim 3 is characterized in that: The decay rate analysis module generates a corresponding decay curve based on the decay index output by the three-layer evaluation module. The specific process is as follows: First, we obtain the dialogue logic dependency data output by the associated depth layer in the three-layer evaluation module. This means we construct a dialogue chain with logical dependencies and insert questions that require the previous context in the subsequent dialogue. Then, a semantic similarity calculation model is used to calculate the semantic similarity between the output of the large language model and the standard answer; The decay index Score is then calculated using the formula "Score = cos(E(standard answer), E(large language model output))*α, α = 1-current round / total rounds"; Finally, a correlation decay curve is generated with rounds as the horizontal axis and the semantic similarity decay index Score as the vertical axis, which intuitively shows the declining trend of the large language model's ability to maintain logical relevance and semantic understanding ability during the topic evolution process.
6. The quantitative evaluation system for multi-round dialogue context preservation capability of a large language model according to claim 3 is characterized in that: The attenuation rate analysis module generates an anti-interference attenuation curve based on the anti-interference coefficient output by the three-layer evaluation module. The specific process is as follows: First, obtain the conversation data after the interference injection module inserts three types of interference into the conversation flow: topic jump, irrelevant Q&A pairs, and long text noise; Then, the accuracy of the large language model before and after interference is recorded, and the anti-interference coefficient κ is calculated by "κ = accuracy after interference / accuracy before interference", and the anti-interference attenuation rate η is further calculated based on "η = 1-κ"; Finally, an anti-interference attenuation curve is generated with rounds as the horizontal axis and the anti-interference attenuation rate η as the vertical axis to show the changing trend of the context retention ability and robustness of the large language model under controllable interference.
7. A quantitative evaluation method for the multi-round dialogue context preservation capability of a large language model, characterized by: The steps include: S1, dialogue flow generation stage: Based on the collaboration of rule templates and large language model (LLM), multi-round dialogue flows with marked information points are generated to provide standardized test materials for subsequent interference injection; S2, interference injection stage: insert controllable interference into the generated dialogue flow, and the output interference dialogue flow is directly used as input data for the quantitative evaluation stage; S3, Quantitative Evaluation Stage: Through the memory accuracy layer, association depth layer, and interference resistance layer, the basic memory decay, topic logical relevance, and anti-interference robustness of the large language model (LLM) are quantified. By calculating the information decay rate, decay index, and anti-interference coefficient, quantitative data support is provided for the decay rate analysis stage. S4, decay rate analysis phase: Based on the calculated information decay rate, decay exponent, and anti-interference coefficient, three types of curves are generated: memory decay, association decay, and anti-interference decay. These curves visually demonstrate the performance trends of the large language model (LLM) in different dimensions as the number of conversation turns increases or interference is introduced, providing intuitive data analysis support for LLM optimization. S5. Evaluation report generation phase: This phase integrates the curves generated in the decay rate analysis phase, the quantitative results output in the quantitative evaluation phase, and the original dialogue flow generated in the dialogue flow generation phase to output an interactive evaluation report that clearly demonstrates the overall performance of the large language model (LLM) in contextual evaluation.
8. The quantitative evaluation method for the multi-round dialogue context preservation capability of a large language model according to claim 7 is characterized in that: In the interference injection phase, based on the marked information points in the dialogue flow generated in the dialogue flow generation phase, three types of interference are injected at specified turns or nodes: topic jumps, irrelevant QA pairs, and long text noise. This simulates the complex scenarios in real interactions. The output of the interference-containing dialogue flow is directly used as input data for the quantitative evaluation phase.
9. The quantitative evaluation method for the multi-round dialogue context preservation capability of a large language model according to claim 8 is characterized in that: In the quantitative evaluation phase, the memory accuracy layer, the association depth layer, and the interference resistance layer take the interference-injection module outputted dialogue stream as input, where: The memory accuracy layer focuses on key information points in the conversation flow, detects the memory decay of the large language model (LLM) on historical information through verification questions, and calculates the information decay rate δ, where δ = (1-correctly recalled information points / total information points) × 100%; The association depth layer focuses on the logical dependencies of the dialogue flow. It uses a semantic similarity algorithm to calculate the semantic match between the model output and the standard answer. It also introduces a decay factor α based on the dialogue flow turns to calculate the semantic similarity decay index Score; Score = cos(E(standard answer), E(large language model output)) * α, where α = 1-current turn / total turns. The interference resistance layer directly tests the robustness of the model under three types of interference: topic jumping, irrelevant QA pairs, and long text noise. The anti-interference ability is quantified by comparing the accuracy before and after interference, and the anti-interference coefficient κ is calculated, κ = accuracy after interference / accuracy before interference.
10. The quantitative evaluation method for the multi-round dialogue context preservation capability of a large language model according to claim 9 is characterized in that: In the decay rate analysis phase, a memory decay curve is generated based on the information decay rate output in the quantitative evaluation phase. The specific process is as follows: (4.1.1) First, obtain the dialogue information point data output by the memory accuracy layer in the three-layer evaluation module. That is, the dynamic scene generation module inserts N labeled information points in each round of dialogue. After an interval of K rounds, verification questions are inserted to count the number of information points correctly recalled by the large language model M. (4.1.2) Then calculate the information decay rate δ for each round using the formula "δ = (1 - correctly recalled information points / total information points) × 100%"; (4.1.3) Finally, we plot a memory decay curve with rounds as the horizontal axis and the information decay rate δ as the vertical axis to visually demonstrate how the large language model's ability to retain historical key information points decreases with each round of conversation. In the decay rate analysis phase, the associated decay curve is generated based on the decay index output in the quantitative evaluation phase. The specific process is as follows: (4.2.1) First, obtain the dialogue logic dependency data output by the associated depth layer in the three-layer evaluation module. This means constructing a dialogue chain with logical dependencies and inserting questions that require the previous context in the subsequent dialogue. (4.2.2) Then, use the semantic similarity calculation model to calculate the semantic similarity between the large language model output and the standard answer; (4.2.3) Calculate the decay index Score using the formula "Score = cos(E(standard answer), E(large language model output)) * α, α = 1 - current round / total rounds"; (4.2.4) Finally, a correlation decay curve is plotted with rounds as the horizontal axis and the semantic similarity decay index Score as the vertical axis. This curve visually demonstrates the decline in the large language model's ability to maintain logical relevance and semantic understanding during topic evolution. In the attenuation rate analysis phase, the anti-interference attenuation curve is generated based on the anti-interference coefficient output in the quantitative evaluation phase. The specific process is as follows: (4.3.1) First, obtain the conversation data after the interference injection module inserts three types of interference into the conversation flow: topic jump, irrelevant QA pairs, and long text noise; (4.3.2) Then, record the accuracy of the large language model before and after the interference, calculate the anti-interference coefficient κ using "κ = accuracy after interference / accuracy before interference", and further calculate the anti-interference attenuation rate η using "η = 1-κ"; (4.3.3) Finally, an anti-interference attenuation curve is generated with rounds as the horizontal axis and the anti-interference attenuation rate η as the vertical axis to show the changing trend of the context retention ability and robustness of the large language model under controllable interference.
Citation Information
Cited By
Assessment method for cross-modal input generation based on weak association rule and dynamic gradient
CN121412090A
Assessment and enhancement method and device for LLMs concept mutual exclusion recognition capability
CN121543698A
A large language model system prompt priority alignment benchmarking method
CN122432053A
A benchmarking method for prioritization alignment in a large language model system
CN122432053B