Database-based multi-source seed fusion evaluation rag knowledge base test case self-evolution method
Patent Information
- Application Number
- CN202611304011.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]1.新文档、新版本入库后缺少可稳定复现的验证手段,测试用例难以及时覆盖新增知识点,低质量知识易直接进入生产使用范围
[0033]通过将静态种子问题、会话种子问题和准入种子问题三类来源统一融合,并经归一化去重、确定性采样构建统一测试输入集,再经证据分级标注、评测基准构建、质量与性能联合判定以及闭环回灌,实现静态回归测试、真实会话覆盖、新知识入库验证和检索性能看护的统一框架。使测试集能够随知识库内容和用户问题变化持续演进,同时通过证据驱动的评测基准隔离生成模型随机性对测试结论的影响,显著提升知识库测试的覆盖率和结论可靠性。
Smart Images

Figure CN122838291A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database management technology, and in particular to a self-evolution method for test cases of RAG knowledge base evaluation based on database multi-source seed fusion assessment. Background Technology
[0002] After receiving a user's question, a RAG (Retrieval-Augmented Generation) system first retrieves relevant documents or text slices from the knowledge base, and then provides the retrieval results as context to a large language model to generate an answer. Therefore, the output quality of a RAG system depends not only on the generation model itself, but also on multiple factors such as the quality of the knowledge base content, the granularity of document slicing, the indexing status, the recall strategy, candidate ranking, and the performance of the retrieval service. New documents or new versions of knowledge, once added to the database, only have practical application value if they can be correctly sliced, indexed, and recalled, and can provide sufficient and accurate evidence to support the user's question. A comprehensive testing framework is needed to verify the new knowledge throughout the entire process.
[0003] In practical applications, existing testing frameworks have the following shortcomings:
[0004] 1. New documents and versions lack reproducible verification methods after being added to the database, making it difficult for test cases to cover new knowledge points in a timely manner, and low-quality knowledge can easily be directly used in production. 2. Real-world conversations contain a large number of high-quality questions, but the lack of a systematic mechanism for capturing, filtering, and preserving these questions leads to a gradual disconnect between the test set and online user expressions. 3. Issues such as whether new documents can be correctly recalled, whether the segmentation is reasonable, and whether there is interference in the search ranking lack independent verification steps, making it impossible to confirm the retrieval usability of new knowledge. 4. Focusing only on whether questions are generated, without determining whether the questions belong to the target knowledge domain, whether they have stable answers, or whether they can form valid evidence, easily leads to the inclusion of low-quality or noisy issues in the regression scope. 5. Failure to link candidate fragment recall, evidence relevance labeling, baseline answer construction, quality scoring, admission determination, and asset backfeeding makes it difficult to pinpoint the root cause of problems when tests fail. 6. Dynamically generated test cases lack stable identification and version management, failing to form a reproducible, auditable, and rollbackable test asset management method.
[0005] In practical applications, changes in knowledge base content can easily lead to the following risks: newly added documents not being indexed, excessively long or short slices causing core facts to be fragmented, similar documents interfering with search ranking, lack of reproducible test cases for new knowledge, inability of recalled content to form core evidence, and increased tail latency in the search service. Without framework-based testing, these risks often only become apparent after users actually ask questions.
[0006] Furthermore, directly using large language models to evaluate the final generated answers has the following limitations: multiple evaluations of the same question may yield different conclusions, lacking certainty and unable to serve as an objective basis for admission; the "illusion" problem of the large model itself may mistakenly allow low-quality knowledge to pass; the evaluation process is a black box operation, and the root cause cannot be traced when anomalies occur; the randomness of the generation model and the retrieval quality are coupled and their respective influences cannot be isolated.
[0007] Chinese patents CN121144337A and CN121722885A disclose an intelligent evaluation method for RAG systems based on large models, respectively. CN121722885A discloses a method for RAG knowledge base evaluation and self-optimization based on three-dimensional dynamic calibration. CN119226753A discloses a large language model-driven method for evaluating the quality of RAG knowledge bases. All three methods rely on LLM-based offline, one-time generation of evaluation datasets and lack a continuous evolution mechanism. They cannot incorporate real user conversation expressions to keep the test set synchronized with online scenarios, nor do they possess the ability to drive dynamic backfeeding of test cases through new knowledge entry verification. Essentially, they are all "static evaluation" solutions, failing to address the closed-loop management problem of continuously iterating and evolving test assets as the knowledge base content changes.
[0008] Therefore, how to provide a knowledge base testing framework that can continuously evolve while maintaining determinism and traceability has become an urgent technical problem to be solved. Summary of the Invention
[0009] In view of this, in order to overcome the shortcomings of the prior art, the present invention aims to provide a database-based method for the self-evolution of test cases for multi-source seed fusion evaluation of RAG knowledge base.
[0010] This invention provides a database-based method for the self-evolution of test cases in a multi-source seed fusion evaluation RAG knowledge base, the method comprising:
[0011] Step S1: Obtain static seed questions, session seed questions, and admission seed questions from the static seed pool, real user sessions, and newly added or updated knowledge base documents within the target time window, respectively, and construct a unified test input set through normalization deduplication and deterministic sampling;
[0012] Step S2: Based on the unified test input set, call the RAG knowledge base retrieval interface to build a candidate text fragment pool, perform relevance classification and labeling on the candidate fragments, extract core evidence with the strong relevance threshold as the boundary, and generate the evaluation benchmark for this batch based on the core evidence of each test case;
[0013] Step S3: Perform retrieval quality evaluation based on the evaluation criteria, calculate the final admission score by weighting the retrieval quality indicators and retrieval time indicators, and output the admission conclusion based on the final admission score;
[0014] Step S4: Feed back the test cases that meet the preset conditions and originate from the session seed problem or admission seed problem, and are deduplicated, to the static seed pool.
[0015] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S1, the session seed questions are obtained as follows: real user session records are captured from a specified time window and decomposed into message sequences, and knowledge-type question statements directly raised by users are identified; rule filtering is performed on the identified knowledge-type question statements, which includes removing system prompt words, batch prompts, irrelevant chatter, and format noise; the filtered knowledge-type question statements are quality-scored using preset quality scoring rules, and a preset number of candidate questions are retained according to the scores; a semantic analysis big model is called to review the retained candidate questions to form a session seed question candidate set.
[0016] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S1, the admission seed questions are obtained in the following manner: newly created or newly updated documents in the knowledge base are filtered according to the target document date, and the document title, paragraph, text slice content, product version, knowledge module and key terms are extracted as the generation context. Natural language questions for testing the retrieval of new knowledge are generated according to the generation context, forming a candidate set of admission seed questions.
[0017] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S1, a unified test input set is constructed as follows: each seed question is labeled with its source type; the question text of each seed question is normalized, including at least removing leading and trailing whitespace characters, merging consecutive whitespace characters, and unifying letter case; the normalized question text is deduplicated, and when the same question appears in multiple sources simultaneously, a single record is retained according to a preset priority; a hash sampling key is calculated, which is the hash value obtained by concatenating a fixed seed value with the normalized question text; test cases are selected according to the hash sampling key, not exceeding the configured limit, to form a unified test input set.
[0018] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S2, a candidate text fragment pool is constructed as follows: for each test case in the unified test input set, the RAG knowledge base retrieval interface is called to obtain an internal candidate set, the number of which is greater than the final retention number; a preset number of final candidate sets with the highest ranking are retained from the internal candidate sets, and the candidate fragment identifier, document identifier, document name, product version, retrieval ranking, similarity, content summary and retrieval time of each candidate fragment in the final candidate set are recorded to form a candidate text fragment pool.
[0019] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, step S2 involves classifying and labeling candidate fragments based on relevance, and extracting core evidence with a strong relevance threshold as the boundary. This includes: using a large model to batch score each candidate text fragment, with relevance scores including at least four levels from 0 to 3, where a score of 0 indicates irrelevant to the question and is discarded, a score of 1 indicates weak relevance and does not enter any evidence set, a score of 2 indicates relevance and enters the relevant evidence set, and a score of 3 indicates strong relevance and enters the core evidence candidate set; setting a relevance threshold of 2 and a strong relevance threshold of 3, allowing only candidate fragments with relevance scores reaching the strong relevance threshold as core evidence; when there are no candidate fragments reaching the strong relevance threshold, the corresponding test case is marked as requiring manual review and is not included in the evaluation benchmark.
[0020] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S2, the evaluation benchmark for this batch is generated based on the core evidence of each test case in the following manner: For each test case with core evidence, a correspondence is established between three types of information: question information, evidence information, and answer reference information. The question information includes at least a question identifier, question text, module, and keywords. The evidence information includes at least a set of related fragments, a set of core fragments, and corresponding document identifiers. The answer reference information is a reference answer and factual points formed based on the core evidence. The question text, the set of related evidence identifiers, and the set of core evidence identifiers are concatenated into an original concatenated string. The hash value of the original concatenated string is calculated using a standard hash algorithm. The hash value is used as the benchmark summary value. The set of related evidence identifiers is a set of identifiers of all candidate fragments whose relevance scores reach the relevance threshold and a set of identifiers of all candidate fragments whose relevance scores reach the strong relevance threshold.
[0021] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S3, when performing retrieval quality evaluation according to the evaluation benchmark, the RAG knowledge base retrieval interface is called to obtain the returned context and determine the context precision and context recall for each test case in the evaluation benchmark. The context precision is the proportion of relevant content in the returned context, and the context recall is the proportion of core evidence in the evaluation benchmark that is successfully recalled.
[0022] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, step S3 involves weighting the retrieval quality index and the retrieval time index to calculate the final admission score, and outputting the admission conclusion based on the final admission score, specifically including:
[0023] Set the context precision score to context precision multiplied by 100, and the context recall score to context recall multiplied by 100;
[0024] The average time consumption score and P95 time consumption score are calculated using a linear decay function. When the actual time consumption is less than or equal to the preset excellent time consumption, the value is 100; when it is greater than or equal to the preset deteriorated time consumption, the value is 0. Otherwise, the score is calculated by multiplying 100 by the ratio of the difference between the deteriorated time consumption and the actual time consumption to the difference between the deteriorated time consumption and the excellent time consumption.
[0025] The final admission score is equal to the sum of the context precision score multiplied by 0.40, the context recall score multiplied by 0.40, the average execution time score multiplied by 0.10, and the P95 execution time score multiplied by 0.10.
[0026] When the final admission score is greater than or equal to the preset pass threshold, the pass conclusion is output; when it is greater than or equal to the preset conditional pass threshold but less than the preset pass threshold, the conditional pass conclusion is output; when it is less than the preset conditional pass threshold, the failure conclusion is output.
[0027] Optionally, in the database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method of the present invention, in step S4, the test cases are fed back to the static seed pool in the following manner:
[0028] Set backfeed admission criteria, which include at least the following: test cases have entered the evaluation benchmark, have valid problem identifiers, have been deduplicated by text and do not have the same normalized problem text in the static seed pool, and the admission conclusion for this batch is passed;
[0029] For test cases filtered by the re-feedback admission criteria, a sedimentation issue identifier is generated. This sedimentation issue identifier is a concatenation of the source prefix, document date, source type, and issue text hash value.
[0030] Before writing to the static seed pool, back up the original static seed file; after the backfilling, the test cases should at least record the source, the type of source to be merged, the problem identifier before sedimentation, the sedimentation batch directory and the sedimentation time;
[0031] Record the batch catalog, the number of new additions, and the identification of new issues in the daily sedimentation history.
[0032] The self-evolution method for test cases of RAG knowledge base fusion evaluation based on database multi-source seed fusion in this invention has the following beneficial technical effects:
[0033] By unifying and integrating three types of seed questions—static seed questions, session seed questions, and admission seed questions—and constructing a unified test input set through normalization deduplication and deterministic sampling, a unified framework for static regression testing, real-world session coverage, new knowledge entry verification, and retrieval performance monitoring is achieved through evidence-driven hierarchical annotation, benchmark construction, joint quality and performance evaluation, and closed-loop feedback. This allows the test set to continuously evolve with changes in knowledge base content and user questions. Furthermore, the evidence-driven benchmark isolates the impact of randomness in the generation model on test conclusions, significantly improving the coverage and reliability of knowledge base testing. Attached Figure Description
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart illustrating the self-evolution method of test cases for RAG knowledge base test case based on database multi-source seed fusion evaluation according to exemplary embodiment 1 of the present invention.
[0036] Figure 2 This is a flowchart illustrating the process of classifying and labeling the relevance of candidate text fragments and extracting core evidence according to the method of Exemplary Embodiment 1 of the present invention.
[0037] Figure 3 This is a schematic diagram illustrating the process of constructing evaluation benchmarks and calculating benchmark summary values according to the method of Exemplary Embodiment 1 of the present invention;
[0038] Figure 4 This is a flowchart illustrating the process of calculating a weighted admission score and outputting three levels of admission conclusions according to the method of Exemplary Embodiment 1 of the present invention.
[0039] Figure 5 This is a flowchart illustrating the process of determining and managing the test case reinjection admission conditions according to the method of Exemplary Embodiment 1 of the present invention. Detailed Implementation
[0040] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0041] It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other; and, based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0042] It should be noted that various aspects of the embodiments described below are within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0043] Example 1
[0044] Exemplary embodiment 1 of the present invention provides a database-based method for the self-evolution of test cases in a multi-source seed fusion evaluation RAG knowledge base. Figure 1 This is a flowchart illustrating the self-evolution method for RAG knowledge base test cases based on database-driven multi-source seed fusion evaluation according to Exemplary Embodiment 1 of the present invention. Figure 1 As shown, in this embodiment, the method of the present invention is implemented in the following manner:
[0045] Step S1: Obtain static seed questions, session seed questions, and admission seed questions from the static seed pool, real user sessions, and newly added or updated knowledge base documents within the target time window, respectively. Construct a unified test input set through normalization deduplication and deterministic sampling.
[0046] In this embodiment, the session seed questions are obtained as follows: real user session records are captured from a specified time window and broken down into message sequences; knowledge-related questions directly raised by users are identified; rule filtering is performed on the identified knowledge-related questions, which includes removing system prompts, batch prompts, irrelevant chatter, and format noise; the filtered knowledge-related questions are scored using preset quality scoring rules, and a preset number of candidate questions are retained according to the scores; a semantic analysis model is called to review the retained candidate questions to form a session seed question candidate set.
[0047] This embodiment employs a multi-level screening mechanism that involves message decomposition, knowledge-based question identification, rule filtering, quality scoring and ranking, and large-scale model verification of real user conversations. This mechanism can automatically extract high-quality knowledge-based questions from massive amounts of raw conversations, effectively filtering out interference information such as system prompts, batch prompts, irrelevant chatter, and format noise. This ensures that high-quality questions from real users are stably accumulated as test assets, keeping the test set synchronized with online user expressions and preventing the test set from becoming disconnected from real-world scenarios.
[0048] In this embodiment, admission seed questions are obtained as follows: newly created or updated documents in the knowledge base are filtered based on the target document date. The document title, paragraphs, text slice content, product version, knowledge module, and key terms are extracted as the generation context. Natural language questions for verifying the retrievability of newly added knowledge are generated based on this generation context, forming a candidate set of admission seed questions. This ensures that newly added knowledge receives corresponding test case coverage in a timely manner, providing a standardized source of questions for independent admission verification of new knowledge.
[0049] It should be noted that in this embodiment, the static seed problem is a long-term retained regression test asset used to cover stable knowledge points such as product documentation, configuration parameters, backup and recovery, migration compatibility, primary and standby high availability, SQL optimization, and fault diagnosis, serving as a baseline for quality comparison between batches.
[0050] In this embodiment, a unified test input set is constructed as follows: each seed question is labeled with its source type; the question text of each seed question is normalized, including at least removing leading and trailing whitespace characters, merging consecutive whitespace characters, and unifying letter case; the normalized question text is deduplicated, and when the same question appears in multiple sources simultaneously, a single record is retained according to a preset priority; a hash sampling key is calculated, which is the hash value obtained by concatenating a fixed seed value with the normalized question text; test cases are selected according to the hash sampling key, up to a maximum configured number, to form a unified test input set.
[0051] In practical applications, this embodiment uses a deterministic sampling method to form a unified test input set for this batch. Deterministic sampling is based on calculating a hash sampling key using a fixed seed value and normalized question text: Sampling key = Hash(fixed seed value + normalized question text).
[0052] This embodiment eliminates duplication among multi-source problems by normalizing and deduplicating three types of seed problems, pre-setting priority retention, and hash deterministic sampling, thus avoiding resource waste caused by repeated evaluation of the same problem. By sorting the hash sampling keys, it ensures that the same input produces completely consistent sampling results when executed in different batches, fundamentally eliminating the uncertainty brought about by random sampling, and making the evaluation results between batches strictly comparable and reproducible.
[0053] Step S2: Based on the unified test input set, call the RAG knowledge base retrieval interface to build a candidate text fragment pool, perform relevance classification and labeling on the candidate fragments, extract core evidence with the strong relevance threshold as the boundary, and generate the batch evaluation benchmark based on the core evidence of each test case.
[0054] The testing framework in this embodiment does not directly use the final answer generated by the large model as the sole criterion for judgment. Instead, it first performs relevance labeling and evidence screening on the search candidates to form relevant evidence, core evidence, and evaluation benchmarks. Then, it calculates the search quality index based on the evaluation benchmarks. By judging "whether it recalls effective knowledge" and "whether it can support the target answer" in a hierarchical manner, the impact of the randomness of the generation model on the knowledge base test conclusions is reduced.
[0055] In this embodiment, the candidate text fragment pool is constructed as follows: For each test case in the unified test input set, the RAG knowledge base retrieval interface is called to obtain an internal candidate set, the number of which is greater than the final retention number; a preset number of top-ranked final candidate sets are retained from the internal candidate set, which ensures the coverage of candidate evidence and avoids missing potentially relevant fragments, while also controlling the final number of candidates to avoid wasting computational resources due to excessive input in the subsequent annotation stage; the candidate fragment identifier, document identifier, document name, product version, retrieval ranking, similarity, content summary, and retrieval time of each candidate fragment in the final candidate set are recorded to form the candidate text fragment pool, providing a sufficient data foundation for subsequent evidence annotation, problem localization, and performance analysis.
[0056] Figure 2 This is a flowchart illustrating the process of relevance classification and labeling of candidate text fragments and extraction of core evidence according to the method of Exemplary Embodiment 1 of the present invention, as follows: Figure 2 As shown in this embodiment, a large model is used to score each candidate text fragment in batches. The relevance score includes at least four levels from 0 to 3, where a score of 0 indicates irrelevant to the question and is discarded, a score of 1 indicates weak relevance and is not included in any evidence set, a score of 2 indicates relevance and is included in the relevant evidence set, and a score of 3 indicates strong relevance and is included in the core evidence candidate set. A relevance threshold of 2 and a strong relevance threshold of 3 are set, allowing only candidate fragments with a relevance score reaching the strong relevance threshold to be used as core evidence. When no candidate fragment reaches the strong relevance threshold, the corresponding test case is marked as requiring manual review and is not included in the evaluation benchmark. If the large model does not explicitly select core evidence but multiple full-score candidates exist, a fallback selection is made based on confidence level and retrieval ranking.
[0057] This embodiment employs a large model to grade and score candidate segments, and sets relevance thresholds and strong relevance thresholds to strictly separate "weakly relevant" from "core evidence," allowing only segments that reach the strong relevance threshold as core evidence. For test cases without strong relevance candidates, they are automatically marked for manual review and excluded from the evaluation benchmark, thereby ensuring the rigor and credibility of the admission evaluation benchmark, effectively avoiding the misjudgment of weakly relevant search results as successful knowledge base retrieval, and significantly improving the reliability of test conclusions.
[0058] Figure 3 This is a schematic diagram illustrating the process of constructing evaluation benchmarks and calculating benchmark summary values according to the method of Exemplary Embodiment 1 of the present invention, as follows: Figure 3 As shown in this embodiment, for each test case with core evidence, a correspondence is established between three types of information: question information, evidence information, and answer reference information. Question information includes at least a question identifier, question text, the module it belongs to, and keywords. Evidence information includes at least a set of related fragments, a set of core fragments, and corresponding document identifiers. Answer reference information consists of reference answers and key facts based on the core evidence. The question text, the set of related evidence identifiers, and the set of core evidence identifiers are concatenated into an original concatenated string. The hash value of this original concatenated string is calculated using a standard hash algorithm. This hash value is used as a baseline summary value. This summary value is used to determine whether the evaluation benchmark has changed between different batches, avoiding misjudgments of changes in knowledge base quality when the evaluation benchmark itself changes. The set of related evidence identifiers consists of the identifiers of all candidate fragments whose relevance scores reach a relevance threshold, and the set of identifiers of all candidate fragments whose relevance scores reach a strong relevance threshold.
[0059] This embodiment establishes a correspondence between three types of information: question information, evidence information, and answer reference information, thereby constructing a structured evaluation benchmark for each test case. Furthermore, by concatenating the question text, the set of relevant evidence identifiers, and the set of core evidence identifiers and calculating a hash value as the benchmark summary value, a unique and stable digital fingerprint is provided for the evaluation benchmark. This allows for quick determination of whether the evaluation benchmark itself has changed between different batches by comparing the benchmark summary value. Thus, when quality indicators fluctuate, it can be accurately attributed to changes on the retrieval side or changes in the benchmark, providing a clear basis for judgment in problem localization.
[0060] Step S3: Perform retrieval quality evaluation based on the evaluation criteria, calculate the final admission score by weighting the retrieval quality index and the retrieval time index, and output the admission conclusion based on the final admission score.
[0061] In this embodiment, when performing retrieval quality evaluation based on the evaluation benchmark, the RAG knowledge base retrieval interface is called to obtain the returned context and determine the context precision and context recall for each test case in the benchmark. Context precision is the proportion of relevant content in the returned context, used to measure the degree of retrieval noise; context recall is the proportion of core evidence successfully recalled in the benchmark, used to measure the coverage of core evidence. Low context precision indicates a large number of irrelevant fragments in the retrieval results, while low context recall indicates that core evidence has not been recalled or is ranked too low. These two indicators are complementary, enabling a quantitative evaluation of retrieval quality from two orthogonal dimensions: the degree of retrieval noise and the coverage of core evidence. This provides a comprehensive and quantitative evaluation method for knowledge base retrieval quality, facilitating the rapid identification of specific problems in the retrieval process.
[0062] Figure 4 This is a schematic diagram illustrating the process of calculating a weighted admission score and outputting three levels of admission conclusions according to the method of Exemplary Embodiment 1 of the present invention, as follows: Figure 4 As shown, in this embodiment, the final admission score is calculated by weighting the retrieval quality index and the retrieval time index, and the admission conclusion is output based on the final admission score, specifically including:
[0063] Set the context precision score to context precision multiplied by 100, and the context recall score to context recall multiplied by 100;
[0064] The average time consumption score and P95 time consumption score are calculated using a linear decay function. When the actual time consumption is less than or equal to the preset excellent time consumption, the value is 100; when it is greater than or equal to the preset deteriorated time consumption, the value is 0. Otherwise, the score is calculated by multiplying 100 by the ratio of the difference between the deteriorated time consumption and the actual time consumption to the difference between the deteriorated time consumption and the excellent time consumption.
[0065] The final admission score is equal to the sum of the context precision score multiplied by 0.40, the context recall score multiplied by 0.40, the average execution time score multiplied by 0.10, and the P95 execution time score multiplied by 0.10.
[0066] When the final admission score is greater than or equal to the preset pass threshold, a pass conclusion is output, the test assets can be reinjected, and new knowledge is admitted; when the score is greater than or equal to the preset conditional pass threshold but less than the preset pass threshold, a conditional pass conclusion is output, and failed samples can be manually checked or retested after knowledge is supplemented; when the score is less than the preset conditional pass threshold, a failure conclusion is output, and the score is not included in the long-term regression range, triggering problem localization.
[0067] This embodiment calculates the final admission score by weighting the context precision score, context recall score, average latency score, and P95 latency score with weights of 0.40, 0.40, 0.10, and 0.10, respectively. It also sets three levels of conclusions: pass threshold and conditional pass threshold. Quality indicators are given the main weight to ensure the effectiveness of the retrieval as the primary goal, while latency indicators are used as constraints to prevent service degradation at the cost of high performance. The three-level conclusion division avoids simplistic black-and-white judgments and provides a buffer space for manual review in conditional pass scenarios, thereby improving the rationality and operability of the admission judgment.
[0068] In practical applications, the testing framework in this embodiment controls both retrieval quality and service performance. Regarding retrieval quality, metrics such as contextual precision, contextual recall, retrieval precision, retrieval recall, average reciprocal rank, or normalized depreciation cumulative gain can be measured. Regarding service performance, metrics such as average retrieval time, P95 / P99 tail latency, success rate, requests processed per second, and concurrent failed samples can also be measured.
[0069] Step S4: Feed back the test cases that meet the preset conditions and originate from the session seed problem or admission seed problem, and are deduplicated, to the static seed pool.
[0070] In this embodiment, the testing framework archives the input parameters, test case sources, candidate pool, annotation results, evaluation benchmarks, evaluation reports, admission scores, stress test results, and execution steps for each batch, and establishes the correlation between the artifacts through a run list. For high-value issues and stable evidence identified during the evaluation process, the framework preserves them according to their source, batch, and original identifier, forming a closed-loop evolution of test assets.
[0071] Figure 5 This is a flowchart illustrating the process of determining and managing the test case reinjection admission conditions according to the method of Exemplary Embodiment 1 of the present invention, as follows: Figure 5 As shown, in this embodiment, test cases are fed back to the static seed pool in the following manner:
[0072] Set up backfeed admission criteria, which include at least the following: test cases have entered the evaluation benchmark, have valid issue identifiers, have been deduplicated from text and do not have the same normalized issue text in the static seed pool, and the admission conclusion for this batch is passed; generate sedimentation issue identifiers for test cases that pass the backfeed admission criteria, which are concatenated values of source prefix, document date, source type, and issue text hash value; back up the original static seed file before writing it to the static seed pool; record at least the source, merged source type, issue identifier before sedimentation, sedimentation batch directory, and sedimentation time for the backfeed test cases, and record the batch directory, number of new additions, and new issue identifiers in the daily sedimentation history.
[0073] This embodiment employs a multi-condition joint judgment system, including criteria such as being included in the evaluation benchmark, having a valid issue identifier, being deduplicated from text and not having the same issue in the static seed pool, and having a pass admission conclusion. This rigorously selects only high-quality test cases that pass quality control and feeds them back into the static seed pool, ensuring the quality purity of static regression assets from the source. By backing up the original static seed files and fully recording information such as the source of the backfeed, fusion type, sedimentation identifier, batch directory, and time, a complete backfeed audit chain is formed. This enables test assets to have standardized management capabilities that are reproducible, auditable, and rollbackable, providing a guarantee for the long-term and robust evolution of the testing framework.
[0074] In practical applications, this embodiment generates stable precipitated issue identifiers for the selected test cases. The precipitated issue identifier = source prefix_document date_source type_issue text hash value. Issues from the same source, in the same batch, and in the same issue have stable identifiers. When identifiers conflict, a unique identifier is generated by appending a sequence number.
[0075] Example 2
[0076] Exemplary Example 2 of the present invention provides a database-based method for the self-evolution of test cases in a multi-source seed fusion evaluation RAG knowledge base. This embodiment further describes the method of the present invention in a specific scenario.
[0077] This embodiment uses a daily admission batch triggered by a Jenkins daily scheduled task on August 19, 2026, as an example to illustrate the overall execution process of the method of this invention. The target document date was not explicitly passed in for this trigger; the framework defaults to using the previous day (August 18, 2026) as the batch date, and the running root directory is runs / 20260818 / . Jenkins is an open-source automation server primarily used for continuous integration and continuous delivery. In this embodiment, Jenkins is used to automatically trigger the entire admission testing process for new knowledge ingestion daily.
[0078] The real-world session capture window was from 00:00 to 24:00 on August 18, 2026, capturing a total of 86 sessions and 312 messages. After rule filtering, 54 candidates remained, with the top 30 retained based on quality scores. 24 passed the large model review, and 6 were rejected and added to the rejection details. Three documents were added or updated within the target date, generating 18 admission seed questions, of which 15 were retained after deduplication. The static seed pool contained 520 questions, with a deterministic sampling limit of 200. The unified test input set for this batch after fusion totaled 239 questions, with the source composition shown in Table 1 below.
[0079] Table 1
[0080]
[0081] The testing framework performed a search on 239 questions, with an internal candidate count of 20, ultimately retaining the Top-10 candidates. Tiered labeling used a relevance threshold of 2 and a strong relevance threshold of 3. The labeling results showed that 231 questions formed valid core evidence and entered the evaluation benchmark, while 8 questions had no strong relevance candidates and entered the manual review list.
[0082] Then, RAGAS evaluation was performed based on 231 benchmark questions. The evaluation results and score calculation are shown in Table 2 below (time consumption scoring parameters: average time excellent 0.5 seconds, poor 2.0 seconds; P95 time excellent 1.0 second, poor 3.0 seconds; pass threshold 85, conditional pass threshold 70). RAGAS (Retrieval Augmented Generation Assessment) is an open-source evaluation framework designed specifically for evaluating the performance of RAG (Retrieval Augmented Generation) systems.
[0083] Table 2
[0084]
[0085] The final score of 86.4 is greater than or equal to the passing threshold of 85, and the admission conclusion for this batch is passed. Meanwhile, the baseline summary value for this batch is consistent with the previous day's batch, indicating that the evaluation benchmark itself has not changed, and the change in quality indicators can be attributed to the retrieval side. Gradient stress testing was performed on the RAGFlow retrieval interface using the knowledge base's mandatory question set, and the results are shown in Table 3 below.
[0086] Table 3
[0087]
[0088] The success rate and tail latency of each concurrent gradient all meet the threshold, and the degradation judgment is not triggered.
[0089] P50 (50th Percentile) is commonly referred to as the median in performance testing. It represents the time taken by 50% of requests in a dataset that is lower than this value, and the time taken by the other 50% that is higher than this value.
[0090] The 95th Percentile Latency (P95) is a percentile metric in search service performance, indicating that in a given test, 95% of search requests took less than this value. For example, if a batch of tests shows a P95 latency of 1.08 seconds, it means that 95% of the search requests in that batch were completed within 1.08 seconds, and only 5% of the requests took longer than 1.08 seconds.
[0091] P99 (99th Percentile) is a key metric for measuring system performance, particularly tail latency. It indicates that in a test, 99% of requests take less time than this value, and only 1% of requests take more time than it.
[0092] Of the eight issues in the manual review list, two came from the newly added "V6.5.2 Version Difference Explanation" document. Through candidate pool analysis, it was found that the document's core facts were fragmented due to its excessively long slice. Only weakly relevant candidates (relevance score of 1) were recalled for these two issues, failing to form core evidence. This document has been added to the slice review list, requiring adjustments to the slicing strategy and a re-execution of the database entry test. The corresponding two admission seed issues will not participate in this batch of refactoring.
[0093] This batch of dynamic source test cases totaled 39 (24 session source and 15 admission source). Of these, 8 underwent manual review, 3 were skipped due to duplication with static library normalized text, and the remaining 28 met all the backfeeding conditions. The test framework backed up the original static seed file to backup / static_seeds_20260818.json, and wrote the 28 new test cases (18 session source and 10 admission source) into the static seed pool, expanding the static seed pool from 520 to 548. The batch directory, number of new cases, and new issue identifiers were recorded in the daily accumulation history. The backfeeding results are shown in Table 4 below.
[0094] Table 4
[0095]
[0096] Thus, this batch completes a full closed loop from new knowledge identification, multi-source use case fusion, evidence-level evaluation, admission determination, performance monitoring to test asset accumulation, and all intermediate products are archived in the batch catalog, which can be reproduced and audited through the run list.
[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0098] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A database-based method for evaluating the self-evolution of test cases in a multi-source seed fusion assessment RAG knowledge base, characterized in that... The method includes: Step S1: Obtain static seed questions, session seed questions, and admission seed questions from the static seed pool, real user sessions, and newly added or updated knowledge base documents within the target time window, respectively, and construct a unified test input set through normalization deduplication and deterministic sampling; Step S2: Based on the unified test input set, call the RAG knowledge base retrieval interface to build a candidate text fragment pool, perform relevance classification and labeling on the candidate fragments, extract core evidence with the strong relevance threshold as the boundary, and generate the evaluation benchmark for this batch based on the core evidence of each test case; Step S3: Perform retrieval quality evaluation based on the evaluation criteria, calculate the final admission score by weighting the retrieval quality indicators and retrieval time indicators, and output the admission conclusion based on the final admission score; Step S4: Feed back the admission conclusions that meet the preset conditions and originate from the session seed problem or admission seed problem, and which have been deduplicated, to the static seed pool.
2. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S1, the session seed question is obtained as follows: real user session records are captured from a specified time window and broken down into message sequences, and knowledge-related question statements directly raised by users are identified; rule filtering is performed on the identified knowledge-related question statements, which includes removing system prompt words, batch prompts, irrelevant chatter and format noise; The filtered knowledge-based question statements are scored using a pre-defined quality scoring rule. A pre-defined number of candidate questions are retained based on the scores. A semantic analysis model is then invoked to review the retained candidate questions, forming a candidate set of conversation seed questions.
3. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S1, the admission seed questions are obtained as follows: newly created or newly updated documents in the knowledge base are filtered according to the target document date, and the document title, paragraphs, text slice content, product version, knowledge module and key terms are extracted as the generation context. Natural language questions for testing the retrieval of new knowledge are generated based on the generation context, forming a candidate set of admission seed questions.
4. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S1, a unified test input set is constructed as follows: the source type is labeled for each seed question, and the question text of each seed question is normalized. The normalization process includes at least removing leading and trailing whitespace characters, merging consecutive whitespace characters, and unifying the case of letters. The normalized question text is deduplicated. When the same question appears in multiple sources at the same time, a single record is retained according to a preset priority. A hash sampling key is calculated. The hash sampling key is the hash value of the concatenation of a fixed seed value and the normalized question text. Test cases are selected according to the hash sampling key, and the number of test cases does not exceed the configured limit, forming a unified test input set.
5. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S2, a candidate text fragment pool is constructed as follows: For each test case in the unified test input set, the RAG knowledge base retrieval interface is called to obtain an internal candidate set, the number of which is greater than the final retention number; a preset number of final candidate sets with the highest rankings are retained from the internal candidate set, and the candidate fragment identifier, document identifier, document name, product version, retrieval ranking, similarity, content summary and retrieval time of each candidate fragment in the final candidate set are recorded to form a candidate text fragment pool.
6. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S2, candidate fragments are labeled with relevance levels, and core evidence is extracted based on a strong relevance threshold. This includes: using a large model to score each candidate text fragment in batches, with relevance scores ranging from 0 to 3. A score of 0 indicates irrelevant to the question and is discarded; a score of 1 indicates weak relevance and is not included in any evidence set; a score of 2 indicates relevance and is included in the relevant evidence set; and a score of 3 indicates strong relevance and is included in the core evidence candidate set. A relevance threshold of 2 and a strong relevance threshold of 3 are set, allowing only candidate fragments with relevance scores reaching the strong relevance threshold as core evidence. When no candidate fragment reaches the strong relevance threshold, the corresponding test case is marked as requiring manual review and is not included in the evaluation benchmark.
7. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S2, the batch evaluation benchmark is generated based on the core evidence of each test case in the following manner: For each test case with core evidence, a correspondence is established between three types of information: question information, evidence information, and answer reference information. Question information includes at least a question identifier, question text, module, and keywords. Evidence information includes at least a set of relevant fragments, a set of core fragments, and corresponding document identifiers. Answer reference information consists of reference answers and factual points formed based on the core evidence. The question text, the set of relevant evidence identifiers, and the set of core evidence identifiers are concatenated into an original concatenated string. The hash value of the original concatenated string is calculated using a standard hash algorithm. This hash value is used as the benchmark summary value. The set of relevant evidence identifiers consists of the identifiers of all candidate fragments whose relevance scores reach the relevance threshold and the set of identifiers of all candidate fragments whose relevance scores reach the strong relevance threshold.
8. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S3, when performing retrieval quality evaluation based on the evaluation benchmark, the RAG knowledge base retrieval interface is called to obtain the returned context and determine the context precision and context recall for each test case in the evaluation benchmark. The context precision is the proportion of relevant content in the returned context, and the context recall is the proportion of core evidence in the evaluation benchmark that is successfully recalled.
9. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S3, the final admission score is calculated by weighting the retrieval quality index and the retrieval time index. Based on the final admission score, the admission conclusion is output, specifically including: Set the context precision score to context precision multiplied by 100, and the context recall score to context recall multiplied by 100; The average time consumption score and P95 time consumption score are calculated using a linear decay function. When the actual time consumption is less than or equal to the preset excellent time consumption, the value is 100; when it is greater than or equal to the preset deteriorated time consumption, the value is 0. Otherwise, the score is calculated by multiplying 100 by the ratio of the difference between the deteriorated time consumption and the actual time consumption to the difference between the deteriorated time consumption and the excellent time consumption. The final admission score is equal to the sum of the context precision score multiplied by 0.40, the context recall score multiplied by 0.40, the average execution time score multiplied by 0.10, and the P95 execution time score multiplied by 0.
10. When the final admission score is greater than or equal to the preset pass threshold, the pass conclusion is output; when it is greater than or equal to the preset conditional pass threshold but less than the preset pass threshold, the conditional pass conclusion is output; when it is less than the preset conditional pass threshold, the failure conclusion is output.
10. The database-based multi-source seed fusion evaluation RAG knowledge base test case self-evolution method according to claim 1, characterized in that, In step S4, the test cases are fed back to the static seed pool in the following manner: Set backfeed admission criteria, which include at least the following: test cases have entered the evaluation benchmark, have valid problem identifiers, have been deduplicated by text and do not have the same normalized problem text in the static seed pool, and the admission conclusion for this batch is passed; For test cases filtered by the re-feedback admission criteria, a sedimentation issue identifier is generated. This sedimentation issue identifier is a concatenation of the source prefix, document date, source type, and issue text hash value. Before writing to the static seed pool, back up the original static seed file; after the backfilling, the test cases should at least record the source, the type of source to be merged, the problem identifier before sedimentation, the sedimentation batch directory and the sedimentation time; Record the batch catalog, the number of new additions, and the identification of new issues in the daily sedimentation history.
Citation Information
Patent Citations
Method and system for evaluating quality of RAG knowledge base driven by large language model
CN119226753A
RAG system intelligent evaluation method and device based on large model, and computer equipment
CN121144337A
RAG knowledge base evaluation and self-optimization method and system based on three-dimensional dynamic calibration
CN121722885A