RAG agent system automatic evaluation method and system

By using automated evaluation methods and multi-dimensional indicators, the problems of low evaluation efficiency and insufficient applicability of the RAG system have been solved, achieving efficient and comprehensive evaluation in the financial field and improving the application effect of the system in financial business.

CN120011186BActive Publication Date: 2025-10-21BEIJING ZHONGKE JINCAI TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411904870.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-21
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing RAG system evaluation methods are inefficient, require a lot of manual intervention, lack systematic evaluation, and are difficult to adapt to the high precision and compliance requirements of the financial sector.

Method used

An automated evaluation method is adopted, which uses the BERT model and K-means clustering algorithm for semantic segmentation to generate multiple semantically complete document blocks. It combines implicit semantic reasoning to generate questions, covering multi-dimensional evaluation indicators, including relevance, accuracy and similarity scores, and generating a comprehensive evaluation report to support customized evaluation in the financial industry.

Benefits of technology

Significantly reduce human intervention, shorten evaluation time, build a comprehensive evaluation framework, improve the applicability and accuracy of the RAG system in financial business, meet the needs of rapid iteration and optimization, and enhance intelligent question answering and data processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011186B_ABST
    Figure CN120011186B_ABST
Patent Text Reader

Abstract

The application discloses an RAG intelligent agent system automatic evaluation method and system, relates to the technical field of artificial intelligence, and comprises the following steps: receiving original document data uploaded by a user, performing semantic segmentation on the original document based on a natural language processing technology, generating a plurality of semantically complete document blocks, automatically generating questions and corresponding answers in different scenes according to different conditions and in combination with the content of the document blocks, including questions generated based on a single document block, cross-block questions generated in combination with a plurality of document blocks, and implicit questions generated based on implicit semantic reasoning; the RAG intelligent agent system automatic evaluation method and system significantly reduce the dependence on manual intervention, reduce the complexity of data labeling and result verification, and greatly shorten the evaluation time, thereby meeting the demand for rapid iteration and optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an automated evaluation method and system for a RAG intelligent agent system. Background Art

[0002] In the current technical field, RAG (Retrieval-Augmented Generation) intelligent agent system, as an emerging artificial intelligence technology, has been widely used in scenarios such as text generation and question answering. However, in the specific application process, especially in the financial field.

[0003] The existing RAG system evaluation methods have the following major problems: First, the evaluation efficiency is low. In existing technologies, the performance evaluation of RAG systems usually requires a lot of manual intervention, such as data labeling and result verification, which makes the evaluation process time-consuming and labor-intensive, and difficult to meet the needs of rapid iterative optimization; second, there is a lack of systematic evaluation. Existing evaluation methods mostly focus on performance indicators in a single dimension, such as accuracy or generation quality, and fail to form a comprehensive systematic evaluation framework. This limitation makes it impossible for existing technologies to fully identify and optimize the potential deficiencies of the system; in addition, existing technologies are difficult to adapt to the complexity of financial professional scenarios. Application scenarios in the financial field usually require processing high-precision and high-reliability data content, while also meeting the compliance requirements of specific industries. Existing RAG system evaluation methods lack special designs for financial scenarios, which limits their application effects in financial business. Summary of the Invention

[0004] The purpose of the present invention is to provide a RAG intelligent agent system automatic evaluation method and system to solve the problems of the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an automated evaluation method for a RAG agent system, the method comprising:

[0006] Receive original document data uploaded by users;

[0007] Perform semantic segmentation on the original document based on natural language processing technology to generate multiple semantically complete document blocks;

[0008] Based on different conditions and the content of document blocks, questions and corresponding answers for different scenarios are automatically generated, including questions generated based on a single document block, cross-block questions generated by combining multiple document blocks, and implicit questions generated based on implicit semantic reasoning;

[0009] The RAG system is automatically evaluated using the questions and answers generated above. The evaluation indicators include the relevance score of the search results and the relevance of the questions, the accuracy score of the search results and the standard answers, the precision score of the document re-ranking ability, and the similarity score of the generated answers and the standard answers.

[0010] Generate a comprehensive evaluation report, providing quantitative indicators and system optimization suggestions.

[0011] Preferably, the semantic segmentation uses a BERT model to encode the document and combines it with a K-means clustering algorithm to generate multiple semantically complete document blocks.

[0012] Preferably, the question generation includes generating questions based on subject keywords, document types, and application scenarios input by the user.

[0013] Preferably, the cross-block question generated by combining multiple document blocks includes deducing complex questions based on comprehensive information of different document blocks, and obtaining answers through semantic matching between multiple documents.

[0014] Preferably, the implicit question is generated by analyzing the implicit semantic relationship between document blocks and combining conditional reasoning to generate questions and answers.

[0015] Preferably, the evaluation indicators include the recall rate and precision of the document retrieval component, the language fluency score of the answer generation component, and the semantic similarity score between the user question and the system answer.

[0016] Preferably, the comprehensive evaluation report includes visual charts and quantitative data, showing the performance of the system in evaluations of different dimensions and providing optimization suggestions.

[0017] Preferably, the evaluation method supports dynamic adjustment of evaluation indicators according to the needs of the financial industry, including loan approval, risk analysis and report generation application scenarios.

[0018] Preferably, the generated evaluation data is stored and archived so as to track and analyze the historical performance of the system.

[0019] The RAG agent system automated evaluation system is used to implement the steps of the RAG agent system automated evaluation method. The system includes:

[0020] A data receiving module is used to receive the original document data uploaded by the user;

[0021] The semantic segmentation module is connected to the data receiving module and is used to perform semantic segmentation on the original document based on natural language processing technology to generate multiple semantically complete document blocks;

[0022] The question generation module is connected to the semantic segmentation module and is used to automatically generate questions and corresponding answers for different scenarios based on different conditions and the content of the document block. The questions include: questions generated based on a single document block, cross-block questions generated by combining multiple document blocks, and implicit questions generated based on implicit semantic reasoning;

[0023] An automated evaluation module, connected to the question generation module, for automatically evaluating the RAG system using the generated questions and answers, wherein the evaluation includes: a relevance score of the relevance between the search results and the questions, an accuracy score of the search results and the standard answers, a precision score of the document re-ranking capability, and a similarity score between the generated answers and the standard answers;

[0024] The evaluation report generation module is connected to the automated evaluation module to generate a comprehensive evaluation report, providing quantitative indicators and system optimization suggestions.

[0025] It can be seen from the above technical solution that the present invention has the following beneficial effects:

[0026] The automated evaluation method and system for the RAG intelligent agent system significantly reduces reliance on human intervention through automated evaluation processes and question generation methods, reduces the complexity of data labeling and result verification, and significantly shortens evaluation time, thereby meeting the needs of rapid iterative optimization. A comprehensive evaluation framework covering multiple dimensions has been constructed, including performance indicators such as retrieval relevance, answer accuracy, document reordering capabilities, and similarity of generated answers. It can comprehensively identify and optimize potential deficiencies in the system and improve the overall performance of the RAG system. An evaluation method is designed specifically for the high-precision and high-reliability requirements of the financial field, supporting complex professional terminology processing, contextual semantic analysis, and industry compliance requirements, thereby enhancing the applicability and accuracy of the RAG system in financial business. The modular design of the framework and the flexible evaluation indicator definition capabilities allow customized evaluation based on financial business scenarios (such as loan approval, risk analysis, report generation, etc.), effectively meeting the needs of specific application scenarios. Through automated evaluation and optimization, the intelligent question-answering and data processing capabilities of the RAG system are enhanced, thereby promoting the development of financial business towards intelligence and automation, and improving overall business efficiency and competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Flow chart of the method of the present invention;

[0028] Figure 2 This is a schematic diagram of the system module connection of the present invention. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] like Figure 1 As shown, the present invention provides a technical solution: an automated evaluation method for a RAG intelligent agent system, comprising receiving original document data uploaded by a user; semantically segmenting the original document based on natural language processing technology to generate a plurality of semantically complete document blocks; automatically generating questions and corresponding answers in different scenarios based on different conditions and in combination with the content of the document blocks, including questions generated based on a single document block, cross-block questions generated in combination with multiple document blocks, and implicit questions generated based on implicit semantic reasoning; utilizing the questions and answers generated above to automatically evaluate the RAG system, wherein the evaluation indicators include a relevance score of the relevance between the retrieval results and the questions, an accuracy score of the retrieval results and the standard answers, an accuracy score of the document rearrangement capability, and a similarity score between the generated answers and the standard answers; generating a comprehensive evaluation report to provide quantitative indicators and system optimization suggestions.

[0031] In the above method, raw document data is received through a user-uploaded interface. Natural language processing techniques (such as word segmentation, syntactic analysis, and semantic modeling) are used to semantically segment the document data, ensuring the semantic integrity of each document block and providing a foundation for the subsequent evaluation and assessment process. During the question and answer generation phase, the method combines various criteria, such as the characteristics of individual document blocks, the logical relationships between multiple document blocks, and implicit semantic reasoning, to generate questions and answers tailored to different scenarios. Single-block questions are generated based on independent document blocks; cross-block questions are generated based on the relevance of different document blocks; and implicit questions utilize implicit semantic reasoning techniques (such as semantic embedding or pre-trained language models) to explore deep semantic relationships, ensuring comprehensive question generation. For the evaluation phase, the method uses the generated questions and standard answers to comprehensively evaluate the RAG system. Evaluation metrics cover the following aspects: Relevance score: assesses the relevance of the system's search results to the question; Accuracy score: measures the consistency of search results with standard answers; Document reranking accuracy score: analyzes the system's ability to rank search results. Similarity scoring: Evaluate the textual similarity between the generated answer and the standard answer, for example, based on natural language generation evaluation indicators such as BLEU and ROUGE. Ultimately, a comprehensive evaluation report is generated based on the above indicators, providing quantitative results and optimization suggestions for system performance. Through natural language processing and automated processes, the need for manual participation is significantly reduced, and evaluation efficiency is greatly improved. Based on different types of question generation methods, the method can cover a variety of evaluation scenarios to ensure the comprehensiveness of the evaluation results. The RAG system is quantitatively evaluated through multi-dimensional indicators to provide a scientific basis for system optimization. The comprehensive evaluation report provides guiding suggestions for the improvement of the RAG system and enhances the actual application effect.

[0032] Semantic segmentation uses the BERT model to encode the document and combines it with the K-means clustering algorithm to generate multiple semantically complete document blocks. In the above embodiment, the semantic segmentation process consists of the following steps: using the BERT model to process the input document, and using its bidirectional context modeling capability to generate a high-dimensional semantic vector representation of the document content. The BERT model can effectively capture the semantic features of the document content, especially in long text processing, with excellent performance. The K-means clustering algorithm is applied to the generated high-dimensional semantic vector representation, and the document is divided into several clusters based on the similarity of the vectors. Each cluster corresponds to a semantically complete document block, ensuring that the sentences within the block are semantically consistent, and the content between blocks is independent of each other. Based on the clustering results, the document content is split into multiple semantically complete document blocks, providing a basis for subsequent question generation and evaluation. The key to this embodiment is that BERT provides high-quality semantic representation, and the K-means clustering algorithm uses the similarity of semantic vectors to effectively segment the document content, so that the document blocks can accurately express the semantic theme of the content. The above method further optimizes the semantic segmentation process in claim 1 by combining the BERT model and the K-means clustering algorithm. Its advantages include: the deep semantic representation ability of the BERT model effectively improves the accuracy of document block segmentation and reduces the loss of semantic information. Combined with K-means clustering, the method can automatically adapt to documents of different types and structures without excessive manual adjustments. Higher-quality document block segmentation provides more stable input data for question generation and evaluation, further improving the reliability and accuracy of the evaluation. Both the BERT model and the K-means clustering algorithm have high computational efficiency and are suitable for processing large-scale document data.

[0033] Question generation includes generating questions based on the subject keywords, document types, and application scenarios input by the user. In the above embodiment, the question generation process is carried out based on the specific conditions input by the user, and mainly includes the following steps: the user inputs the subject keywords, and the system performs semantic analysis on the keywords through natural language processing technology to understand its core meaning and related concepts. During the parsing process, the system will expand or perform association analysis on the keywords, such as generating synonyms, hyponyms, and hyponyms related to the topic based on the pre-trained language model. The system adjusts the question generation rules according to the document type (such as technical documents, news reports, legal terms, etc.). For example, for technical documents, the main questions generated focus on operating procedures or technical features; while for news documents, they may tend to be more about event background or development trends. Combined with application scenario information (such as education scenarios, medical scenarios, or business scenarios), adjust the tone, scope, and focus of question generation. For example, in an education scenario, questions may be mainly guiding questions, while in a business scenario, they may focus more on decision support questions. The specific questions generated based on the above conditions are divided into the following categories: questions closely generated around the keywords input by the user, such as "What is the core technology about [keyword]?" Detailed questions related to keywords generated based on document content, such as "What are the implementation steps of [keyword]?" Expanded questions generated based on scenario characteristics, such as "How to apply [keyword] in [scenario]?" Through the above steps, the system ensures that the generated questions can cover the core content of user needs, while adapting to document characteristics and usage scenarios. Based on the keywords, document types and scenarios input by the user, the system can generate customized questions that better meet user needs, improve the pertinence and practicality of question generation, and the generated questions cover multi-level needs from the core of the topic to scenario applications, enhancing the integrity and depth of the evaluation process. Through automated condition adaptation and question generation processes, manual intervention is greatly reduced, and the efficiency of question generation is improved. The scenario adaptation function makes the generated questions more practical, which helps users achieve more precise evaluation goals in specific application fields.

[0034] Cross-block questions generated by combining multiple document blocks include deducing complex questions based on the comprehensive information of different document blocks, and obtaining answers through semantic matching between multiple documents. In this embodiment, the system analyzes the semantic relationships between multiple document blocks and uses cross-document semantic reasoning technology (such as context understanding and knowledge integration based on language models) to deduce answers to complex questions. The generated questions can cover the intersection of multi-document information, such as causal relationships in time series or logical associations between concepts. The answers are obtained through semantic matching and reasoning between multiple documents to ensure the accuracy and consistency of the comprehensive information. This method effectively improves the ability to generate questions and extract answers in cross-document scenarios, making the evaluation coverage more comprehensive; at the same time, the generation and answering of cross-block questions utilize the semantic connections between document blocks, improving the system's ability to integrate complex information and logical reasoning, thereby enhancing the depth and professionalism of the evaluation.

[0035] The generation of implicit questions is achieved by analyzing the implicit semantic relationships between document blocks and combining them with conditional reasoning to generate questions and answers. In this embodiment, the system first uses semantic analysis technology to deeply mine the implicit semantic relationships between document blocks, for example, discovering potential associations through logical connections between sentences, semantic similarity or contextual dependencies. Subsequently, combined with conditional reasoning methods (such as rule-based reasoning or causal reasoning based on neural networks), implicit questions are constructed and corresponding answers are derived. The generated implicit questions often reveal deep information that is not directly apparent in the document content but can be logically deduced. This method significantly expands the scope of question generation, can mine potential information that is not explicitly expressed in the document, and enhances the intelligence level of the evaluation; the implicit questions and answers generated by conditional reasoning further improve the depth and coverage of the evaluation, providing strong support for the comprehensive evaluation of the RAG system.

[0036] The evaluation indicators include the recall rate and precision of the document retrieval component, the language fluency score of the answer generation component, and the semantic similarity score between the user question and the system answer. In this embodiment, the system comprehensively evaluates the RAG system through the following evaluation indicators: evaluating the coverage and accuracy of the document retrieval component in retrieving the target document, and the calculation method is based on the standard information retrieval evaluation formula. Performing linguistic analysis on the output of the answer generation component, such as generating a score through a grammar checker or a pre-trained language model to evaluate its grammatical correctness and naturalness of expression. Evaluating the semantic consistency between the user question and the system answer through a semantic matching algorithm (such as cosine similarity based on embedding or a sentence conversion model) to measure the relevance and accuracy of the system answer. This method conducts a detailed evaluation of the RAG system from multiple perspectives of document retrieval, answer generation and user interaction, which helps to discover the strengths and weaknesses of the system in different modules; the quantitative scoring method provides a clear performance measurement standard, provides clear guidance for the optimization direction of the system, and further enhances the scientificity and practicality of the evaluation.

[0037] The comprehensive evaluation report includes visual charts and quantitative data, which show the performance of the system in the evaluation of different dimensions and provide optimization suggestions. In this embodiment, the evaluation report is based on multi-dimensional data analysis, and uses visualization tools (such as bar charts, line charts, radar charts, etc.) to intuitively display the performance of the system in various evaluation indicators (such as recall rate, precision, fluency, similarity, etc.). At the same time, combined with quantitative data, a specific score and its weight are provided for each indicator to ensure that users can fully understand the performance of the system. The report also uses the analysis results to generate targeted optimization suggestions, such as recommending adjustment of model parameters or training data when the retrieval accuracy is insufficient, and recommending improvement of the language generation algorithm when the answer generation fluency is low. This method presents complex evaluation data through visualization, allowing users to quickly and intuitively understand the performance of the system in different dimensions; at the same time, combined with quantitative analysis and optimization suggestions, it significantly improves the guidance and practicality of the report, making it easier for users to implement targeted system improvements.

[0038] The evaluation method supports dynamic adjustment of evaluation indicators based on the needs of the financial industry, including application scenarios such as loan approval, risk analysis, and report generation. In this implementation, the evaluation method dynamically adjusts evaluation indicators to meet the specific needs of the financial industry: evaluating the RAG system's capabilities in processing customer credit information, document retrieval, and automatically generating approval recommendations. New evaluation indicators may include information integrity checks and risk warning accuracy. For risk factor mining related to multiple documents, evaluation indicators can be adjusted to include accuracy of cross-document information extraction, correlation analysis capabilities, and logical consistency of reports. The RAG system is evaluated for language accuracy, format compliance, and correct use of domain terminology when generating specialized reports. The evaluation method dynamically adjusts indicator weights based on predefined rules or user input, and highlights evaluation results and optimization suggestions directly relevant to financial applications in the report. This method enhances the industry adaptability of RAG system evaluation, meeting the diverse application needs of the financial industry. The dynamic adjustment of evaluation indicators ensures high relevance of evaluation results to actual business scenarios, enhancing the value and guidance of the evaluation, and helping to accelerate the implementation of the system in financial scenarios.

[0039] It also includes storing and archiving the generated evaluation data in order to track and analyze the historical performance of the system. In this embodiment, the system stores the data generated by each evaluation, including the specific scores of the evaluation indicators, detailed records of the generated questions and answers, and comprehensive evaluation reports, in a database. When storing data, it is archived in chronological order and marked with information such as application scenarios and configuration versions to facilitate the retrieval and comparison of historical data. Based on the archived data, the system supports multi-dimensional historical performance analysis, such as performance trend evaluation, comparison of improvement effects between different versions, etc., to provide data support for system optimization. This method achieves long-term tracking of the historical performance of the system through the storage and archiving of evaluation data, which facilitates the discovery of performance fluctuations and improvement trends; the archived data can also be used as a reference for subsequent optimization, which improves the continuity and scientific nature of the evaluation work and contributes to the iteration and performance improvement of the system.

[0040] A RAG agent system automated evaluation system is also provided for implementing the steps of the RAG agent system automated evaluation method. The system comprises:

[0041] A data receiving module is used to receive the original document data uploaded by the user;

[0042] The semantic segmentation module is connected to the data receiving module and is used to perform semantic segmentation on the original document based on natural language processing technology to generate multiple semantically complete document blocks;

[0043] The question generation module is connected to the semantic segmentation module and is used to automatically generate questions and corresponding answers for different scenarios based on different conditions and the content of the document block. The questions include: questions generated based on a single document block, cross-block questions generated by combining multiple document blocks, and implicit questions generated based on implicit semantic reasoning;

[0044] An automated evaluation module, connected to the question generation module, for automatically evaluating the RAG system using the generated questions and answers, wherein the evaluation includes: a relevance score of the relevance between the search results and the questions, an accuracy score of the search results and the standard answers, a precision score of the document re-ranking capability, and a similarity score between the generated answers and the standard answers;

[0045] The evaluation report generation module is connected to the automated evaluation module to generate a comprehensive evaluation report, providing quantitative indicators and system optimization suggestions.

[0046] Each module works collaboratively, with a modular design ensuring system flexibility and scalability. Automation is the core goal of every process, from data reception to evaluation report generation. Natural language processing technology and evaluation algorithms are used to achieve efficient data processing and system performance evaluation. System functions are clearly separated, and each module can be independently developed and optimized, facilitating expansion and maintenance. This covers the entire process from data reception and question generation to evaluation report generation, significantly reducing manual intervention. Multi-dimensional indicators combined with dynamic adjustment capabilities adapt to application requirements in different fields. Quantitative evaluation reports and visual displays provide a scientific basis for system optimization and iteration.

[0047] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A RAG agent system automated evaluation method, characterized in that: The method comprises: Receive original document data uploaded by users; Perform semantic segmentation on the original document based on natural language processing technology to generate multiple semantically complete document blocks; Based on different conditions and the content of document blocks, questions and corresponding answers for different scenarios are automatically generated, including questions generated based on a single document block, cross-block questions generated by combining multiple document blocks, and implicit questions generated based on implicit semantic reasoning. The RAG system is automatically evaluated using the questions and answers generated above. The evaluation indicators include the relevance score of the search results and the relevance of the questions, the accuracy score of the search results and the standard answers, the precision score of the document re-ranking ability, and the similarity score of the generated answers and the standard answers. Generate a comprehensive evaluation report, providing quantitative indicators and system optimization suggestions; The semantic segmentation uses the BERT model to encode the document and combines it with the K-means clustering algorithm to generate multiple semantically complete document blocks; The question generation includes generating questions based on the subject keywords, document type and application scenario input by the user; Cross-block questions generated by combining multiple document blocks involve inferring complex questions based on the comprehensive information of different document blocks, and obtaining answers through semantic matching between multiple documents: the system analyzes the semantic relationships between multiple document blocks and uses cross-document semantic reasoning technology to deduce answers to complex questions; the generated questions can cover the intersection of multi-document information, causal relationships in time series, or logical associations between concepts; answers are obtained through semantic matching and reasoning between multiple documents to ensure the accuracy and consistency of the comprehensive information; this effectively improves the ability to generate questions and extract answers in cross-document scenarios, making the evaluation coverage more comprehensive; at the same time, the generation and answering of cross-block questions utilize the semantic connections between document blocks, improving the system's ability to integrate complex information and logically reason, thereby enhancing the depth and professionalism of the evaluation; Implicit questions are generated by analyzing the implicit semantic relationships between document blocks and combining them with conditional reasoning to generate questions and answers: the system first uses semantic analysis technology to deeply explore the implicit semantic relationships between document blocks, and discovers potential associations through logical connections between sentences, semantic similarity or contextual dependencies; then, it combines conditional reasoning methods with rule-based reasoning or neural network-based causal reasoning to construct implicit questions and deduce corresponding answers. The generated implicit questions reveal deep information that is not directly apparent in the document content but can be logically deduced; it significantly expands the scope of question generation, can mine potential information that is not explicitly expressed in the document, and enhances the intelligence level of the evaluation; the implicit questions and answers generated by conditional reasoning further improve the depth and coverage of the evaluation, providing strong support for the comprehensive evaluation of the RAG system.

2. The RAG agent system automated evaluation method according to claim 1, characterized in that: The evaluation indicators include the recall rate and precision of the document retrieval component, the language fluency score of the answer generation component, and the semantic similarity score between the user question and the system answer.

3. The RAG agent system automated evaluation method according to claim 1, characterized in that: The comprehensive evaluation report includes visual charts and quantitative data, showing the performance of the system in evaluations of different dimensions and providing optimization suggestions.

4. The RAG agent system automated evaluation method according to claim 1, characterized in that: The evaluation method supports dynamic adjustment of evaluation indicators according to the needs of the financial industry, including loan approval, risk analysis and report generation application scenarios.

5. The RAG agent system automatic evaluation method according to claim 1, characterized in that: It also includes storing and archiving the generated evaluation data to track and analyze the historical performance of the system.

6. A RAG agent system automated evaluation system, used to implement the steps of the RAG agent system automated evaluation method according to any one of claims 1 to 5, characterized in that: The system comprises: A data receiving module is used to receive the original document data uploaded by the user; The semantic segmentation module is connected to the data receiving module and is used to perform semantic segmentation on the original document based on natural language processing technology to generate multiple semantically complete document blocks; The question generation module is connected to the semantic segmentation module and is used to automatically generate questions and corresponding answers for different scenarios based on different conditions and the content of the document block. The questions include: questions generated based on a single document block, cross-block questions generated by combining multiple document blocks, and implicit questions generated based on implicit semantic reasoning; An automated evaluation module, connected to the question generation module, for automatically evaluating the RAG system using the generated questions and answers, wherein the evaluation includes: a relevance score of the relevance between the search results and the questions, an accuracy score of the search results and the standard answers, a precision score of the document re-ranking capability, and a similarity score between the generated answers and the standard answers; The evaluation report generation module is connected to the automated evaluation module to generate a comprehensive evaluation report, providing quantitative indicators and system optimization suggestions.

Citation Information

Patent Citations

  • Automatic evaluation method and system for retrieval enhancement generation system

    CN119166785A