Knowledge question-answering method based on improved RAG and agent workflow
By improving the knowledge question-answering method of RAG and intelligent agent workflow, the semantic fragmentation and missing reference problems in the field of engineering structure are solved, and the professional knowledge retrieval with high accuracy is achieved, which is applicable to technical specification documents in the field of engineering construction.
Patent Information
- Application Number
- CN202511447700.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing RAG technology suffers from semantic breaks, version confusion, and missing references in the field of engineering structures, making it difficult to meet the demand for highly accurate professional knowledge retrieval.
By employing structured document preprocessing, semantic enhancement, metadata binding, parent-child segmentation vectorization, and agent-driven recursive retrieval workflow, and by improving the RAG and agent workflow, we can achieve accurate retrieval and reliable generation of technical specification documents.
It significantly improves the retrieval accuracy of technical specification documents to over 90%, and solves problems such as semantic breaks, version confusion, and missing references. It is suitable for practical engineering scenarios such as design review, appraisal and evaluation, and education and training.
Smart Images

Figure CN120929577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of artificial intelligence and engineering construction, and in particular to a knowledge question answering method based on an improved RAG (Retrieval-Augmented Generation) and intelligent agent workflow. It is geared towards the field of engineering structures and is especially suitable for the accurate retrieval and reliable generation of high-density, highly authoritative professional knowledge documents such as technical specifications and laws and regulations. Background Technology
[0002] With the widespread application of Large Language Models (LLMs) across various industries, Retrieval-Enhanced Generation (RAG) technology has become one of the core methods for reducing model illusion and improving the credibility of generated content. In the field of engineering construction, especially in structural engineering and geotechnical engineering, designers, reviewers, and construction managers have high requirements for the accuracy, authority, and timeliness of their searches for technical specifications and standard provisions. While current RAG technology has achieved certain results in general scenarios, it still faces significant bottlenecks in its practical application in the field of structural engineering.
[0003] Existing RAG systems suffer from the following prominent problems when processing engineering and construction documents: First, technical specifications have complex structures, including chapters, sections, clauses, items, and numerous appendices, tables, and formulas. Conventional text slicing methods are prone to incorrect segmentation at line breaks or blank lines, leading to semantic breaks and affecting the completeness of retrieval. Second, specification texts often lack explicit expressions of meta-information such as standard name, number, and version, making it difficult to accurately match target documents when users use abbreviations (such as "anti-seismic code" or "high-rise code") to search. Third, different standards have version updates and priority differences (such as mandatory standards taking precedence over recommended standards, and local standards taking precedence over industry standards), making it difficult for conventional retrieval and ranking mechanisms to effectively identify and prioritize valid and high-priority clauses. Finally, specification clauses often refer to other chapters or appendices (such as "shall comply with the provisions of Appendix F.0.1 of this standard"), while traditional RAG can only retrieve directly matching text blocks and cannot achieve recursive retrieval of related clauses, resulting in missing information.
[0004] An analysis of existing related patents reveals significant limitations: CN120316232A focuses on customer service robots and does not address the structuring of professional documents; CN120579616A constructs a water conservancy knowledge graph, but its reliance on graph construction is costly and makes it difficult to cover multiple versions of specifications; CN120296111A proposes dynamic semantic segmentation and cross-text association, but fails to solve the problem of recursive retrieval of article citations; CN120448510A introduces job-related permission filtering, suitable for financial scenarios, but does not optimize for the hierarchical structure of engineering specifications; CN120448512A supports multimodal retrieval, but focuses on image and text fusion, failing to address the issues of text semantic continuity and version priority; CN120407749A proposes meta-knowledge and dynamic optimization mechanisms, but does not achieve automatic tracking of cited articles; CN120561251A is used for voice question answering, focusing on ASR and question extraction, but does not involve the processing of engineering document characteristics. In summary, existing technologies have failed to systematically solve core problems in knowledge retrieval in the field of engineering structures, such as semantic breaks, version confusion, and missing references, and are unable to meet the requirements of high accuracy (usually >90%) for practical engineering applications.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to solve technical problems such as semantic fragmentation, version confusion, and missing references in knowledge retrieval in the field of engineering structures. To this end, a knowledge question answering method based on an improved RAG and agent workflow is provided. Through structured document preprocessing, semantic enhancement, metadata binding, parent-child segmentation vectorization, and agent-driven recursive retrieval workflow, the retrieval accuracy and generation credibility of technical specification documents are significantly improved, enabling them to be applied in actual engineering scenarios.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A knowledge question answering method based on improved RAG and agent workflow includes the following steps: Step S1, Document Preprocessing: Convert the collected technical specification documents into plain text format (such as Markdown) using OCR or a format conversion tool, and establish a five-level hierarchical structure: the first level is chapters, the second level is sections, the third level is subsections, the fourth level is clauses, and the fifth level is clauses. Step S2, Enhanced Hierarchical Headings: For headings at levels one through three, the text of each heading references the content of the heading at the previous level to achieve direct indexing; for clauses at level four and items at level five, semantic anchoring information including the standard name, standard number, year and version, and clause number is appended before the main text, such as "Article 3.0.1 of the Code for Seismic Design of Buildings GB 50011-2010"; Step S3, Complex Content Processing: Record the table content in the document as HTML tables, with each page of tables as an independent item, and name them according to "Standard Name + Appendix Number + Continued Table Number", such as "Appendix A.1-Continued Table 2 of the Code for Seismic Design of Buildings GB50011-2010". Step S4, Text Cleanup: Remove all extra blank lines from the document, leaving only line breaks before headings at each level, ensuring that text slices are continuously divided according to the granularity of the text. Step S5, Document Information Enhancement: For each text block (i.e., a continuous paragraph without blank lines), generate a summary and a typical question. The summary summarizes the core points of the article and does not repeat the article number. The question is based on the article content and includes the full name of the standard, the number, and the version information. Step S6, Metadata Binding: Collect and bind the metadata of each document, including document type, name, official number, validity (current / expired), priority (laws and regulations > management regulations > industry technical documents > technical specifications > enterprise documents), keywords, etc., and load them together during vectorization; Step S7, Segmentation and Vectorization: A parent-child segmentation method is adopted, where the parent segment is a complete text block and the child segment is a single line of text, summary, or question within the parent segment; the maximum length of the parent segment is set to a 95% guarantee threshold based on the log-normal distribution of the text block's word count; a general vectorization model is used to vectorize all segments. Step S8, Retrieval and Recall: After receiving the user's question, the LLM first rewrites and breaks down the question, converting the abbreviation into the standard full name, and breaking down complex questions into multiple sub-questions; for each sub-question, a sub-segment retrieval is performed, and after secondary sorting through the re-sorting model, the complete text is recalled according to the parent segment, and the bound metadata is returned together; Step S9, Recursive Search Trigger: The LLM analyzes the content of the initially recalled clauses. If it finds a pre-defined reference expression such as "should meet the requirements of Article XXX" or "should comply with the requirements of Appendix XX of this standard", and the referenced clause is not in the recall results, a new search statement is generated (such as "What is the content of Article XXX of the XXX standard?"), and step S8 is executed again for a second search. Step S10, Result Fusion and Deduplication: Merge and deduplicate multiple search results according to text block ID to avoid duplicate information occupying the context window; Step S11, Final Generation and Sorting: LLM combines the user's original question with the fused search results, filters outdated documents based on the validity and priority information in the metadata, organizes the answers in priority order, and generates a structured and interpretable final response.
[0008] The following is a further defined technical solution of the present invention, wherein step S1 includes: Step S101: Extract text from the scanned PDF document using OCR technology and verify the accuracy of the recognition. Step S102: Use a format parsing tool to extract structured text from electronic documents (such as Word and PDF) and retain heading level information; Step S103: Convert the extracted text into Markdown format for easier subsequent processing.
[0009] The following is a further defined technical solution of the present invention, wherein step S5 includes: Step S501: Summary generation is automatically generated by LLM based on the content of the clauses. It requires concise language and does not contain specific numerical values or table references. Step S502: Question generation is performed by LLM simulating user questions. The question must include the full name, number, and version information of the standard to ensure that different versions of the standard can be distinguished.
[0010] The following is a further technical solution of the present invention. In step S9, the termination condition of recursive retrieval is: when no new reference text is found in a certain retrieval, or when the preset maximum recursion depth (such as 3 levels) has been reached, the recursion stops.
[0011] Compared with the prior art, the present invention has the following technical effects: This invention addresses the retrieval failures caused by semantic fragmentation and missing titles in technical specification texts within the engineering structure field through structured preprocessing and semantic enhancement. Metadata binding and priority sorting ensure the authority and timeliness of the output results. An agent-driven recursive retrieval workflow enables automatic tracking and complete recall of citation relationships between clauses, significantly improving the retrieval accuracy and generation completeness of the RAG system in professional fields. Preliminary verification shows that this method achieves a retrieval accuracy of over 90% in typical engineering specification question-and-answer tasks, a 20%–30% improvement over conventional RAG methods, demonstrating its value for widespread application in practical engineering scenarios such as design review, evaluation, and training.
[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a system flowchart of the present invention; Figure 2 This is a diagram of the recursive retrieval workflow. Detailed Implementation
[0015] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0016] See Figure 1 and Figure 2 This embodiment provides a knowledge question answering method based on an improved RAG and agent workflow, the implementation process of which is as follows: 1. Preprocess the collected paper or electronic technical specification documents by using OCR technology or document parsing tools to convert them into plain text in Markdown format while preserving the original hierarchical structure.
[0017] 2. In accordance with the "Regulations for the Compilation of Engineering Construction Standards", a five-level structure was established for the document: chapters, sections, subsections, clauses, and provisions. The semantics of each level of heading were enhanced to ensure that the content of the higher-level heading is referenced by the lower-level heading.
[0018] 3. For clauses and provisions, add semantic anchor information before the main text, in the format of "《Standard Name Standard Number-Year》Article XXX", to ensure that the standard can be accurately located through its full name during retrieval.
[0019] 4. Convert complex tables in the document to HTML format, treat each table as an independent item, and name them as "Appendix XX of the Standard Name - Continued Table Y".
[0020] 5. Clean up all unnecessary blank lines in the document, leaving only line breaks before headings at each level to ensure continuous text blocks.
[0021] 6. For each complete text block, call the large language model to generate a summary and a typical question. The summary is used to summarize the key points, and the question is used to enhance the search matching.
[0022] 7. Configure metadata for each document, including name, number, validity, priority, etc., and store it together with the text segments during vectorization.
[0023] 8. A parent-child segmentation strategy is adopted, with the parent segment being a complete clause and the child segment being a single-line text, summary, or question. The upper limit of the parent segment length is set to the 95th percentile based on historical document statistics.
[0024] 9. For user-input questions, LLM first rewrites them, converting abbreviations such as "seismic design code" to "《Code for Seismic Design of Buildings GB 50011-2010》", and decomposes complex questions into multiple independent sub-questions.
[0025] 10. Perform vector retrieval for each sub-problem, optimize the sorting based on the sub-segment matching results through a re-sorting model, and finally recall the complete text and metadata by parent segment.
[0026] 11. If the LLM analysis recall results find references such as "should comply with Clause 5.2.3 of this standard" but the corresponding clauses are not recalled, a new query statement is generated and the search is performed again.
[0027] 12. Multiple search results are merged and deduplicated by text block ID. LLM filters outdated standards based on the validity and priority information in the metadata and organizes the answers in the order of "mandatory > local > industry > country > organization" (unless the user specifies to search according to a specific standard, even if the document is obsolete or has a low priority).
[0028] 13. Generate the final response, outputting a structured answer that includes the original text of the clauses, a statement of document validity, and answers to the questions.
[0029] The following functions can be achieved through the method of this embodiment: 1. Structured hierarchical slicing and semantic anchoring mechanism: Based on the "Regulations for the Compilation of Engineering Construction Standards" (Jianbiao
[2008] No. 182) and the GB / T 20001 series of standards, a five-level document hierarchical structure (chapter, section, subsection, clause, item) is established, and semantic anchoring information including the standard name, number, version, and clause number is added before each level title and clause to ensure that the slicing is semantically complete and can be accurately located.
[0030] 2. Semantic enhancement of text blocks and pre-generation of questions: A semantic enhancement layer consisting of a summary and typical questions is added before each text block. The summary summarizes the core points, while the questions are used to improve the search matching degree, especially to enhance the robustness to users asking questions using abbreviations.
[0031] 3. Metadata-driven priority sorting and validity filtering: Each document is bound to metadata such as type, validity, and priority. After retrieval and recall, dynamic sorting is performed in combination with a large language model to prioritize the output of valid and high-priority standard clauses, and automatically filter out expired or low-priority standards.
[0032] 4. Parent-child segmentation vectorization and sub-segment retrieval-parent segment recall mechanism: The parent-child segmentation strategy is adopted, with the parent segment being a complete text and the sub-segment being a single-line text, summary or question. Parent segment recall is triggered by sub-segment retrieval, taking into account both fine-grained matching and semantic integrity.
[0033] 5. Agent-driven recursive retrieval workflow: Introducing LLM as an agent, it standardizes and rewrites user questions, decomposes them, identifies the reference relationships in the text after the initial retrieval, automatically initiates secondary or multiple retrievals, achieves complete recall of appendices and related texts, and finally merges the results and outputs them in order of priority.
[0034] In summary, this invention systematically improves the application process of RAG in the field of engineering structures, solves key problems such as semantic breaks, missing references, and version confusion in professional document retrieval, significantly improves the accuracy and practicality of the knowledge question-and-answer system, and provides a reliable technical path for intelligent knowledge services in the field of engineering construction.
[0035] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any person skilled in the art can make many possible variations and modifications to the technical solution of the present invention, or modify it into equivalent embodiments, without departing from the scope of the present invention's technical solution. Therefore, all equivalent changes made based on the shape, structure, and principle of the present invention without departing from the scope of the present invention's technical solution should be covered within the protection scope of the present invention.
Claims
1. A knowledge question-answering method based on improved RAG and agent workflow, characterized in that, Includes the following steps: Step S1, Document Preprocessing: Convert the collected technical specification documents into plain text format using OCR or format conversion tools, and establish a five-level hierarchical structure: the first level is chapters, the second level is sections, the third level is subsections, the fourth level is clauses, and the fifth level is clauses; Step S2, Enhanced Hierarchical Headings: For the first to third level headings, the text of each level heading references the content of the previous level heading to achieve direct indexing; for the fourth level articles and the fifth level clauses, semantic anchoring information including the standard name, standard number, year and version, and article number is added in text form before the main text. Step S3, Complex Content Processing: Record the table content in the document as HTML tables, with each page of tables as an independent item, and name it according to "Standard Name + Appendix Number + Continued Table Number"; Step S4, Text Cleanup: Remove all extra blank lines from the document, leaving only line breaks before headings at each level, ensuring that text slices are continuously divided according to the granularity of the text. Step S5, Document Information Enhancement: For each text block, generate a summary and a typical question. The summary summarizes the core points of the article and does not repeat the article number. The question is based on the article content and includes the full name of the standard, the number, and the version information. Step S6, Metadata Binding: Collect and bind the metadata of each document, including document type, name, regular number, validity, priority, and keywords, and load them together during vectorization; Step S7, Segmentation and Vectorization: A parent-child segmentation method is adopted, where the parent segment is a complete text block and the child segment is a single line of text, summary, or question within the parent segment; the maximum length of the parent segment is set to a 95% guarantee threshold based on the log-normal distribution of the text block's word count; a general vectorization model is used to vectorize all segments. Step S8, Retrieval and Recall: After receiving the user's question, the LLM first rewrites and breaks down the question, converting the abbreviation into the standard full name, and breaking down complex questions into multiple sub-questions; for each sub-question, a sub-segment retrieval is performed, and after secondary sorting through the re-sorting model, the complete text is recalled according to the parent segment, and the bound metadata is returned together; Step S9, Recursive Search Trigger: The LLM analyzes the content of the initially recalled clauses. If a preset reference expression is found and the referenced clause is not in the recall results, a new search statement is generated, and step S8 is executed again for a second search. Step S10, Result Fusion and Deduplication: Merge and deduplicate multiple search results according to text block ID to avoid duplicate information occupying the context window; Step S11, Final Generation and Sorting: LLM combines the user's original question with the fused search results, filters outdated documents based on the validity and priority information in the metadata, organizes the answers in priority order, and generates a structured and interpretable final response.
2. The knowledge question answering method based on improved RAG and agent workflow as described in claim 1, characterized in that, Step S1 includes: Step S101: Extract text from the scanned PDF document using OCR technology and verify the accuracy of the recognition. Step S102: Use a format parsing tool to extract structured text from the electronic document and retain heading level information; Step S103: Convert the extracted text into Markdown format for easier subsequent processing.
3. The knowledge question answering method based on improved RAG and agent workflow as described in claim 1, characterized in that, Step S5 includes: Step S501: Summary generation is automatically generated by LLM based on the content of the clauses. It requires concise language and does not contain specific numerical values or table references. Step S502: Question generation is performed by LLM simulating user questions. The question must include the full name, number, and version information of the standard to ensure that different versions of the standard can be distinguished.
4. The knowledge question answering method based on improved RAG and agent workflow as described in claim 1, characterized in that, The termination condition for recursive retrieval in step S9 is: when no new reference is found in a certain retrieval, or when the preset maximum recursion depth has been reached, the recursion stops.
Citation Information
Patent Citations
Knowledge construction method and system based on large model and RAG technology
CN120296111A
Customer service robot system integrating knowledge base question and answer retrieval and work order processing
CN120316232A
Document knowledge management method and system based on text retrieval enhancement generation
CN120407749A
Modularized knowledge graph and retrieval enhanced large model fusion interaction method and system oriented to financial branch mechanism
CN120448510A
Railway multi-mode knowledge base question-answering system construction method of hybrid RAG architecture
CN120448512A
Cited By
Retrieval enhancement generation-based system provision processing method, system and equipment
CN121434392A