Retrieval enhancement generation method based on large language model structured document routing

By constructing a document structure tree and a retrieval subtree, and combining it with a large language model, the shortcomings of the retrieval enhancement system in understanding complex queries are addressed, achieving more efficient utilization of document knowledge and improved question-answering accuracy.

CN121070950APending Publication Date: 2025-12-05BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958334.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing retrieval enhancement systems are insufficient in understanding complex queries and struggle to effectively utilize document structure information, leading to recall noise and semantic incompleteness.

Method used

By constructing a document structure tree and retrieval subtree, and combining it with a large language model, the system dynamically organizes structured text blocks through the automatic construction of routing instruction sets and evaluation metrics, thereby improving the end-to-end performance of retrieval and generation.

Benefits of technology

By effectively utilizing document structure information, the accuracy of the question-and-answer system and the efficiency of data production were improved, while reducing manpower and annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070950A_ABST
    Figure CN121070950A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation method based on a large language model structured document route, and belongs to the field of natural language processing. According to the method, related information and structure information are comprehensively utilized, and gains are generated for retrieval and end-to-end effect generation; the method specifically comprises the steps of modeling a document routing task, automatically constructing a routing instruction set, determining end-to-end evaluation indexes of a routing module and retrieval, constructing a data structure of a document structure tree and a retrieval sub-tree, finely tuning a large language model, and constructing a text block with complete perceptual ability to a document, complete dynamic organization structure and effective content. According to the method, the document structure tree and the structured document routing big language model are combined, so that the capabilities of acquiring documents and utilizing document knowledge are improved, and the correctness of questions and answers is improved; according to the method, the structured document routing instruction set is automatically constructed, the multi-dimensional automatic evaluation index is determined, construction and evaluation of large-scale data can be efficiently supported, the data production efficiency is improved, and manpower training and labeling cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a retrieval enhancement generation method based on structured document routing of a large language model, belonging to the field of natural language processing. Background Technology

[0002] Retrieval augmentation is an important technological direction in the field of natural language processing. A basic retrieval augmentation system includes two stages: retrieval and understanding. By combining the model's internal parameters with knowledge of the external world, it significantly improves the ability of large language models to handle knowledge-intensive tasks and promotes the application of large language models in industry.

[0003] A basic search enhancement system comprises three key stages: index building, retrieval, and understanding. With technological advancements, the basic "retrieval-understanding" architecture suffers from weaknesses in understanding and planning complex queries, and insufficient robustness to recall noise (such as false positives and duplicates), leading to varying degrees of performance degradation on various challenging tasks. To address this issue, this paper focuses on the retrieval stages, promoting innovation and upgrades to each component in the search enhancement chain.

[0004] In practical applications of retrieval enhancement and question answering tasks, user queries often exhibit complex intents, requiring retrieval enhancement systems to comprehensively utilize factual knowledge from multiple documents to generate answers. Therefore, the quality of retrieval results is crucial. Vector retrieval methods discard document structural information by slicing text into fixed-length segments, which can easily lead to semantic incompleteness or redundancy, introducing noise into the reader's response. Naive retrieval enhancement methods ignore the connections between text blocks, relying entirely on relevance order to organize contextual input, making it difficult to effectively utilize information from multiple documents. Summary of the Invention

[0005] To address the lack of document structure awareness in retrieval enhancement methods, this invention aims to provide a retrieval enhancement generation method based on structured document routing using a large language model. This method comprehensively utilizes relevant and structural information to improve the retrieval enhancement method's understanding of document knowledge, thereby enhancing the end-to-end performance of retrieval and generation.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] This invention discloses a retrieval enhancement generation method based on structured document routing using a large language model. The method models the document routing task, automatically constructs a routing instruction set, determines the evaluation metrics for the routing module and the end-to-end retrieval process, constructs the data structure of the document structure tree and the retrieval subtree, fine-tunes the large language model, builds the ability to perceive documents, and dynamically organizes text blocks with complete structure and effective content. The method includes the following steps:

[0008] Step 1: Build the document structure tree offline;

[0009] The data structure for constructing a document structure tree transforms the raw document data in the knowledge base into a hierarchical document structure tree form, providing a structured foundation for subsequent retrieval.

[0010] Construct a data structure for the document tree structure to fully represent the relationship between structure and content in the document;

[0011] A document is abstracted as a document structure tree, which includes: headings, paragraphs, and their relationships.

[0012] The document structure tree nodes are divided into heading nodes and paragraph nodes: heading nodes represent the structural information in the document, and paragraph nodes represent the content information in the document;

[0013] For a document structure tree D, its nodes n∈D are represented as six-tuples, as shown in equation (1):

[0014] n=(id,text,type,parent,children,lighted) (1)

[0015] Where id∈D is the unique identifier of the node, and text∈∑ * This represents the text content of the node, where type∈{title,content} indicates the node type, and parent and children represent the relationships between the nodes. Represents a reference to the parent node. This represents the set of child nodes, where `lighted∈{True,False}` indicates the lit-up state of the nodes, and is adapted to the routing method on the document structure tree.

[0016] For all document data in the knowledge base, an offline document structure tree is constructed to represent its structure and content information. For document sources such as web pages that have their own structure, the relationship between their titles and paragraphs is directly extracted to construct a document structure tree. For document sources such as PDFs, a parsing tool is used to obtain the relationship between their structure and paragraphs to construct a document structure tree.

[0017] Step 2: Illuminate the search subtree online;

[0018] Construct a data representation of the retrieval subtree, and illuminate relevant subtree nodes in real time from the document structure tree based on query relevance;

[0019] The length of a complete document structure tree exceeds the context window size of a large language model and contains a lot of secondary information, making it unsuitable as direct input to a routing model.

[0020] In the document structure tree, the title node is a concise and refined title, which is difficult to model through relevance, while the paragraph node is a detailed and specific text description, which is usually recalled through relevance.

[0021] A node in the document structure tree is highlighted if and only if: the node in the document structure tree is a title node or the node in the document structure tree is a paragraph node and is recalled by relevance;

[0022] By using shorter heading structures to form a continuous subtree from the paragraphs scattered throughout the document, and by expanding the longer, more detailed paragraph nodes through relevance recall, we can make full use of structural information and details that are helpful to the answer, and keep the input of the routing model within a reasonable length.

[0023] Based on the characteristics of nodes in the document structure tree, the set of nodes that are highlighted in the document structure tree D is constructed as the derived retrieval subtree T, as shown in equation (2):

[0024] T={n∈D|n.lighted=True} (2)

[0025] Where n is a node in the document structure tree D, and its attribute lited=True indicates that the node is lit up;

[0026] During the retrieval enhancement generation process, the retrieval subtree is illuminated online, and a recall node is provided. Retrieve the corresponding document structure tree D from the knowledge base. i Activate the document structure tree D i All title nodes and recall nodes in the list. All sibling content nodes are thus illuminated online to obtain the corresponding retrieval subtree N. i ;

[0027] Step 3: Construct structured document routing data;

[0028] Automatically construct the instruction set required for training a large language model for structured document routing, which is divided into document-centric instruction construction method and query-centric instruction construction method;

[0029] Using a document-centric instruction construction approach: Given a document structure tree D, preferably, construct a query set Q = {q} using the Large Language Model API. i} i=1:n =LLM(D), when prompting large language models, we use the thinking chain technique to improve the ability to follow complex instructions. Based on the retrieval behavior, the query is broken down into four categories: details, integration, logic, and overview to improve the instruction set coverage. We also use few-sample prompting technology to constrain the language style to be consistent with online queries.

[0030] Obtain large language model data <D,q i > A document can correspond to multiple queries, and the results of each query can be answered by the document, thus building a model that understands the document;

[0031] Employing a query-centric instruction construction method: Given a query q, retrieve a fixed set of documents D = {D i} i=1:k Use step 2 to construct the online retrieval subtree set T = {T} i} i=1:k For each query-subtree pair <q,T i > Constructing routing results r using the Large Language Model API i =LLM(q,T) i When proposing large language models, the system uses thought chain technology to improve the ability to follow complex instructions, constrain formats, and automatically extract batch-generated results.

[0032] Obtain large language model query data <q,T i ,r i > A query corresponds to multiple documents. The results of each query can be answered by the documents or cannot be answered by the documents, which allows the model to become familiar with the online distribution and acquire the ability to refuse to answer.

[0033] Based on the query flattening, the instruction set consists of...<q,T> →r constitutes the input and output, and divides the training and testing sets;

[0034] Step 4: Determine the evaluation index system for structured document routing;

[0035] Construct a multidimensional automatic evaluation model for structured document routing, evaluate the training status of the large language model for structured document routing, select the best save point, and finally evaluate the performance.

[0036] The evaluation of end-to-end retrieval and generation includes metrics across two dimensions: end-to-end retrieval and end-to-end generation.

[0037] In the end-to-end retrieval dimension, using three manually labeled tags, the recall and precision of the retrieved text blocks in the node dimension are automatically evaluated, as shown in Equations (3) and (4):

[0038]

[0039] Among them, R ans This represents the set of nodes in the end-to-end output of the retrieval, while GT represents the set of nodes in the standard answer.

[0040] The end-to-end dimensions are generated, and the double-blind win rate of response A compared to response B is manually evaluated, as shown in equations (5) and (6):

[0041]

[0042] Where A represents the final response of method A in the end-to-end generation, B represents the final response of method B in the end-to-end generation, and Test table is the test dataset;

[0043] Step 5: Train a structured document routing large language model;

[0044] The instruction set constructed in step 3 is used to train a structured document routing large language model, and the best model storage point is selected by retrieving end-to-end metrics in step 4.

[0045] Input to the model<q,T> It consists of a query q and a retrieval subtree T. In the retrieval subtree T, a serial number is added before each lit node. The indentation length before the serial number represents the node relationship on the document structure tree, and the structural information is modeled into a text representation that is conducive to the understanding of large language models.

[0046] The output r of the model consists of the routing result r. Each node in the routing result r is processed into the form of node number and the first few characters to ensure that the large language model understands the correspondence between the number and the text content, control the decoding length, and reduce time and computing power overhead.

[0047] The structured document routing large language model is trained by obtaining only the cross-entropy loss of the output part in a supervised fine-tuning manner, as shown in Equation (7).

[0048] L=-log P(r|q,T) (7)

[0049] Where q represents the query input by the user, T represents the retrieval subtree corresponding to the data, and r represents the model output;

[0050] Step 6: Construct a retrieval enhancement model based on structured document routing;

[0051] The retrieval enhancement model based on structured document routing includes a recall module, a routing module, and a generation module;

[0052] Retrieval Module: Retrieves relevant text blocks through the retrieval tool, associates them with the corresponding document structure tree, highlights nodes, generates a retrieval subtree, and, for the user-input query q, calls the retrieval tool Retriever to retrieve k text blocks from the knowledge base. As shown in equation (8), the mapping yields the corresponding document structure tree set D = {D1, ..., D...} k}, as shown in equation (9);

[0053] C re =Retriever(q,k) (8)

[0054]

[0055] Where q represents the query entered by the user, and k represents the number of text blocks retrieved by the search engine. D represents the i-th text block to be recalled. i This represents the corresponding document structure tree;

[0056] The routing module retrieves relevant text blocks through the search engine, associates them with the corresponding document structure tree, highlights nodes, generates search subtrees, and obtains a set of search subtrees T = {T1, ..., T...}. k As shown in equation (10), in the document structure tree, the title node is always highlighted, and the paragraph node is highlighted by relevance recall. The routing module matches the user query q with the corresponding retrieval subtree T. i Input the routing model Router, and get the set of routing result nodes. As shown in equation (11), the routing text block is obtained by combining the structural information. As shown in equation (12);

[0057]

[0058]

[0059]

[0060] Where q represents the query entered by the user. D represents the i-th text block to be recalled. i T represents its corresponding document structure tree. i R represents the retrieval subtree derived from the document's structure tree. i This represents the set of nodes output by the routing model.

[0061] The generation module takes the user query and the route text block as input to the reader, generates the final response content, and combines the user's question 'q' with the set of route text blocks. Input the reading model Reader, understand and output the final answer a, as shown in Equation (13);

[0062] a = Reader(q, C ro (13)

[0063] Where q represents the query entered by the user. This represents the set of text blocks obtained during the routing phase.

[0064] It also includes step 7, applying the retrieval enhancement model obtained in step 6 to handle question-answering tasks in the field of natural language processing;

[0065] The retrieval enhancement model applied in step 6 can retrieve relevant knowledge based on user questions. By combining the document structure tree and the structured document routing big language model, it can dynamically organize text knowledge blocks that are helpful in answering questions and have a complete structure. This enhances the retrieval enhancement model's ability to acquire and utilize document knowledge, outputs answers with higher information content, and improves the accuracy of the question answering system.

[0066] Beneficial effects:

[0067] 1. The present invention provides a retrieval enhancement generation method based on structured document routing of a large language model. By combining a document structure tree with a structured document routing large language model, it dynamically organizes text knowledge blocks that are helpful to answering questions and have a complete structure, effectively utilizes document structure information, improves the system's ability to obtain documents and utilize document knowledge, and improves the accuracy of question answering.

[0068] 2. The present invention provides a retrieval enhancement generation method based on structured document routing of a large language model, which automatically constructs a set of structured document routing instructions and determines multi-dimensional automatic evaluation indicators. Through the structured document routing training task, it can efficiently support the construction and evaluation of large-scale data, improve data production efficiency, and save manpower training and annotation costs. Attached Figure Description

[0069] Figure 1 This is a flowchart of a retrieval enhancement generation method based on structured document routing of a large language model according to the present invention.

[0070] Figure 2 This is a schematic diagram illustrating the automated construction of the document routing instruction set in the embodiment.

[0071] Figure 3 This is a schematic diagram illustrating the process of constructing a document structure tree and retrieving subtrees in the embodiment.

[0072] Figure 4 The diagram shows the model constructed in the embodiment and its overall comparison with the conventional method model.

[0073] Figure 5 This is a schematic diagram illustrating the effect of the question-and-answer task in the embodiment. Detailed Implementation

[0074] To better illustrate the purpose and advantages of the present invention, the invention will be further described below in conjunction with the accompanying drawings and examples.

[0075] Example 1:

[0076] In natural language question answering scenarios, the retrieval enhancement generation method based on structured document routing using a large language model, as proposed in this invention, retrieves relevant knowledge fragments from user-input queries, routes them using document structure, and understands the knowledge to provide answers. Figure 1As shown, it includes the following steps:

[0077] Step 1: Build the document structure tree offline;

[0078] The data structure for constructing a document structure tree transforms the raw document data in the knowledge base into a hierarchical document structure tree form, providing a structured foundation for subsequent retrieval.

[0079] Construct a data structure for the document tree structure to fully represent the relationship between structure and content in the document;

[0080] A document is abstracted as a document structure tree, which includes: headings, paragraphs, and their relationships.

[0081] The document structure tree nodes are divided into heading nodes and paragraph nodes: heading nodes represent the structural information in the document, and paragraph nodes represent the content information in the document;

[0082] For a document structure tree D, its nodes n∈D are represented as six-tuples, as shown in equation (1):

[0083] n=(id,text,type,parent,children,lighted) (1)

[0084] Where id∈D is the unique identifier of the node, and text∈∑ * This represents the text content of the node, where type∈{title,content} indicates the node type, and parent and children represent the relationships between the nodes. Represents a reference to the parent node. This represents the set of child nodes, where `lighted∈{True,False}` indicates the lit-up state of the nodes, and is adapted to the routing method on the document structure tree.

[0085] For all document data in the knowledge base, an offline document structure tree is constructed to represent its structure and content information. For document sources such as web pages that have their own structure, the relationship between their titles and paragraphs is directly extracted to construct a document structure tree. For document sources such as PDFs, a parsing tool is used to obtain the relationship between their structure and paragraphs to construct a document structure tree.

[0086] In the embodiment, the constructed document structure tree is as follows: Figure 3 As shown, the structure and content information of the modeling document are as follows: Figure 3 Taking node 5 as an example, node 5 is a paragraph node containing text content. The parent node of node 5 is the heading node 4. The child nodes of node 5 are empty, and it is initially in an unlit state. The six-tuple c includes: id=5, text="Relative strength index is…", type=content, parent=4. lighted = False;

[0087] All nodes constitute a complete document structure tree, represented as D = {n} i} i=1:10 ;

[0088] Step 2: Illuminate the search subtree online;

[0089] Construct a data representation of the retrieval subtree, and illuminate relevant subtree nodes in real time from the document structure tree based on query relevance;

[0090] The length of a complete document structure tree exceeds the context window size of a large language model and contains a lot of secondary information, making it unsuitable as direct input to a routing model.

[0091] In the document structure tree, the title node is a concise and refined title, which is difficult to model through relevance, while the paragraph node is a detailed and specific text description, which is usually recalled through relevance.

[0092] A node in the document structure tree is highlighted if and only if: the node in the document structure tree is a title node or the node in the document structure tree is a paragraph node and is recalled by relevance;

[0093] By using shorter heading structures to form a continuous subtree from the paragraphs scattered throughout the document, and by expanding the longer, more detailed paragraph nodes through relevance recall, we can make full use of structural information and details that are helpful to the answer, and keep the input of the routing model within a reasonable length.

[0094] Based on the characteristics of nodes in the document structure tree, the set of nodes that are highlighted in the document structure tree D is constructed as the derived retrieval subtree T, as shown in equation (2):

[0095] T={n∈D|n.lighted=True} (2)

[0096] Where n is a node in the document structure tree D, and its attribute lited=True indicates that the node is lit up;

[0097] During the process of enhancing retrieval, the retrieval subtree is illuminated online. In the example, for instance... Figure 3 As shown on the right, given the recall nodes c1 = n1, c2 = n5, c3 = n7, retrieve the corresponding document structure tree D from the knowledge base, and light up all the title nodes {n2, n4, n6, n8} in the document structure tree D, as well as the recalled content and its sibling nodes {n1, n5, n7}, and obtain the corresponding retrieval subtree T = {n1, n2, n4, n5, n6, n7, n8} online;

[0098] Taking node 5 as an example, if the lit state of node 5 is lit=True, then node 5 is updated to a six-tuple.

[0099] Step 3: Construct structured document routing data;

[0100] Automatically build the instruction set required for training large language models for structured document routing, such as Figure 2 As shown, the methods are divided into document-centric command construction methods and query-centric command construction methods.

[0101] Using a document-centric instructional approach: Given a document structure tree D, construct a query set Q = {q} using the Large Language Model API (GPT-4o in this example). i} i=1:n =LLM(D), when prompting large language models, we use the thinking chain technique to improve the ability to follow complex instructions. Based on the retrieval behavior, the query is broken down into four categories: details, integration, logic, and overview to improve the instruction set coverage. We also use few-sample prompting technology to constrain the language style to be consistent with online queries.

[0102] Obtain large language model data <D,q i > A document can correspond to multiple queries, and the results of each query can be answered by the document, thus building a model that understands the document;

[0103] In the embodiment, the specifically constructed query instance, such as q1 = "Stock technical indicator analysis?", is a general question;

[0104] Employing a query-centric instruction construction method: Given a query q, retrieve a fixed set of documents D = {D i} i=1:k Use step 2 to construct the online retrieval subtree set T = {T} i} i=1:k For each query-subtree pair <q,T i > Using the Large Language Model API, specifically GPT-4o in this example, the routing result r is constructed. i =LLM(q,T) i When proposing large language models, the system uses thought chain technology to improve the ability to follow complex instructions, constrain formats, and automatically extract batch-generated results.

[0105] Obtain large language model query data <q,T i ,r i > A query corresponds to multiple documents. The results of each query can be answered by the documents or cannot be answered by the documents, which allows the model to become familiar with the online distribution and acquire the ability to refuse to answer.

[0106] In the embodiment, an example of the routing result specifically constructed is r1 = {n2, n4, n6, n8};

[0107] Based on the query flattening, the instruction set consists of...<q,T> →r constitutes the input and output, and divides the training and testing sets;

[0108] Step 4: Determine the evaluation index system for structured document routing;

[0109] Construct a multidimensional automatic evaluation model for structured document routing, evaluate the training status of the large language model for structured document routing, select the best save point, and finally evaluate the performance.

[0110] In this embodiment, for the query q = "stock technical indicator analysis?", given the node set R output by the end-to-end retrieval... ans ={n2,n3,n4,n5,n6,n7,n8,n9}, where the node set GT = {n2,n4,n6,n8} in the standard answer represents the generation of the end-to-end final response. A = "Stock technical indicator analysis includes tools such as moving averages (MA), relative strength index (RSI), stochastic oscillator (KDJ), and Bollinger Bands, used to identify market trends, momentum, and buy / sell signals.", B = "Stock technical indicator analysis includes Bollinger Bands and the stochastic oscillator KDJ momentum indicator. Correct application of these indicators requires combination with other analyses to form a comprehensive investment strategy.";

[0111] The evaluation of end-to-end retrieval and generation includes two dimensions of metrics: the set of nodes R of the end-to-end retrieval output. ans The standard answer contains the node set GT, which generates the end-to-end final response A, and the test set Test.

[0112] In the end-to-end retrieval dimension, using three manually labeled tags, the recall and precision of the retrieved text blocks in the node dimension are automatically evaluated, as shown in Equations (3) and (4):

[0113]

[0114] Among them, R ans This represents the set of nodes in the end-to-end output of the retrieval, while GT represents the set of nodes in the standard answer.

[0115] The end-to-end dimensions are generated, and the double-blind win rate of response A compared to response B is manually evaluated, as shown in equations (5) and (6):

[0116]

[0117] Where A represents the final response of method A in the end-to-end generation, B represents the final response of method B in the end-to-end generation, and Test table is the test dataset;

[0118] In the embodiment, this data Win-tie A The count is 1, Win-tie B The count is 0;

[0119] Step 5: Train a structured document routing large language model;

[0120] like Figure 4 As shown, the instruction set constructed in step 3 is used to train the structured document routing large language model, and the best model storage point is selected by retrieving end-to-end metrics in step 4.

[0121] Input to the model<q,T> It consists of a query q and a retrieval subtree T. In the retrieval subtree T, a serial number is added before each lit node. The indentation length before the serial number represents the node relationship on the document structure tree, and the structural information is modeled into a text representation that is conducive to the understanding of large language models.

[0122] In the example, if q = "Stock Technical Indicator Analysis", T = "1: In the stock market, technical indicators are... The following are some key technical indicators and their interpretation methods.\n\t2: 1. Moving Average (MA)\n\t4: 2. Relative Strength Index (RSI)\n\t\t5: The Relative Strength Index is a momentum indicator... which may indicate an upcoming rebound.\n\t6: 3. Stochastic Oscillator (KDJ)\n\t\t7: The Stochastic Oscillator is a commonly used momentum indicator... while the J-line value can provide additional information on market overbought or oversold conditions.\n\t8: 4. Bollinger Bands";

[0123] The output r of the model consists of the routing result r. Each node in the routing result r is processed into the form of node number and the first few characters to ensure that the large language model understands the correspondence between the number and the text content, control the decoding length, and reduce time and computing power overhead.

[0124] In the example, r = "2:1. Move...\n4:2. Relative...\n6:3. Random...\n8: Boolean...";

[0125] The structured document routing large language model is trained by obtaining only the cross-entropy loss of the output part in a supervised fine-tuning manner, as shown in Equation (7).

[0126] L=-log P(r|q,T) (7)

[0127] Where q represents the query input by the user, T represents the retrieval subtree corresponding to the data, and r represents the model output;

[0128] Step 6: Construct a retrieval enhancement model based on structured document routing;

[0129] Retrieval enhancement models based on structured document routing, such as Figure 4 As shown, it includes a recall module, a routing module, and a generation module;

[0130] Retrieval Module: Retrieves relevant text blocks through the retrieval tool, associates them with the corresponding document structure tree, highlights nodes, generates a retrieval subtree, and, for the user-input query q, calls the retrieval tool to retrieve k text blocks from the knowledge base. As shown in equation (8), the mapping yields the corresponding document structure tree set D = {D1, ..., D...} k}, as shown in equation (9);

[0131] C re =Retriever(q,k) (8)

[0132] D = {D1, ..., D} k},D i =map(c i (9)

[0133] Where q represents the query entered by the user, and k represents the number of text blocks retrieved by the search engine. D represents the i-th text block to be recalled. i This represents the corresponding document structure tree;

[0134] In the embodiments, such as Figure 5 As shown, for the user-input query q = "stock technical indicator analysis?", the retrieval module retrieves k = 2 text blocks from the knowledge base, mapping them to the corresponding document structure tree. The text blocks retrieved by the retrieval module are the same as those generated by the standard retrieval enhancement method, hitting 3 entities. Figure 5 The yellow markings indicate redundancy, missing information, or out-of-order issues.

[0135] The routing module retrieves relevant text blocks through the search engine, associates them with the corresponding document structure tree, highlights nodes, generates search subtrees, and obtains a set of search subtrees T = {T1, ..., T...}. k As shown in equation (10), in the document structure tree, the title node is always highlighted, and the paragraph node is highlighted by relevance recall. The routing module inputs the user query q and the corresponding retrieval subtree T into the routing model Router to obtain the routing result node set. As shown in equation (11), the routing text block is obtained by combining the structural information. As shown in equation (12);

[0136]

[0137]

[0138]

[0139] Where q represents the query entered by the user. D represents the i-th text block to be recalled. i T represents its corresponding document structure tree. i R represents the retrieval subtree derived from the document's structure tree. i This represents the set of nodes output by the routing model.

[0140] In the embodiments, such as Figure 5 As shown, the routing module inputs the user query q = "stock technical indicator analysis?" and the corresponding retrieval subtree into the routing model, obtaining a set of routing result nodes {n2,n3,n4,n5,n6,n7,n8,n9}. Combining the structural information, it reassembles the routing text block, hitting all four correct entities. Figure 5 The yellow markings in the center serve both practicality and completeness;

[0141] The generation module takes the user query and the route text block as input to the reader, generates the final response content, and combines the user's question 'q' with the set of route text blocks. Input the reading model Reader, understand and output the final answer a, as shown in Equation (13);

[0142] a = Reader(q, C ro (13)

[0143] Where q represents the query entered by the user. This represents the set of text blocks obtained during the routing phase.

[0144] In this embodiment, the user's question q and the set of routing text blocks are input into the reading model, and the model understands and outputs the final answer a = "Stock technical indicator analysis includes tools such as moving average (MA), relative strength index (RSI), stochastic oscillator (KDJ) and Bollinger Bands, which are used to identify market trends, momentum and buy and sell signals.", hitting all four key entities;

[0145] It also includes step 7, applying the retrieval enhancement model obtained in step 6 to handle question-answering tasks in the field of natural language processing;

[0146] The retrieval enhancement model applied in step 6 can retrieve relevant knowledge based on user questions. By combining the document structure tree and the structured document routing big language model, it can dynamically organize text knowledge blocks that are helpful in answering questions and have a complete structure. This enhances the retrieval enhancement model's ability to acquire and utilize document knowledge, outputs answers with higher information content, and improves the accuracy of the question answering system.

[0147] In the embodiments, such as Figure 5 As shown, key correct entities are highlighted in yellow. For the user query q = "stock technical indicator analysis", the present invention's answer A = "Stock technical indicator analysis includes tools such as moving averages (MA), relative strength index (RSI), stochastic oscillator (KDJ), and Bollinger Bands, used to identify market trends, momentum, and buy / sell signals." retrieves 4 key entities, and the answer hits all 4 key entities. The standard search enhancement generates the answer B = "Stock technical indicator analysis includes Bollinger Bands and the stochastic oscillator KDJ momentum indicator. Correct application of these indicators requires combination with other analyses to form a comprehensive investment strategy." retrieves 3 entities, and the answer only hits 2 entities. In summary, the present invention has performance advantages in knowledge acquisition and utilization, as well as improved accuracy of the final answer.

[0148] The above detailed description further illustrates the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for retrieval enhancement generation based on large language model structured document routing, characterized in that: Comprehensive utilization of relevant information and structural information, modeling document routing tasks, automatically constructing routing instruction set, determining routing module and retrieval end-to-end evaluation index, constructing document structure tree and retrieval subtree data structure, fine-tuning large language model, constructing perception ability of document, dynamically organizing text blocks with complete structure and effective content, improving understanding of document knowledge of retrieval enhancement method, and gaining increment of end-to-end effect of retrieval and generation.

2. The retrieval augmentation generation method based on large language model structured document routing of claim 1, wherein: Specifically includes the following steps: Step 1, offline construction of document structure tree; The data structure of the document structure tree is constructed, and the original document data in the knowledge base is converted into a hierarchical document structure tree form to provide a structured basis for subsequent retrieval; The data structure of the document structure tree is constructed to fully represent the relationship between structure and content in the document; A document is abstracted as a document structure tree, which includes: title, paragraph, and the relationship thereon; The document structure tree nodes are divided into title nodes and paragraph nodes: the title nodes represent the structural information in the document, and the paragraph nodes represent the content information in the document; For all document data in the knowledge base, the document structure tree is constructed offline to represent its structure and content information. For document sources such as web pages that have their own structure, the relationship between the title and the paragraph is directly extracted to construct the document structure tree. For document sources such as PDF, parsing tools are used to obtain the structure and paragraph relationship to construct the document structure tree; Step 2, online lighting of retrieval subtree; The data representation form of the retrieval subtree is constructed, and relevant subtree nodes are lighted from the document structure tree in real time based on query relevance; The length of the complete document structure tree exceeds the context window size of the large language model, and there is a lot of secondary information, which is not suitable for direct input of the routing model; The nodes in the document structure tree, the title nodes are concise and refined titles, which are difficult to model through relevance, and the paragraph nodes are detailed and specific text descriptions, which are usually recalled through relevance; Only when the node in the document structure tree is a title node or the node in the document structure tree is a paragraph node and is recalled by relevance, the node in the document structure tree is lighted; Use the shorter title structure to form a continuous subtree whole from the scattered paragraphs in the document, expand the longer detailed paragraph nodes through relevance recall, fully utilize the structural information and the detailed information helpful to the answer, and control the input of the routing model within a reasonable length; In the process of retrieval enhancement generation, online light up the retrieval sub-tree, given the recall node Get its corresponding document structure tree D from the knowledge base i Light up all the title nodes in the document structure tree D i , and all the sibling content nodes of the recall node , thereby online light up the corresponding retrieval sub-tree T i ; Step 3, construction of structured document routing data; Automatically construct the instruction set required for training of the structured document routing large language model, which is divided into document-centered instruction construction method and query-centered instruction construction method; Using a document-centric instruction construction method: given a document structure tree D, as preferred, construct a query set Q = {q i} i=1:n = LLM(D) using a large language model API, use the chain-of-thought technique to improve the compliance of complex instructions, according to the retrieval behavior, decompose the query into four categories: details, integration, logic, and overview, improve the coverage of the instruction set, use the few-shot prompt technique, and constrain the language style to be consistent with the online query; Get large language model data <D, q i >, a document corresponds to multiple queries, and the results of each query can be answered by the document, and the model understands the document; Adopting a query-centered instruction construction method: given a query q, retrieve a fixed number of document sets D = {D i} i=1:k , use step 2 to construct a retrieval sub-tree set T = {T i} i=1:k , for each query-sub-tree pair <q, T i >, use a large language model API to construct a routing result r i = LLM(q, T i ), when prompting the large language model, adopt the thinking chain technology to improve the compliance of complex instructions, constrain the format, and automatically extract the batch generated results; Obtain large language model query data <q, T i , i >, one query corresponds to multiple documents, and the result of each query can be answered by the document or cannot be answered by the document, so that the model is familiar with the distribution on the line and obtains the ability to refuse to answer. According to query flattening, the instruction set is composed of <q, T>→r input and output, and is divided into training and testing sets; Step 4, determination of structured document routing evaluation index system; Construct a multi-dimensional automatic evaluation model for structured document routing to evaluate the training of the structured document routing large language model, select the best save point, and finally evaluate the performance; The evaluation and generation end-to-end includes two dimensions of indicators: the node set R of the retrieval end-to-end output ans The node set GT in the standard answer direct The node set A of the generation end-to-end final reply, and the test set D Step 5, training of the structured document routing large language model; Train the structured document routing large language model using the instruction set constructed in step 3, and select the best model checkpoint by the retrieval end-to-end index in step 4; The input of the model is <q, T>, which is composed of the query q and the retrieval subtree T. In the retrieval subtree T, a serial number is added before each lit node. The indentation length before the serial number represents the node relationship on the document structure tree. The structure information is modeled as a text representation that is easy for the large language model to understand. The output of the model is r, which is composed of the routing result r. Each node in the routing result r is processed into the form of node serial number and the first few characters to ensure that the large language model understands the correspondence between the serial number and the text content, control the decoding length, and reduce the time and computational overhead. Train the structured document routing large language model in a supervised instruction fine-tuning manner, only obtain the cross-entropy loss of the output part. Step 6, construct a retrieval enhancement model based on structured document routing; The retrieval enhancement model based on structured document routing includes a recall module, a routing module, and a generation module. Retrieval module: retrieve relevant text blocks by retriever, associate corresponding document structure tree, light up nodes, generate retrieval sub-tree, call retriever Retriever to retrieve k text blocks from knowledge base for user input query q Mapping to get corresponding document structure tree set D = {D1,..., D k}; Router module: through the retriever to recall relevant text blocks, associate the corresponding document structure tree, light up the node, generate the retrieval subtree, get the retrieval subtree set T = {T1, …, T k} on the document structure tree, the title node is always lit up, the paragraph node is lit up by relevance recall, the router module combines the user query q with the corresponding retrieval subtree T i Input the routing model Router, get the routing result node set Recombine the routing text blocks obtained in combination with the structure information Generation module: input the user query and the routing text block into the reader together to generate the final reply content, and input the user question q and the routing text block set Input the reading model Reader and understand the output final answer a.

3. The retrieval augmentation generation method based on large language model structured document routing of claim 2, wherein: In step 1, for the document structure tree D, the node n ∈ D is represented as a six-tuple, as shown in equation (1): n = (id, text, type, parent, children, lighted) (1) where id 2 D is the unique identifier of the node, text 2 5 * is the text content of the node, type 2 {title, content} indicates the node type, parent, children are the relationships between nodes, represent the reference of the parent node, represent the child node set, lighted 2 {True, False} represents the node lighting state, and the routing method on the adaptive document structure tree.

4. The retrieval augmentation generation method based on large language model structured document routing of claim 2, wherein: In step 2, based on the characteristics of the nodes in the document structure tree, the set of lit nodes in the document structure tree D is constructed as the derived retrieval subtree T, as shown in equation (2): T = {n ∈ D | n.lighted = True} (2) Where n is a node on the document structure tree D, and the attribute lighted = True indicates that the node is in the lit state.

5. The retrieval enhancement generation method based on large language model structured document routing of claim 2, wherein: In step 4, the retrieval end-to-end dimension, three kinds of labels annotated by humans are used to automatically evaluate the recall of the text block in the node dimension, as shown in equations (3) and (4): wherein R ans denotes the set of nodes retrieved for the end-to-end output, GT denotes the set of nodes in the ground truth.

6. The retrieval enhancement generation method based on large language model structured document routing of claim 2, wherein: In step 4, the generation end-to-end dimension, the double-blind win rate of reply A over reply B is evaluated by humans, as shown in equations (5) and (6): Where A represents method A in generating the end-to-end final reply, B represents method B in generating the end-to-end final reply, and Test represents the test dataset.

7. The retrieval enhancement generation method based on large language model structured document routing of claim 2, wherein: In step 5, train the structured document routing large language model in a supervised instruction fine-tuning manner, only obtain the cross-entropy loss of the output part, as shown in equation (7): L = -log P(r|q, T) (7) Where q represents the user's input query, T represents the retrieval subtree corresponding to the data, and r represents the model output.

8. The retrieval enhancement generation method based on large language model structured document routing of claim 2, wherein: In step 6, for a user input query q, the retriever is invoked to recall k text blocks from the knowledge base As shown in equation (8), the mapping results in a corresponding set of document structure trees D = {D1,..., D k} as shown in equation (9); C re = Retriever(q, k) (8) where q represents a query input by a user, k represents a number of text blocks recalled by a retriever, represents the i-th text block recalled, D i represents its corresponding document structure tree; The search sub-tree set T = {T1,..., T k} is retrieved, and the routing result node set The routing text block is obtained as shown in formula (12). wherein q represents a query input by a user, the i-th text block representing a recall, D i a document structure tree corresponding thereto, T i a retrieval sub-tree derived from the document structure tree, R i a node set output by the routing model; The final answer a is shown in equation (13): a = Reader(q, C ro ) (13) wherein q represents a query input by a user, representing a set of text blocks obtained in the representative routing stage.

9. The retrieval enhancement generation method based on large language model structured document routing of claim 1 or 2, characterized in that: Step 7, apply the retrieval enhancement model obtained in step 6 to process the question and answer tasks in the natural language processing field. The retrieval enhancement model of step 6 can retrieve relevant knowledge according to user questions, combine the document structure tree with the structured document routing large language model, dynamically organize text knowledge blocks that are helpful to answering questions, improve the ability of the retrieval enhancement model to obtain and utilize document knowledge, output answers with higher information content, and improve the correctness of the question and answer system.