RAG information retrieval method based on tree-shaped mixed retrieval

By using a tree-like hybrid search method in RAG information retrieval, initializing the hierarchical tree and maintaining node priority, and dynamic pruning and priority updates are performed in combination with DFS strategy, the problem of slow and inaccurate search speed in massive documents is solved, and the information retrieval efficiency and accuracy are improved.

CN120371937APending Publication Date: 2025-07-25JIANGSU AEROSPACE DAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510477958.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing RAG information retrieval methods are slow and inaccurate in massive documents, especially in scenarios where data volume is huge and updates are frequent, which affects the answer accuracy and efficiency of the intelligent question-and-answer system.

Method used

Using a method based on tree-like hybrid retrieval, the hierarchical tree is initialized and the node priority is maintained, and the traversal is traversed in combination with DFS strategy, and the node priority is dynamically pruned and updated during the traversal process, the better paths are searched first, and nodes with lower correlation with user query are pruned.

Benefits of technology

It improves the efficiency of RAG information retrieval, can quickly locate the associated information queried by users in large-scale documents, combines the advantages of recall and response speed, and avoids information loss and local optimal traps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371937A_ABST
    Figure CN120371937A_ABST
Patent Text Reader

Abstract

The invention discloses an RAG information retrieval method based on tree-shaped mixed retrieval, and relates to the technical field of intelligent questions.The method includes the steps that after a to-be-traversed structure tree is initialized according to a hierarchical structure tree of document content in an external knowledge base, the priority of all nodes in the to-be-traversed structure tree is maintained; according to the RAG information retrieval method, traversal is carried out by combining the priority of each node on the basis of a traditional DFS method, and a to-be-traversed structure tree and the priority of the nodes are dynamically pruned and updated in the traversal process, so that the RAG information retrieval method can preferentially retrieve a better traversal path and dynamically prune the nodes with relatively low association degree with user query Query; according to the method, the related information of the Query queried by the user can be quickly read and effectively positioned in the document content with a large amount of information, the advantages of the recall rate and the response speed are integrated, the RAG information retrieval efficiency can be improved, and particularly, the outstanding effect is achieved in a large-scale document quick retrieval scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent question answering, and in particular to a RAG information retrieval method based on tree-shaped hybrid retrieval. Background Art

[0002] With the rapid development of natural language processing technology, intelligent question answering systems and dialogue systems driven by large-scale language models (such as GPT, deepseek, etc.) have received extensive attention and research in various fields.

[0003] Currently, more and more intelligent question answering systems have introduced the RAG (Retriever-Augmented Generation, information retrieval technology) framework. The RAG technology is an AI technology that combines retrieval and generation. In the retrieval stage, it first extracts relevant information from a large number of documents in an external knowledge base based on the user query Query, and then in the generation stage, it uses a large language model to generate answers based on the user query Query and the retrieved relevant information. The RAG technology can combine the advantages of information retrieval and generation models, extract relevant information from a large number of documents to generate more accurate answers than traditional methods, aiming to enhance the accuracy and factuality of the large language model, thereby facilitating the improvement of the reliability, consistency, and answer accuracy of intelligent question answering.

[0004] Currently, the retrieval stage of most RAG methods relies on traditional information retrieval technologies (such as inverted indexes or dense vector-based retrieval). These methods mostly rely on simple keyword matching, resulting in retrieval results often lacking semantic depth. Therefore, there are problems of slow retrieval speed and inaccuracy, which are particularly prominent in scenarios with a large number of documents and frequent data updates. Although some optimization algorithms such as FAISS can accelerate the retrieval process, in the face of a large number of documents, the retrieval accuracy and efficiency still cannot meet the real-time requirements, directly affecting the answer accuracy and answer efficiency of the intelligent question answering system. Summary of the Invention

[0005] In view of the above problems and technical requirements, this application proposes a RAG information retrieval method based on tree-shaped hybrid retrieval. The technical solution of this application is as follows:

[0006] A RAG information retrieval method based on tree-shaped hybrid retrieval, the RAG information retrieval method includes:

[0007] Initialize the structure tree to be traversed according to the hierarchical structure tree of the document content in the external knowledge base, and initialize the priority of each node in the structure tree to be traversed; the hierarchical structure tree of the document content represents the hierarchical relationship of the information paragraphs in the document content, each node has its corresponding information paragraph, and the node content of each node is a title vector generated based on the corresponding information paragraph.

[0008] Start traversing each node in the structure tree to be traversed in sequence from the root node. For any node node traversed, calculate the text similarity between the user query Query and the node content of node as the similarity score S of node. node ;

[0009] When it is detected that the pruning condition is satisfied according to the similarity score S of node node Delete all the child nodes of node from the structure tree to be traversed to update the structure tree to be traversed, otherwise keep the current structure tree to be traversed unchanged;

[0010] Update the priority of each node in the structure tree to be traversed according to the similarity score S of node node Based on the DFS strategy and the priority of each node in the structure tree to be traversed, determine the next node to be traversed. After traversing the structure tree to be traversed, output the information paragraph included in the path with the highest path score as the associated content of the retrieved user query Query.

[0011] A further technical solution thereof is that determining the next node to be traversed based on the DFS strategy and the priority of each node in the structure tree to be traversed includes:

[0012] When it is determined that node is in the active state according to the priority P of node node Determine the next node to be traversed according to the traversal order of the DFS strategy;

[0013] When it is determined that node is in the inactive state according to the priority P of node node and it is detected that there are untraversed sibling nodes of node in the structure tree to be traversed, use the untraversed sibling nodes of node in the structure tree to be traversed as the next node to be traversed;

[0014] When it is determined that node is in the inactive state according to the priority P of node node and it is detected that there are no untraversed sibling nodes of node in the structure tree to be traversed, use the node with the highest priority and in the active state in the structure tree to be traversed as the next node to be traversed.

[0015] Its further technical solution is that, according to the priority P of the node node node Detecting whether the node node is in an active state includes:

[0016] When the priority of the node node is, determine that the node node is in an active state, otherwise determine that the node node is in an inactive state;

[0017] where P queue represents the sum of the priorities of all nodes on the current traversal path from the root node to the node node, and n represents the total number of nodes included in the current traversal path from the root node to the node node.

[0018] Its further technical solution is that, according to the similarity score S of the node node node Detecting whether the pruning condition is satisfied includes:

[0019] According to the similarity scores of each node on the current traversal path from the root node to the node node, calculate the path score L of the current traversal path node ;

[0020] When it is determined that the pruning condition is satisfied, otherwise it is determined that the pruning condition is not satisfied; where L max is the maximum value of the path scores of all the paths that have been traversed, and C th is the pruning threshold.

[0021] Its further technical solution is that calculating the path score L of the current traversal path node includes calculating according to the following formula:

[0022]

[0023] where n represents the total number of nodes included in the current traversal path, and S i is the similarity score of the i-th node starting from the root node on the current traversal path.

[0024] Its further technical solution is that the pruning threshold W layer is the hierarchical weight of the level where the currently traversed node node is located in the structure tree to be traversed.

[0025] Its further technical solution is that the calculation method of the hierarchical weight W of the level where the currently traversed node node is located in the structure tree to be traversed layer is:

[0026]

[0027] Among them, d is the node depth of the currently traversed node node in the structure tree to be traversed, and the node depth of the root node is 0. sibling_count is the number of sibling nodes of node in the structure tree to be traversed whose similarity scores reach the similarity threshold.

[0028] Its further technical solution is that according to the similarity score S of node node Updating the priorities of each node in the structure tree to be traversed includes:

[0029] According to the similarity score S of node node For the priority P of node node Perform an update;

[0030] Using the updated priority of node Update the priority P of any sibling node siblings of node in the structure tree to be traversed to siblings Update to

[0031] Among them, cos(node, siblings) represents the text similarity between the node content of node and its sibling node siblings.

[0032] Its further technical solution is that according to the similarity score S of node node For the priority P of node node Perform an update including obtaining the updated priority of node according to the following formula

[0033]

[0034] Among them, S parent Is the similarity score of the parent node of node, and max(S siblings ) is the maximum value of the similarity scores of all sibling nodes of node in the structure tree to be traversed. α, β, and γ are weighting parameters and α + β + γ = 1.

[0035] Its further technical solution is that constructing the hierarchical structure tree of the document content in the external knowledge base includes:

[0036] Analyze the document content to determine the information paragraphs corresponding to the original text titles at multiple levels included in the document content, and generate a hierarchical structure tree. Each original text title is used as a node, and the information paragraph included in the original text title is used as the information paragraph corresponding to the node;

[0037] Use the BERT model to extract the entity core words of the information paragraphs corresponding to each node, generate a question-style text title for the entity core words based on the Prompt template, and use the Chinese text embedding model based on CoSENT to generate the title vector of the question-style text as the node content.

[0038] The beneficial technical effects of this application are:

[0039] This application discloses a RAG information retrieval method based on tree-shaped hybrid retrieval. After initializing the structure tree to be traversed according to the hierarchical structure tree of the document content in the external knowledge base, the method also maintains the priorities of each node in the structure tree to be traversed. Subsequently, based on the traditional DFS method, the traversal is performed in combination with the priorities of each node, and during the traversal process, the structure tree to be traversed and the priorities of the nodes are dynamically pruned and updated. This RAG information retrieval method can preferentially retrieve better traversal paths and dynamically prune the nodes with lower correlation with the user query Query, so as to quickly and effectively locate the associated information of the user query Query in the huge amount of document content. This method combines the advantages of recall rate and response speed, can improve the efficiency of RAG information retrieval, and has outstanding effects especially in the scenario of rapid retrieval of large-scale documents.

[0040] When dynamically updating the priorities of the nodes, the method updates its own priority by combining the similarity scores of each node and its associated nodes. In addition, the priority of the current traversed node is used to spread the correlation degree to the priorities of other associated nodes. This approach can better strengthen the association relationship between the nodes, thereby better comprehensively considering the overall similarity between the nodes with an association relationship, improving the accuracy of the priorities, and alleviating the local optimal trap.

[0041] During the traversal and retrieval process, the method implements pruning by comparing the ratio of the path score of the current traversal path to the maximum value of the path scores of the historical traversal paths with the pruning threshold. As the retrieval traversal progresses, the maximum value of the path scores of the historical traversal paths is continuously updated, and the pruning threshold is also dynamically calculated based on the node depth and the similarity scores of the valid sibling nodes at the same layer. Therefore, it can comprehensively consider the historical traversal paths and the overall remaining tree structure for dynamic pruning, which can improve the pruning accuracy and avoid information loss caused by pruning. Brief Description of the Drawings

[0042] Figure 1 is the method flow chart of the RAG information retrieval method in an embodiment of this application.

[0043] Figure 2 is the method flow chart of the RAG information retrieval method in another embodiment of this application.

[0044] Figure 3It is a schematic diagram of the structure of the structure tree to be traversed initialized in an example.

[0045] Figure 4 It is for Figure 3 A schematic diagram of the structure of the structure tree to be traversed obtained after dynamically pruning and updating the structure tree to be traversed in the embodiment. Specific implementation manners

[0046] The following further describes the specific implementation manners of the present application in conjunction with the accompanying drawings.

[0047] The present application discloses a RAG information retrieval method based on tree - shaped hybrid retrieval. Please refer to Figure 1 The flowchart shown. The RAG information retrieval method includes the following steps:

[0048] Step 110, initialize the structure tree to be traversed according to the hierarchical structure tree of the document content in the external knowledge base, and initialize the priorities of each node in the structure tree to be traversed.

[0049] The hierarchical structure tree of the document content represents the hierarchical relationship of the information paragraphs in the document content. The hierarchical structure tree includes multiple nodes, and each node has its corresponding information paragraph in the document content. That is, the nodes and the information paragraphs are mapped to each other to form an index, which is convenient for quickly locating the information paragraph according to the node. The multiple nodes in the hierarchical structure tree are connected to form a hierarchical structure. Each node contains the information paragraphs corresponding to all its child nodes, and the root node contains all the information paragraphs of the document content. Each node also has its own node content, and the node content of each node is a title vector generated based on the information paragraph corresponding to the node. This title vector is a high - quality vector representation of the information paragraph corresponding to the node.

[0050] The hierarchical structure tree of the document content is pre - constructed and stored according to the document content for repeated use in the RAG information retrieval process. Constructing the hierarchical structure tree of the document content in the external knowledge base includes:

[0051] The document content usually directly contains multiple levels of headings and paragraphs, such as the general heading, first - level heading, second - level heading, third - level heading, etc. from top to bottom. Therefore, first directly parse the document content to determine the information paragraphs corresponding to the multiple levels of original text headings contained in the document content and generate a preliminary hierarchical structure tree. Here, the original text heading is the text heading directly used in the document content, and each original text heading is used as a node in the preliminary hierarchical structure tree, and the information paragraph contained in the original text heading is used as the information paragraph corresponding to the node.

[0052] However, the original text title may contain redundant information, or the similarity between the original text title and the title text may not fit the actual retrieval scenario. Therefore, the hierarchical structure tree is further optimized, including: using the BERT model to extract the entity core words of the information paragraphs corresponding to each node, and generating a question-based text title for the entity core words based on the Prompt template. Through instruction fine-tuning, the text title is made closer to the user query distribution. For example, using the BERT model to extract the entity core words of the information paragraphs corresponding to each node as "Subway Environment and Equipment Monitoring System (BAS)", and further generating a question-based text title based on the Prompt template as "What is the Subway Environment and Equipment Monitoring System (BAS)?".

[0053] Then, use the Chinese text embedding model based on CoSENT (Contextualized Sentence Embedding Transformer) to generate the title vector of the question-based text as the node content of the node. CoSENT generates a high-quality vector representation of the text through in-depth understanding of the context, enabling effective capture of synonymous sentences, grammatical differences, etc. Its working principle is to use a pre-trained Chinese text embedding model to generate context-sensitive embedding vectors at the sentence level as the node content, facilitating subsequent information retrieval.

[0054] When obtaining the input user query Query and needing to retrieve the relevant content of the user query Query from the document content, first initialize the traversal structure tree according to the hierarchical structure tree of the document content. The initialized traversal structure tree has the same structure as the hierarchical structure tree, and the node content and corresponding information paragraphs of each node in the traversal structure tree are also the same as those in the hierarchical structure tree. In addition, it is necessary to initialize the priority of each node in the traversal structure tree. In one embodiment, the priorities of all nodes in the traversal structure tree are initialized to be equal, such as 0.5.

[0055] Step 120, starting from the root node, traverse each node in the traversal structure tree in turn. For any node node traversed, calculate the text similarity between the user query Query and the node content of the node node as the similarity score S of the node node node 。

[0056] First, use the Chinese text embedding model based on CoSENT to generate the embedding vector of the user query Query, and then calculate the cosine similarity between the embedding vector of the user query Query and the node content of the node node as the text similarity, so as to obtain the similarity score S of the node node node , and the similarity scores of other nodes are calculated in the same way, which will not be elaborated later.

[0057] Step 130, according to the similarity score S of node node node Detect whether the pruning condition is satisfied.

[0058] The information retrieval method of this application is implemented based on tree-shaped depth-first search (DFS). Tree-shaped depth-first search DFS is a common method for traversing tree structures. The traversal logic of traditional DFS strategies starts from the root node of the tree structure, traverses each node in the tree structure until the leaf nodes, and then backtracks to the upper-level nodes to continue the search. However, in the RAG information retrieval scenario, the hierarchical structure tree constructed from the huge amount of document content is also very large. Directly using the DFS strategy to traverse the hierarchical structure tree will cause too much time to be wasted on wrong paths, thus affecting the retrieval efficiency.

[0059] To improve the retrieval efficiency, when traversing based on the DFS strategy, this application will detect whether the pruning condition is satisfied according to the similarity score S of node node node In another embodiment, detecting whether the pruning condition is satisfied includes the following steps. Please refer to Figure 2 the flowchart of

[0060] (1) First, calculate the path score L of the current traversal path according to the similarity scores of each node on the current traversal path from the root node to node node node . One way is to directly add the similarity scores of each node on the current traversal path to obtain the path score L node . However, to better reflect the importance of nodes at different levels, in another embodiment, according to calculate to obtain the path score L node , n represents the total number of nodes included in the current traversal path, S i is the similarity score of the i-th node in the order from the root node to node node on the current traversal path. In this embodiment, when calculating the path score L node , a corresponding weighting coefficient (i + 1) is added to the i-th node starting from the root node, so that nodes at lower levels have higher weights, thus being able to better measure the content correlation degree between the current traversal path and the user query Query.

[0061] (2) When , it is determined that the pruning condition is satisfied, otherwise it is determined that the pruning condition is not satisfied. L max in this formula is the maximum value of the path scores among all the paths that have been traversed, and it changes dynamically as the traversal process progresses. C th is the pruning threshold.

[0062] In one embodiment, the pruning threshold W layeris the level weight of the currently traversed node node in the structure tree to be traversed. And further, d is the node depth of the currently traversed node node in the structure tree to be traversed. The node depth of the root node is 0, and the node depths of other nodes increase successively with the hierarchical structure. sibling_count is the number of sibling nodes of node whose similarity scores in the structure tree to be traversed reach the similarity threshold. The sibling nodes of node are the nodes in the structure tree to be traversed that are directly connected to the same parent node as node. The similarity threshold can be customized, such as set to 0.5. Based on this, when node is closer to the root node and in the shallow layer (the smaller d) or the number of valid (similarity scores reach the similarity threshold) sibling nodes is larger, the level weight W of node layer is higher, and the pruning threshold C th will automatically decrease, making it easier to retain the current traversal path and avoid pruning affecting the retrieval accuracy. Correspondingly, when node is farther from the root node or the number of valid sibling nodes is smaller, the pruning threshold C th will automatically increase to perform efficient pruning and avoid the time consumption caused by traversing wrong nodes.

[0063] Step 140, when it is detected that the pruning condition is met, it means that the current traversal path has too little relevance to the user query Query, and it is no longer necessary to continue traversing downward along the current traversal path. At this time, all child nodes of node will be directly deleted to implement dynamic pruning, avoiding wasting time traversing on the wrong path. Therefore, all child nodes of node will be deleted from the structure tree to be traversed to update the structure tree to be traversed. Otherwise, the current structure tree to be traversed remains unchanged.

[0064] Step 150, regardless of whether the structure tree to be traversed is updated, update the priorities of each node in the current structure tree to be traversed according to the similarity score S of node node including updating the priority P of node itself according to the similarity score S of node node and performing the associated diffusion of priorities. Use the updated priority of node node to update the priority P of any sibling node siblings of node in the structure tree to be traversed siblings .

[0065] (1) First, update the priority P of node itself according to the similarity score S of node node node In one embodiment, directly convert the similarity score S node into the updated priority of node​​ Similarity score S node The higher it is, the higher the obtained conversion priority However, to alleviate the local optimal trap, in another embodiment, according to the similarity score S of the node node node and the similarity scores of other associated nodes, the updated priority is calculated including obtaining the updated priority of the node node according to the following formula

[0066]

[0067] where S parent is the similarity score of the parent node of the node node. When the node node has no parent node, S is taken parent to be 0. max(S siblings ) is the maximum value of the similarity scores of all sibling nodes of the node node in the structure tree to be traversed. When the node node has no sibling nodes, take max(S siblings ) to be 0. α, β, and γ are weighting parameters and α + β + γ = 1. In one instance, α = 0.6, β = 0.3, γ = 0.1. It can be seen that the higher the priority of the node node, the higher the comprehensive level of the text similarity between the node node and its associated nodes (including its parent node and sibling nodes) and the user query Query.

[0068] (2) Then use the updated priority of the node node to update the priority P of any sibling node siblings of it in the structure tree to be traversed siblings as follows

[0069]

[0070] where cos(node, siblings) represents the text similarity between the node contents of the node node and its sibling node siblings.

[0071] Step 160, based on the DFS strategy and combined with the priorities of each node in the structure tree to be traversed, determine the next node to be traversed. Until the structure tree to be traversed is completely traversed, output the information paragraph contained in the path with the highest path score as the associated content retrieved for the user query Query.

[0072] In addition to dynamically pruning and updating the structure tree to be traversed during the traversal process, compared with the traditional DFS method, the present application also maintains and updates the priorities of each node in the structure tree to be traversed. When traversing each node in the structure tree to be traversed in sequence, it does not always default to traversing to the leaf nodes in sequence according to the DFS strategy and then returning to the upper level. Instead, it determines the next node to be traversed based on the DFS strategy combined with the priorities of each node in the structure tree to be traversed, including:

[0073] First, according to the priorities of each node in the structure tree to be traversed, it can be determined whether each node is activated. According to the priority P of node node Detecting whether node is in the activated state includes: when the priority of node is, it is determined that the node is in the activated state, otherwise it is determined that the node is in the unactivated state. Among them, P queue represents the sum of the priorities of all nodes on the current traversal path from the root node to node, and n represents the total number of nodes included on the current traversal path from the root node to node. The priority P of node node will be updated when traversing node or its sibling nodes, and the priorities of other nodes on the path where node is located will also change, resulting in a change in P queue Therefore, the activation state of the nodes in the structure tree to be traversed is dynamically changing.

[0074] Through the above method, the activation state of node can be determined according to the priority P of node node and different situations are classified as follows:

[0075] (1) When it is determined that node is in the activated state according to the priority P of node node the next node to be traversed is determined according to the traversal order of the DFS strategy. This situation is the same as the traversal order of the classic DFS strategy and will not be elaborated here.

[0076] (2) When it is determined that node is in the unactivated state according to the priority P of node node it means that the priority of node is relatively low at this time. Then, it no longer continues to traverse downward according to the traditional DFS strategy, but tends to traverse other more appropriate paths to avoid wasting too much time on the current traversal path. At this time, it is further detected whether all the sibling nodes of node in the structure tree to be traversed have been traversed.

[0077] (3) When it is detected that there are unvisited sibling nodes of node in the structure tree to be traversed, the unvisited sibling nodes of node in the structure tree to be traversed are used as the next node to be traversed. When there are multiple unvisited sibling nodes, one sibling node with a higher priority is selected.

[0078] (4) When it is detected that there are no unvisited sibling nodes of node in the structure tree to be traversed, the node with the highest priority and in the active state in the structure tree to be traversed is used as the next node to be traversed.

[0079] For example, in an instance, the initialized structure tree to be traversed is as Figure 3 shown. The structure tree to be traversed contains a total of 12 nodes, namely node A to node L. Each node has its corresponding information paragraph, and the priority of each node is initialized to 0.5. It should be noted that Figure 3 the structure tree to be traversed is only exemplary. In the case of a very large actual document, the tree structure of the structure tree to be traversed is also very complex. Then, according to the traditional DFS traversal method, each node will be traversed in the order of A, B, E, I, L, J, F, K, C, G, D, H, and the text similarity with the user query Query will be calculated respectively until the optimal path is found.

[0080] However, according to the information retrieval method of the present application, first traverse node A, calculate the similarity score of node A. Since there is only node A on the current traversal path, the path score L A is calculated according to the similarity score of node A. Since it is the first traversal, L max =L A . At this time and the calculated C th =0.6×(1 - 0.1×1)=0.54. At this time the pruning condition is not satisfied, so the current structure tree to be traversed remains unchanged. In fact, in the first few traversals, since most of the traversed nodes are in the shallow layer and is generally relatively large, the probability of pruning is relatively small. Since node A has no sibling nodes, only the priority of node A is updated.

[0081] When it is determined that node A is in the active state according to the priority of node A, continue to traverse the next layer of node B according to the traversal order of the DFS strategy. When traversing to node B, calculate the similarity score of node B. The current traversal path becomes A - B, then calculate the path score L B of the current traversal path according to the similarity scores of node A and node B. When the path score L B >L max , update L max to LB , at this time In addition, calculate the similarity scores of the sibling nodes of node B (node C and node D) separately and update them dynamically to obtain C th . Assume that at this time The pruning condition is still not satisfied, and the current structure tree to be traversed is kept unchanged. Then, update the priority of node B according to the similarity score of node B, the similarity score of its parent node, that is, node A, and the similarity scores of its sibling nodes (node C and node D), and then further update the priorities of node C and node D.

[0082] When it is determined that node B is in the active state according to the priority of node B, continue to traverse the next-level node E in the traversal order of the DFS strategy. When traversing to node E, calculate the similarity score of node E, and the current traversal path becomes A - B - E. Then calculate the path score L of the current traversal path according to the similarity scores of node A, node B, and node E E . When the path score L E < L max At this time, keep L max = L B unchanged. In addition, calculate the similarity score of the sibling node of node E, that is, node F, so as to update C dynamically th . Assume that at this time The pruning condition is satisfied, then directly delete the child nodes of node E, that is, node I, node J, and node L, to update the structure tree to be traversed. The updated structure tree to be traversed is as shown in Figure 4 . Subsequently, continue to traverse directly based on the structure tree to be traversed. Compared with the traditional DFS method, when using the method of this application, these child nodes I, J, and L with relatively low correlation degrees will not be traversed subsequently, which is beneficial to improving the retrieval efficiency. Then, update the priority of node E according to the similarity score of node E, the similarity score of its parent node, that is, node B, and the similarity scores of its sibling nodes, that is, node F, and then further update the priority of node F.

[0083] When it is determined that node E is in the unactivated state according to the priority of node E, further detect that its sibling node, that is, node F, has not been traversed yet, then directly traverse node F. When traversing to node F, calculate the similarity score of node F, and the current traversal path becomes A - B - F. Then calculate the path score L of the current traversal path according to the similarity scores of node A, node B, and node F F . When the path score L F < L max At this time, continue to keep L max = L B unchanged. In addition, calculate the similarity score of the sibling node of node F, that is, node E, so as to update C dynamically th . Assume that at this time If the pruning condition is not met, the current structure tree to be traversed remains unchanged. Then, the priority of node F is updated based on the similarity score of node F, the similarity score of its parent node, i.e., node B, and the similarity score of its sibling node, i.e., node E. Then, the priority of node F is further updated.

[0084] When it is determined that node F is in an inactive state according to the priority of node F, the child node K of node F is no longer directly traversed according to the traditional DFS logic. And in this instance, the sibling node of node F, i.e., node E, has already been traversed. Then, directly jump to traverse the node with the highest current priority, assuming it is node D. It can be seen from this that at this time, the structure tree to be traversed still contains node K. However, since node F is in an inactive state, it will not continue to search deeper into the lower layer like traditional DFS. Instead, it will first turn to traverse a more appropriate path. Traverse node D and calculate the similarity score of node D. The current traversal path becomes A-D. Then, calculate the path score L of the current traversal path according to the similarity scores of node A and node D. D When the path score L D >L max then update L max =L D At this time In addition, C is dynamically updated according to the similarity scores of the sibling nodes of node D, i.e., node B and node C. th Assume that at this time the pruning condition is not met, then continue to keep Figure 4 the structure tree to be traversed unchanged. Then, the priority of node D is updated based on the similarity score of node D, the similarity score of its parent node, i.e., node A, and the similarity scores of its sibling nodes, i.e., node D and C. Then, the priorities of node B and node C are further updated.

[0085] Continue traversing in this way until the traversal is completed. It should be noted that when traversing to node D, the priority of node B will be updated accordingly, resulting in a change in the sum P queue of the priorities of all nodes on the path A-D-F where node F is located. Subsequently, the activation state of node F may change. When node F returns to the active state during this process, node F may still be traversed again during subsequent traversal processes, and then node K may be traversed continuously.

[0086] From the above examples, it can be seen that the DFS combined with the priority traversal strategy can give priority to retrieving better paths during the traversal and retrieval process, instead of always traversing to the leaf node like the traditional DFS method. In the traversal process, not only the priority and activation status of the nodes will be dynamically updated, but also dynamic pruning will be performed. The pruning process will change the nodes included in the structure tree to be traversed and then affect the traversal order. The traversal order will affect the maximum path score of the historical traversal path and thus affect the pruning process. These two aspects are coupled with each other, so that they are more inclined to retrieve paths with higher correlation and remove nodes with lower correlation, which is conducive to quickly and effectively finding the optimal path to output related information and improve information retrieval efficiency.

[0087] The above is only a preferred embodiment of the present application, and the present application is not limited to the above embodiments. It is understood that other improvements and changes directly derived or associated by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included in the protection scope of the present application.

Claims

1. A RAG information retrieval method based on tree-like hybrid retrieval, characterized in that, The described RAG information retrieval method includes: Initializing the structure tree to be traversed according to the hierarchical structure tree of the document content in the external knowledge base, and initializing the priorities of each node in the structure tree to be traversed; the hierarchical structure tree of the document content represents the hierarchical relationship of information paragraphs in the document content, each node has its corresponding information paragraph, and the node content of each node is a title vector generated based on the corresponding information paragraph; Traverse each node in the structure tree to be traversed sequentially starting from the root node. For any node node traversed, calculate the text similarity between the user query Query and the node content of node as the similarity score S of node node ; When the pruning condition is detected according to the similarity score S of the node node node all child nodes of the node node are deleted from the structure tree to be traversed to update the structure tree to be traversed, otherwise the current structure tree to be traversed remains unchanged; According to the similarity score S of the node node Update the priorities of each node in the structure tree to be traversed. Based on the DFS strategy and the priorities of each node in the structure tree to be traversed, determine the next node to be traversed. After traversing the entire structure tree to be traversed, output the information paragraph contained in the path with the highest path score as the associated content of the retrieved user query Query.

2. The RAG information retrieval method according to claim 1, wherein Determining the next node to be traversed based on the DFS strategy in combination with the priorities of each node in the structure tree to be traversed includes: When determining that the node is in an active state according to the priority P of the node node determine the next traversed node according to the traversal order of the DFS strategy; When according to the priority P of the node node node it is determined that the node node is in an inactive state, and when it is detected that there are un-traversed sibling nodes of the node node in the structure tree to be traversed, the un-traversed sibling nodes of the node node in the structure tree to be traversed are used as the next node to be traversed; When, according to the priority P of the node node node it is determined that the node node is in an inactive state, and it is detected that there are no un-traversed sibling nodes of the node node in the structure tree to be traversed, the node with the highest priority and in an active state in the structure tree to be traversed is used as the next node to be traversed.

3. The RAG information retrieval method according to claim 2, wherein According to the priority P of the node node node Detecting whether the node node is in an active state includes: When the priority of node node is, it is determined that node node is in an active state, otherwise it is determined that node node is in an inactive state; Among them, P queue represents the sum of the priorities of all nodes on the current traversal path from the root node to node node, and n represents the total number of nodes included on the current traversal path from the root node to node node.

4. The RAG information retrieval method according to claim 1, wherein According to the similarity score S of the node node node Detecting whether the pruning condition is satisfied includes: Calculate the path score L of the current traversal path based on the similarity scores of each node on the current traversal path from the root node to node node ; When it is determined that the pruning condition is satisfied, otherwise it is determined that the pruning condition is not satisfied; where L max is the maximum value of the path scores among all the paths that have been traversed, and C th is the pruning threshold.

5. The RAG information retrieval method according to claim 4, wherein Calculate the path score L of the currently traversed path node including calculating according to the following formula: where n represents the total number of nodes included in the currently traversed path, and S i is the similarity score of the i-th node starting from the root node on the currently traversed path.

6. The RAG information retrieval method according to claim 4, wherein Pruning threshold W layer is the level weight of the currently traversed node node in the structure tree to be traversed at the level where it is located.

7. The RAG information retrieval method according to claim 6, wherein The level weight W of the level in the structure tree to be traversed where the currently traversed node is located layer is calculated as follows: where d is the node depth of the currently traversed node node in the structure tree to be traversed and the node depth of the root node is 0, and sibling_count is the number of sibling nodes of node node in the structure tree to be traversed whose similarity scores reach the similarity threshold.

8. The RAG information retrieval method according to claim 1, wherein, According to the similarity score S of the node node Update the priorities of each node in the structure tree to be traversed, including: Based on the similarity score S of node node node update the priority P of node node node ; Using the updated priority of node Update the priority P of any sibling node siblings of node in the structure tree to be traversed siblings to where cos(node, siblings) represents the text similarity between the node content of node and its sibling node siblings 9. The RAG information retrieval method according to claim 8, wherein According to the similarity score S of the node node node the priority P of the node node node is updated, including obtaining the updated priority of the node node according to the following formula Among them, S parent is the similarity score of the parent node of node, and max(S siblings ) is the maximum value of the similarity scores of all sibling nodes of node in the structure tree to be traversed. α, β, and γ are weighting parameters and α + β + γ = 1.

10. The RAG information retrieval method according to claim 1, characterized in that, Constructing the hierarchical structure tree of the document content in the external knowledge base includes: Parsing the document content to determine the information paragraphs corresponding to the original text titles at multiple levels included in the document content to generate a hierarchical structure tree, each original text title is used as a node, and the information paragraph included in the original text title is used as the information paragraph corresponding to the node; Using the BERT model to extract the entity core words of the information paragraph corresponding to each node, and generating an interrogative text title for the entity core words based on the Prompt template, and using the Chinese text embedding model based on CoSENT to generate the title vector of the interrogative text as the node content of the node.

Citation Information

Cited By

  • Knowledge base construction method and device, computer equipment and storage medium

    CN120804234A

  • Medical drug knowledge RAG optimization method based on Trie tree

    CN121331497A