Large model RAG homogeneous tool query routing method based on comparative learning

Dual-directional contrastive learning and cost-efficient allocation enhance query routing in multi-document RAG systems, improving precision and reducing costs by selecting optimal document retrieval tools.

CN120316239APending Publication Date: 2025-07-15NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365619.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the multi-document RAG scenario, the router selection strategy lacks effective supervision signals, resulting in low computing efficiency and unreasonable cost, making it difficult to achieve a balance between performance and cost in a multi-tool collaboration scenario.

Method used

Using a method based on contrast learning, the router is trained through two-way contrast learning, combining bundle sampling and cost-reducing allocation inference strategies, optimized the matching of query and document search tools, and designed the router to select the best-performing document search tools and reduce the calculation cost.

Benefits of technology

It significantly improves the accuracy and efficiency of query routing, simplifies task goals, improves the scalability of the method, and realizes efficient query routing in multi-document RAG scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316239A_ABST
    Figure CN120316239A_ABST
Patent Text Reader

Abstract

The invention discloses a large model RAG homogeneous tool query routing method based on comparative learning, which comprises the following steps: collecting key information of each document in a document set, generating a corresponding abstract, and packaging the abstract as a tool description into an independent document retrieval tool; performing question and answer performance scoring on each document retrieval tool by adopting beam sampling, distinguishing a positive document retrieval tool index set and a negative document retrieval tool index set of a current retrieval problem according to the score, and constructing a positive query index set and a negative query index set required by a training router; training a router; bidirectional comparative learning is realized by taking query and document retrieval tools as anchor points respectively; and ensuring that the average value of the similarities between the new retrieval question and all the document retrieval tools in the document retrieval tool set is greater than a preset threshold value by adopting a cost reduction allocation reasoning strategy, and meanwhile, minimizing the average tool calling cost to obtain a final retrieval enhancement answer. According to the method, the routing query precision can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of natural language processing and large model retrieval enhanced generation, and particularly relates to a large model RAG homogeneous tool query routing method based on contrast learning. Background Art

[0002] With the application of the Retrieval Augmented Generation (RAG) method, large language models (LLMs) have made significant progress in natural language understanding and interactive knowledge query. In practical applications, user queries often require comprehensive analysis of multiple documents. Tool interface call is one of the typical inference applications of multi-document RAG. Given a user query, the RAG application first encapsulates all background documents into retrieval tools, the LLM executes the tool to retrieve relevant documents, and then gives a response after comprehensive analysis of the query and the documents. Considering the homogeneous and heterogeneous nature of various document retrieval tools, the ability to match queries with the best retrieval tool is crucial. Through an adaptive query routing strategy, not only can the efficiency of document retrieval and generation be optimized, but also the balance between cost and performance can be achieved in complex query tasks and large-scale generation applications, providing irreplaceable technical support for solving diverse requirements.

[0003] Early studies used simple assembly methods to make decisions on user queries uniformly after calling all documents, which not only introduced a large amount of redundant information but also had low computational efficiency. Recent studies introduced a routing selection strategy to learn and optimize the performance of the router by fitting the distribution of the selection probability of the router and the true scoring preference, and route the query tools by evaluating the normalized scores. However, when multiple documents perform well for a query, this method is inappropriate because the scores of multi-class normalization tend to be evenly distributed, which is not a strong supervision signal for learning the router. This limitation hinders the router's ability to further optimize query allocation, which is particularly prominent in multi-tool collaboration scenarios. Summary of the Invention

[0004] In view of this, this application provides a large model RAG homogeneous tool query routing method based on contrast learning.

[0005] This application discloses a large model RAG homogeneous tool query routing method based on contrast learning, which includes:

[0006] Step 1: Collect the key information of each document in the document set, generate the corresponding abstract, and use the abstract as the tool description to encapsulate it into an independent document retrieval tool; the key information includes the title, author, keywords, and core content;

[0007] Step 2: Use beam sampling to score the Q&A performance of each document retrieval tool, distinguish the positive document retrieval tool index set and the negative document retrieval tool index set for the current retrieval problem according to the score, and simultaneously construct the positive query index set and the negative query index set required for training the router;

[0008] Step 3: Use the positive and negative document retrieval tool index sets and the positive and negative query index sets to train the router to capture the association between the query and the document retrieval tool; implement two-way contrastive learning with the query and the document retrieval tool as the anchors respectively;

[0009] Step 4: In the set of document retrieval tools, define the time required for each document retrieval tool to be called and answered as the tool call cost; use the cost reduction allocation inference strategy to ensure that the average similarity between the new retrieval problem and all document retrieval tools in the set of document retrieval tools is greater than a preset threshold, while minimizing the average tool call cost to obtain the final retrieval enhanced answer.

[0010] Further, the said Step 1 includes:

[0011] Suppose there are n query questions, and there are m documents in the document set Each query question x i and its answer y i together constitute the training set

[0012] The method for processing the m documents includes:

[0013] Collect the key information of each document, define the name of the document as N j and the generated abstract as D j Take N j as the tool name and D j as the tool description, encapsulate each document retrieval function to form each document retrieval tool in the set of document retrieval tools Each query in only involves a subset of the document set The router takes the query x i as the input and generates the probability distribution of the selection of m document retrieval tools.

[0014] Further, the said Step 2 includes:

[0015] Input the query x i into the set of document retrieval tools K times to obtain the predicted output each time Define the score of each retrieval tool t j on the query x i as: If the predicted output is correct, the evaluation value is 1, otherwise it is 0.

[0016] Furthermore, for query x i , according to the score construct its positive document retrieval tool index set and negative document retrieval tool index set which are composed of document retrieval tool indexes with scores, and the remaining document retrieval tool indexes form Considering two-way contrastive learning, for document retrieval tool t j , construct positive and negative query index sets according to the query samples it processes; the positive query index set is composed of query indexes with scores, and the negative query index set is composed of query indexes with scores.

[0017] Furthermore, step 3 includes:

[0018] Train the router using contrastive learning methods, select the document retrieval tool with the best performance; design a two-way contrastive learning scheme; this two-way contrastive learning scheme includes contrastive learning with the query as the anchor and contrastive learning with the document retrieval tool as the anchor; contrastive learning with the query as the anchor aims to make the query representation tend to the positive document retrieval tool index set while moving away from the negative document retrieval tool index set; contrastive learning with the document retrieval tool as the anchor is expected to make the document retrieval tool tend to the positive query index set while moving away from the negative query index set.

[0019] Furthermore, the contrastive learning with the query as the anchor and the contrastive learning with the document retrieval tool as the anchor include:

[0020] Input query x i into the encoder ε, and the encoded representation is x i = ε(x i ; w), where w is a learnable parameter; concatenate the name N J of the document retrieval tool and the abstract D j and input them into the encoder ε to generate the vector representation t j = ε(t j ; w); the contrastive learning loss function with query x i as the anchor is defined as:

[0021]

[0022] where sim(·, ·) represents cosine similarity, t + is the positive document retrieval tool representation, and t - is the negative document retrieval tool representation;

[0023] The contrastive learning loss function with the document retrieval tool t j is defined as:

[0024]

[0025] where sim(·,·) represents the cosine similarity, x + is the positive query representation, and x - is the negative query representation;

[0026] By minimizing the weighted loss of and learn the router

[0027]

[0028] where λ is a hyperparameter used to balance and

[0029] Furthermore, the step 4 includes:

[0030] Step 41: In the set of document retrieval tools the time taken by each document retrieval tool to call the answer is defined as the tool call cost

[0031] Step 42: In the LLM inference stage, input each test query x′ i into the router Calculate the similarity score between each test query x′ i and each document retrieval tool in the set of document retrieval tools

[0032] Step 43: Obtain the optimal tool allocation scheme according to the similarity score and the tool call cost

[0033] Step 44: Based on the optimal tool allocation scheme, merge the results of all document retrieval tool calls to answer, and jointly input them into the LLM with the test query problem x′ i to obtain the final retrieval enhanced answer.

[0034] Furthermore, the step 43 includes:

[0035] The optimization objective of the cost reduction allocation inference strategy is:

[0036]

[0037] where respectively indicate whether to query x′ i is assigned to the document retrieval tool t j ;

[0038] Based on the constraint conditions for the objective function is solved to obtain the optimal tool allocation scheme.

[0039] Due to the adoption of the above technical solution, the present application has the following advantages: The present application addresses the problem of intelligent routing of large models in the multi-document RAG scenario and proposes a method for query routing of homogeneous tools for large model RAG based on contrastive learning. By using bidirectional contrastive learning to train a query-based router and introducing a cost-reducing allocation inference strategy, this method significantly improves the accuracy and efficiency of document routing retrieval, has good scalability, and can be applied to any application scenario of homogeneous tool routing. Compared with the prior art, it has the following beneficial effects and advantages:

[0040] 1. Abstracts the complex multi-document RAG problem into a tool call problem, simplifies the task objective, and improves the scalability of the method;

[0041] 2. Adopts a routing scoring mechanism based on the beam sampling strategy, breaks through the limitations of the traditional greedy decoding scoring strategy, and obtains better-quality answers by exploring a wider LLM decoding space;

[0042] 3. Proposes a novel bidirectional contrastive learning training method to co-optimize the contrastive learning objective with the query as the anchor point and the contrastive learning objective with the document retrieval tool as the anchor point, and significantly improves the accuracy and robustness of the router's matching of queries and document retrieval tools while ensuring training stability;

[0043] 4. Introduces a cost-reducing allocation inference mechanism to balance the computational cost and retrieval performance, and effectively reduces the tool call cost while ensuring the quality of retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0045] Figure 1 is a schematic flow chart of a method for query routing of homogeneous tools for large model RAG based on contrastive learning according to an embodiment of the present application;

[0046] Figure 2 is a schematic diagram of the bidirectional contrastive learning training architecture according to an embodiment of the present application. Detailed implementation manners

[0047] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0048] Therefore, the present application proposes a large model RAG homogeneous tool query routing method based on contrastive learning. Its core idea is to learn a router to improve the matching accuracy between queries and document retrieval tools. For each query, first, by comparing the output results of candidate document retrieval tools with the true answers, candidate document scores are generated. Instead of directly aligning the score distributions, the best and worst performing candidate documents are selected based on the scores. Then, a query-tool contrastive loss is proposed to pull the query representation (extracted by the encoder) closer to the representations of relevant document retrieval tools while moving away from the representations of irrelevant document retrieval tools. Through this contrastive learning strategy, the best-performing document retrieval tools can be equally selected, significantly improving the accuracy of query routing.

[0049] The present application aims to solve the problem of optimizing tool query routing in the multi-document RAG scenario, and specifically proposes solutions for the following aspects of problems:

[0050] 1) The traditional greedy decoding scoring supervision signal lacks a global perspective and easily leads to the problem that the answers of the LLM fall into a single local optimal solution;

[0051] 2) The lack of a correct guiding direction for the routing learning objective makes the tool selection in the multi-document query task lack discrimination, resulting in poor performance;

[0052] 3) The inference process often selects the optimal tool set, ignoring the cost-benefit balance problem in actual applications.

[0053] The query routing of multi-document RAG refers to the intelligent allocation of queries to the most suitable background documents to assist the LLM in achieving retrieval-enhanced question answering. The purpose is to design a low-cost routing function to replace the combined inference of the LLM on all document sets, thereby significantly reducing the computational overhead and achieving more accurate and efficient document retrieval. By optimizing the query routing process, the response speed and query quality of the RAG system can be improved.

[0054] See Figure 1 , the present application provides an embodiment of a large model RAG homogeneous tool query routing method based on contrastive learning. Based on the contrastive learning method, the supervisory role of the reward scoring signal is fully utilized to optimize the routing tool selection and improve the effect and efficiency of query processing. The embodiments of the present application include the following four steps:

[0055] S1. Tool encapsulation: Collect the key information of each document in the document collection to generate the corresponding abstract. And use this abstract as the tool description to encapsulate it into an independent document retrieval tool for subsequent tool calls and query matching operations. The key information includes the title, author, keywords, and core content;

[0056] Optionally, assume there are n query questions and m reference documents. Each query question x i and its answer y i together constitute the training set To achieve routing query optimization, first, the reference documents need to be processed: Collect the key information of each document, define the document name as N j , and generate the abstract defined as D j . Use N j as the tool name and D j as the tool description to encapsulate each document retrieval function and form a document retrieval tool set Generally, each query in only involves a specific subset of the document collection i . To ensure the accuracy of retrieval, a router model needs to be designed that can select the most suitable document retrieval tool for each query question. The router takes the query x

[0057] S2. Beam sampling scoring: Compare with the true answer, use the beam sampling method to score the Q&A performance of each document retrieval tool, and based on this score, distinguish the positive and negative document retrieval tool samples (positive document retrieval tool index set and negative document retrieval tool index set) for the current retrieval question, and construct the sample list set (positive query index set and negative query index set) required for training the router.

[0058] Optionally, to learn the router, an effective scoring method needs to be designed to evaluate the performance of the document retrieval tool when processing queries. Most existing methods use the greedy decoding strategy and directly compare the true value y i with the tool output . Although the greedy decoding strategy is simple and efficient, due to its search horizon limitation, it often fails to find the optimal solution. Therefore, adopting the beam sampling strategy, by expanding the search space and exploring more possible candidate answers, may produce better results. Specifically, input the query x i into the document retrieval tool set K times to obtain the predicted output each time Define the score of each retrieval tool t j on the query x i as: If the predicted output If correct, the evaluation value is 1, otherwise 0.

[0059] Based on the above scoring method, trainable sample pairs can be constructed for contrast learning. For query x i , according to the score construct its positive document retrieval tool index set and negative document retrieval tool index set which are composed of the document retrieval tool indexes with scores, and the rest of the document retrieval tool indexes form Considering two-way contrast learning, similarly, for document retrieval tool t j , construct positive and negative query index sets according to the query samples it processes. The positive query index set is composed of the query indexes with scores, and the negative query index set is composed of the query indexes with scores.

[0060] S3. Two-way contrast learning training (using the positive and negative retrieval tool index sets and the positive and negative query index sets to jointly train the router), match the sub-question list with the document retrieval tool list, and implement two-way contrast learning with each other as the anchor points to train the router (encoder) to comprehensively capture the association between the query and the document retrieval tool and improve the matching accuracy.

[0061] Optionally, for each query, to avoid the problem that the tool selection distribution tends to be uniform, this application uses the contrast learning method to train the router to equally select the document retrieval tool with the best performance. At the same time, to improve the training stability, a two-way contrast learning scheme is designed. This scheme includes two core objectives: contrast learning with the query as the anchor point and contrast learning with the document retrieval tool as the anchor point. Specifically, contrast learning with the query as the anchor point aims to pull the query representation closer to the positive document retrieval tool index set and away from the negative document retrieval tool index set; while contrast learning with the document retrieval tool as the anchor point expects to pull the document retrieval tool closer to the positive query index set and away from the negative query index set.

[0062] See Figure 2 , the specific steps of training the router using the contrast learning method are as follows:

[0063] 1) Input the query x i into the encoder ε, and the encoded representation is x i = ε(x i ; w), where w is a learnable parameter;

[0064] 2) The name and abstract of the document retrieval tool (N j , D jAfter connection, input it into the encoder ε to generate the vector representation t of the document retrieval tool j = ε(t j ; w), where w is also a learnable parameter.

[0065] 3) Based on the above encoding, the contrastive learning loss function with the query x i as the anchor point is defined as follows:

[0066]

[0067] where sim(·, ·) represents the cosine similarity, t + is the positive document retrieval tool representation, and t - is the negative document retrieval tool representation.

[0068] 4) Similarly, the contrastive learning loss function with the document retrieval tool t j as the anchor point is defined as follows:

[0069]

[0070] where sim(·, ·) represents the cosine similarity, x + is the positive query representation, and x - is the negative query representation.

[0071] 5) Learn the router by minimizing the weighted loss of the two contrastive losses in 3) and 4) where λ is a hyperparameter used to balance

[0072]

[0073] and and

[0074] S4. Cost reduction allocation inference. In the inference stage, send the new retrieval problem and the list of retrieval tools into the encoder for encoding respectively, calculate the similarity between them, select the optimal subset of document retrieval tools by combining the similarity scores and cost-benefit, and use the execution results of the tools to assist the large model in answering, reducing the consumption of computing resources and achieving the balance of performance and cost.

[0075] To ensure performance while improving efficiency, the inference process is modeled as an integer linear programming problem to reduce the computational cost. In the set τ of document retrieval tools, the time required for each tool to call and answer is defined as the tool call cost The cost reduction allocation inference strategy aims to ensure that the average similarity between the new retrieval problem and all document retrieval tools in the set of document retrieval tools exceeds a predefined threshold, while minimizing the average call cost. The specific steps are as follows:

[0076] 1) During the inference stage, for each test query x′ i , input it into the router Calculate its similarity score with each document retrieval tool

[0077] 2) The optimization objective of the cost reduction allocation inference strategy is:

[0078]

[0079] where respectively represent whether to assign the query x′ i to the document retrieval tool t j . Based on the constraint condition Solve the objective function to obtain the optimal tool allocation plan.

[0080] Based on the allocation plan in 2), merge the results called by the optimal tool subset, and input it together with the original query problem x′ i into the LLM to obtain the final retrieval-enhanced answer.

[0081] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present application or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present application shall be covered by the protection scope of the claims of the present application.

Claims

1. A large model RAG homogeneous tool query routing method based on contrastive learning, characterized in that Including: Step 1: Collect the key information of each document in the document collection, generate the corresponding abstract, and use the abstract as a tool description to encapsulate it into an independent document retrieval tool; the key information includes the title, author, keywords, and core content. Step 2: Use beam sampling to score the question-answering performance of each document retrieval tool, distinguish the positive document retrieval tool index set and the negative document retrieval tool index set of the current retrieval question according to the score, and construct the positive query index set and the negative query index set required for training the router. Step 3: Use the positive document retrieval tool index set, the negative document retrieval tool index set, the positive query index set, and the negative query index set to jointly train the router to capture the association between the query and the document retrieval tool; implement two-way contrastive learning with the query and the document retrieval tool as the anchors respectively. Step 4: In the document retrieval tool set, define the time required to call each document retrieval tool for answering as the tool call cost; adopt a cost reduction allocation inference strategy to ensure that the average similarity between the new retrieval question and all document retrieval tools in the document retrieval tool set is greater than a preset threshold, while minimizing the average tool call cost to obtain the final retrieval enhanced answer.

2. The method according to claim 1, characterized in that The said Step 1 includes: Suppose there are n query questions, and there are m documents in the document collection altogether. For each query question x i and its answer y i together form the training set The method for processing m documents includes: Collect the key information of each document, and define the name of the document as N j , and define the generated abstract as D j , take N j as the tool name, and D j as the tool description, encapsulate the retrieval function of each document, and form a document retrieval tool set Each query in only involves a subset of the document set ; The router takes query x i as input and generates a probability distribution of the selection of m document retrieval tools.

3. The method according to claim 1, characterized in that, The said Step 2 includes: Query x i Input the document retrieval tool set K times to obtain the predicted output each time Define each retrieval tool t j On query x i The score is: If the predicted output is correct, the evaluation value is 1, otherwise it is 0.

4. The method according to claim 3, characterized in that For query x i , according to the score , construct its positive document retrieval tool index set and negative document retrieval tool index set which are composed of document retrieval tool indexes with scores, and the remaining document retrieval tool indexes form Considering two-way contrastive learning, for document retrieval tool t j , construct positive and negative query index sets according to the query samples it processes; the positive query index set is composed of query indexes with scores, and the negative query index set is composed of query indexes with scores.

5. The method according to claim 1, wherein The said Step 3 includes: Adopt a contrastive learning method to train the router and select the document retrieval tool with the optimal performance; design a two-way contrastive learning scheme; this two-way contrastive learning scheme includes contrastive learning with the query as the anchor and contrastive learning with the document retrieval tool as the anchor; the contrastive learning with the query as the anchor aims to make the query representation tend to the positive document retrieval tool index set and away from the negative document retrieval tool index set at the same time; the contrastive learning with the document retrieval tool as the anchor expects to make the document retrieval tool tend to the positive query index set and away from the negative query index set at the same time.

6. The method according to claim 5, wherein The contrastive learning with the query as the anchor and the contrastive learning with the document retrieval tool as the anchor include: Query x i Input into the encoder ε, and the encoded representation is x i = ε(x i ; w), where w are learnable parameters; The name N J of the document retrieval tool and the abstract D j are concatenated and input into the encoder ε to generate the vector representation t j = ε(t j ; w); The contrastive learning loss function anchored by the query x i is defined as: Among them, sim(·, ·) represents the cosine similarity, and t + represents the positive document retrieval tool, and t - represents the negative document retrieval tool; The contrastive learning loss function anchored by the document retrieval tool t j is defined as: where sim(·, ·) represents the cosine similarity, x + is the positive query representation, and x - is the negative query representation; By minimizing and weighted loss learning router where λ is a hyperparameter used to balance and 7. The method according to claim 1, wherein The said Step 4 includes: Step 41: In the collection of document retrieval tools The time taken by each document retrieval tool to call for an answer is defined as the tool call cost Step 42: In the LLM inference stage, each test query x′ i is input into the router to calculate the similarity score of each test query x′ i with each document retrieval tool in the set of document retrieval tools ​ Step 43: Obtain the optimal tool allocation plan based on the similarity score and the tool call cost ; Step 44: Based on the optimal tool allocation scheme, merge the results answered by all document retrieval tools and co-enter them with the test query question x' i into the LLM to obtain the final retrieval-augmented answer.

8. The method according to claim 7, wherein The said Step 43 includes: The optimization objective of the cost reduction allocation inference strategy is: Among them, respectively indicate whether to assign the query x' i to the document retrieval tool t j ; Based on the constraints Solve the objective function to obtain the optimal tool allocation plan.