Dynamic route based on large model self-interpretation and training method thereof

By employing a self-explanatory dynamic routing method based on a large model with a small number of parameters, combined with supervised fine-tuning and group relative policy optimization algorithms, the problem of accurate routing in complex routing scenarios is solved, achieving a transparent and accurate routing decision-making process that is efficient and low-cost.

CN121390288APending Publication Date: 2026-01-23SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511481497.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies lack methods that can achieve accurate and interpretable routing in complex routing scenarios, while also requiring minimal computing resources.

Method used

Self-explanatory dynamic routing is based on a large model with few parameters. It is trained by supervised fine-tuning and group relative policy optimization algorithms. The output includes intermediate inference process and routing classification results. The LoRA module is used to reduce training cost and improve accuracy.

Benefits of technology

It achieves high-accuracy routing results in specific domains, reduces deployment costs and response time, and improves user experience and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121390288A_ABST
    Figure CN121390288A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic routing based on large model self-interpretation and a training method thereof, and the training method comprises a cold start fine tuning stage: obtaining a pre-trained basic small-parameter large language model and an initialized LoRA module, carrying out the supervision fine tuning training of the basic small-parameter large language model based on a synthetic routing data set, a fine-tuned LoRA module is obtained; and an iterative optimization stage: combining the basic small-parameter large language model and the fine-tuned LoRA module to obtain a dynamic routing model. And based on the iterative optimization data set, optimizing the dynamic routing model through a group relative strategy optimization algorithm, in the optimization process, matching a downstream module according to an output result of the dynamic routing model, and calculating a reward function of the group relative strategy optimization algorithm according to an output result of the downstream module, comprising a result correctness reward function and a format correctness reward function. Compared with the prior art, the method has the advantages of high interpretability, high routing accuracy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a dynamic routing based on large model self-explanation and a training method thereof. BACKGROUND

[0002] In recent years, large language models (LLMs) have made breakthrough progress in natural language processing (NLP) question and answer systems, text summarization, dialogue systems and other tasks, making them widely used in daily conversations, academic research, business intelligence, entertainment content creation and other aspects.

[0003] However, in order to produce relatively accurate results for a variety of user questions, it is usually necessary to use a large parameter LLM to answer, but using these models not only requires a large amount of computing resources, thereby increasing service costs, but also may result in a longer response time, thereby reducing user experience. In order to alleviate this situation, smaller parameter models can be used to solve some simple or professional domain user problems, but such models usually have poor performance in answering complex or non-professional domain questions. This requires dynamic routing of user questions to the most suitable downstream model, while ensuring performance, reducing costs and improving efficiency.

[0004] The existing routing solutions mainly include: the first is rule-based, which can determine which downstream module is suitable for a user question through keyword matching, fixed evaluation indicators, similarity algorithms, etc. For example, Chinese patent CN119167051A discloses a method, device, equipment and computer readable medium of intelligent routing selection model, which includes: performing preliminary evaluation on the model resource pool through the intelligent router to obtain a first evaluation result, which together with a multi-dimensional key feature vector extracted from the user task request determines a model candidate set; performing multi-dimensional element scoring on each candidate large model, and the elements are respectively matched with element weights that are dynamically adjusted based on user demand to obtain a second evaluation result; constructing a target function, and selecting a target large model based on the target function, decision constraints and the second evaluation result for response. The advantage of this method is that it is simple to implement, but the disadvantage is that it has poor generality and poor processing effect for complex situations; the second is based on neural network classifier, which regards the routing problem as a classification problem, and directly obtains the classification result by training a simple neural network or a pre-trained language model (PLM) such as BERT (Bidirectional Encoder Representation Transformer). The advantage of this method is that it uses less resource consumption while obtaining good performance in specific scenarios through fine-tuning, but the disadvantage is that it can only directly obtain the routing result, which is not clear to the user, and is easily affected by vocabulary mismatch; the third is a method based on LLM, which uses a large-parameter general LLM to dynamically route user questions according to the description of the downstream model. The advantage of this method is that the powerful zero-shot or few-shot learning ability of LLM can obtain good generalization ability, and the intermediate text reasoning process can provide better routing result accuracy, but the disadvantage is that the large-parameter model has cost and efficiency problems, and may produce illusions for specific user questions (such as large context injection) to cause inaccurate routing results.

[0005] Therefore, there is currently a lack of a method that can achieve accurate routing in complex routing scenarios, has explainability, and has low requirements for computing resources. SUMMARY

[0006] The purpose of the present application is to overcome the defects of the prior art and provide a dynamic routing based on large model self-explanation and a training method thereof. The small parameter large model after fine-tuning is used to self-explain the routing result. The dynamic routing model constructed based on the specific domain synthetic dataset is trained by combining supervised fine-tuning and group relative strategy optimization algorithm, so that it can output text containing intermediate reasoning process, routing classification result, routing additional information and specific separator. Finally, the routing dispatcher parses the text and allocates it to the specified downstream module. The method applies the self-explanation ability of the large model to the dynamic routing task, solves the problem of poor flexibility and opaque decision-making based on rule or classifier method, and reduces the model deployment cost and user response time by fine-tuning the small parameter large model compared with using a general large parameter large model, and improves the routing accuracy in this specific domain.

[0007] The purpose of the present application can be achieved by the following technical solutions: A dynamic routing training method based on large model self-explanation, the method comprising: Cold start fine-tuning stage: Obtain a pre-trained basic small parameter large language model and an initialized LoRA module, fine-tune the basic small parameter large language model based on the synthetic routing dataset to obtain the LoRA module after fine-tuning; Iterative optimization stage: Combine the basic small parameter large language model and the LoRA module after fine-tuning to obtain a dynamic routing model; Generate an iterative optimization dataset based on the synthetic routing dataset; Based on the iterative optimization dataset, optimize the dynamic routing model by using the group relative strategy optimization algorithm. In the optimization process, the downstream module is matched according to the output result of the dynamic routing model, and the reward function of the group relative strategy optimization algorithm is calculated according to the output result of the downstream module. The reward function includes a result correctness reward function and a format correctness reward function.

[0008] The output result of the dynamic routing model includes an intermediate reasoning process and a routing classification result, and the output format is: <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> The downstream module is matched according to the routing classification result in it.

[0009] The output result of the dynamic routing model also includes routing additional information, and the output format is: <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> <extra>Routing additional information< / extra> .

[0010] If the matched downstream module supports routing additional information, the routing additional information is transmitted to the corresponding downstream module, otherwise the routing additional information is not transmitted.

[0011] If a corresponding downstream module is matched, the routing target is set to the downstream module; if a corresponding downstream module is not matched, a routing default value is used.

[0012] The result correctness reward function judges the routing result according to the target scene corresponding to the downstream module, and if the judgment is correct, the result correctness reward function is 1, and if the judgment is incorrect, the result correctness reward function is 0.

[0013] The format correctness reward function is defined as follows: if the output result of the dynamic routing model successively appears <think> 、< / think> 、 <answer> 、< / answer> , the format correctness reward function is 1, if it successively appears <extra> 、< / extra> on this basis, the format correctness reward function is recorded as 1.25 points, and in other cases, the format correctness reward function is recorded as -1 points.

[0014] The acquisition process of the synthesized routing data set is as follows: An original data set containing user questions and answer pairs is obtained, and the routing classification true results of each question are determined according to the functional boundaries of the downstream modules of the dynamic routing model and the answer conditions, or the classification true results are synthesized with the help of a large parameter LLM, to obtain a first intermediate data set containing user questions, routing classification results and answer pairs; For a specific routing scene, the following two methods or a combination of the following two methods are used to generate an intermediate reasoning process that guides correct routing classification results, to generate a second intermediate data set containing user questions, intermediate reasoning processes, routing classification results and answers: manually design a fixed reasoning process solution, and the intermediate reasoning process is obtained by a deterministic method; use a large parameter reasoning model to synthesize the intermediate reasoning process; Based on the second intermediate data set, the intermediate reasoning process and the routing classification result are spliced to obtain a routing reasoning example, and finally a synthesized routing data set containing user questions and routing reasoning examples is obtained.

[0015] A dynamic routing method based on large model self-explanation, the method comprising the following steps: Obtain user input, downstream module capability description, and format according to the preset prompt word template, and input the dynamic routing model trained based on the training method; According to the output result of the dynamic routing model, a corresponding downstream module is matched; Use the downstream module to output a dynamic routing result.

[0016] The obtained information also includes other constraint information, and the other constraint information includes time constraint conditions.

[0017] Compared with the prior art, the present application has the following beneficial effects: (1) The present application uses LLM for dynamic routing. Compared with rule-based or classifier-based single routing results, the routing results can be self-explained by LLM, thereby transparently showing the routing decision process to the user, improving user trust and facilitating the understanding of the decision process by the debugging personnel.

[0018] (2) The present application uses small parameter LLM for dynamic routing. Compared with using general large parameter LLM, the deployment cost can be reduced, and the routing inference time can be reduced, thereby improving user experience.

[0019] (3) The present application uses LLM based on sufficient knowledge in a specific field for dynamic routing. The advantages of LLM in natural language semantic understanding can produce more accurate and flexible routing results. Through the fine-tuning method, high accuracy of dynamic routing for various user problems in a specific scene can be obtained, and a small amount of computing resources can achieve low response time deployment effect, thereby potentially improving the overall effect of the target scene system.

[0020] (4) The requirement output format of LLM in the dynamic routing task is innovated, and the <think>< / think> thinking area stimulates the thinking ability of LLM, and the <answer>< / answer> answer area facilitates the analysis of dynamic routing dispatchers, and the <extra>< / extra> additional information can potentially improve the execution effect of the downstream module.

[0021] (5) In the dynamic routing task, the present application uses supervised fine-tuning to enable LLM to follow instructions, and reduces the training cost through LoRA technology; uses the GRPO algorithm in reinforcement learning to train LLM to output flexible routing additional information, and further enhances routing accuracy through system result-oriented enhancement; combines supervised fine-tuning and reinforcement learning to fine-tune LLM, which can improve the overall effect of the target scene system compared with using any single method. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a dynamic routing method flowchart of the present application; Figure 2 is a dynamic routing training method flowchart of the present application; Figure 3 is a statistical chart of the change of routing classification accuracy with the number of iterations in the dynamic routing model training process of the present application; Figure 4 is a statistical chart of the change of routing average time and full-process average time with the number of iterations in the dynamic routing model training process of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the protection scope of the present application.

[0024] Unless otherwise defined, technical terms or scientific terms used in the present application should be understood as their common meanings to those of ordinary skill in the art to which the present application pertains. The terms "one", "a", "an", "the", and similar terms in the present application do not denote a singular number or quantity but include a plural number or quantity unless otherwise defined. The terms "comprise", "comprising", "include", "including", "have", and "having" in the present application are intended to cover the inclusions thereof without limitation; for example, a process, method, system, product, or device that includes a list of steps or modules (units) is not limited to the listed steps or units, but can further include other steps or units not listed or can further include other steps or units inherent to such a process, method, product, or device. The terms "connect", "connected", "couple", and similar terms in the present application are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The term "plurality" in the present application refers to two or more. The term "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects. The terms "first", "second", "third", and the like in the present application are only to distinguish similar objects, and do not represent a specific order of the objects.

[0025] Embodiment 1 This embodiment first provides a dynamic routing method based on large model self-explanation, as shown in the following figure, which includes the following steps: Figure 1 Step S101, obtaining user input, downstream module capability description, and optional other constraint information (such as time constraint condition, etc.).

[0026] Step S102, after the information obtained in step S101 is formatted according to the preset prompt word template, it is input into the trained dynamic routing model, and the model is required to output information in the format of "dynamic routing result". <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> <extra>Routing additional information< / extra>

[0027] Step S103, the dynamic routing scheduler parses the output result of the dynamic routing model, identifies <answer>< / answer> ​​The dynamic routing result inside is matched with the corresponding downstream module; if the dynamic routing result matches one of the downstream modules, the routing target is set as the downstream module, otherwise, the routing default value is used; optionally, if the downstream routing module supports <extra>< / extra> the information in <extra>< / extra> , the information in <extra>< / extra> is transmitted to the downstream module, otherwise, the additional information is not transmitted.

[0028] Step S104, the selected downstream routing module is executed, and the number of downstream modules is determined by the target scene system.

[0029] Step S105, the results of the downstream modules in step S104 are reduced to obtain the routing result.

[0030] The embodiment also provides a dynamic routing training method based on large model self-explanation, which is used for training a dynamic routing model, and the method comprises the following steps: Cold start fine-tuning stage: obtain a pre-trained basic small parameter large language model LLM (typical parameter quantity is 3B, 7B, etc.) and an initialized LoRA (Low-Rank Adaptation, low rank adaptation) module, and prepare a synthetic routing data set according to a pre-defined system prompt word template and an answer template with explanation, perform SFT (Supervised Fine-Tuning, supervised fine-tuning) training on the basic small parameter large language model based on the synthetic routing data set, and obtain a fine-tuned LoRA module; Iterative optimization stage: combine the basic small parameter large language model and the fine-tuned LoRA module to obtain a dynamic routing model; generate an iterative optimization data set based on the synthetic routing data set; based on the iterative optimization data set, the dynamic routing model is optimized by a group relative strategy optimization algorithm, and in the optimization process, the downstream module is matched according to the output result of the dynamic routing model, and the reward function of the group relative strategy optimization algorithm is calculated according to the output result of the downstream module.

[0031] As shown in Figure 2 , the training process specifically comprises the following steps: Step S201, for an original data set containing (user question, answer) pairs, determine the routing classification true result of each question according to the function boundary of the downstream module of the dynamic routing model and the answer condition, or generate a first intermediate data set containing (user question, routing classification result, answer) by assisting the synthesis of a higher confidence classification true result through a large parameter LLM.

[0032] Step S202, for the first intermediate data set generated in step S201, an intermediate reasoning process capable of leading to correct routing classification results is generated for a specific routing scenario by one of the following two methods or in combination: a fixed reasoning process solution can be designed manually, and the intermediate result is obtained by a deterministic method; or a reasoning model with a large number of parameters is used to synthesize such an intermediate reasoning process, to generate a second intermediate data set containing (user question, intermediate reasoning process, routing classification result, answer).

[0033] Step S203, for the second intermediate data set generated in step S202, the intermediate reasoning process and the routing classification result are spliced by " to obtain a routing reasoning example, and finally a synthetic routing data set containing (user question, routing reasoning example) is obtained, and is divided into a training supervision fine-tuning data set and a verification supervision fine-tuning data set according to a quantitative proportion. <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer>

[0034] Step S204, add LoRA modules to multiple layers of the basic small-parameter LLM, and initialize the parameters of the LoRA modules to obtain a to-be-trained LLM integrated with the LoRA and the basic small-parameter LLM.

[0035] Step S205, set the system prompt word as a combination of the following three parts: description information of the downstream module, requirement to obtain a module routing result through reasoning, and output in the format of ". Format the system prompt word, the user question text based on the corresponding LLM dialogue template, and input the to-be-trained LLM in step S204 to obtain the reply of the LLM. Train the LoRA module on the training supervision fine-tuning data set through cross-entropy loss. <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer>

[0036] Step S206, according to step S205, the training supervision fine-tuning data set is trained for multiple rounds, and after each round of training, the routing accuracy is verified on the verification supervision fine-tuning data set obtained in step S203, and the checkpoint of the LoRA module with the highest routing accuracy is selected as the LoRA module after the cold start fine-tuning stage.

[0037] Step S207, combine the basic small-parameter LLM with the LoRA module fine-tuned in step S206 to obtain a cold start fine-tuning stage dynamic routing model.

[0038] ​​In step S208, the first intermediate dataset containing (user questions, route classification results, and answers) obtained in step S201 is used, or optionally combined with user questions collected additionally based on the online environment, the labeled answers, and the first intermediate dataset synthesized in step S201, to obtain the iterative optimization dataset. The iterative optimization dataset is then divided into a training iterative optimization dataset and a validation iterative optimization dataset according to a certain ratio.

[0039] Step S209: Install the dynamic routing module obtained in step S207 into the system of the target scenario, and set the system prompt to a combination of the following three parts: description information of the downstream module, request to obtain the model routing result through inference and provide appropriate additional information, and in accordance with " <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> <extra>Routing additional information< / extra> The output is in the format of "". One question from the training iteration optimization dataset obtained in step S208 is taken as the user's question. The intermediate results are obtained through the system front-end module. The system prompt words, user question text and intermediate results from the system front-end module are formatted based on the corresponding LLM dialogue template and input into the dynamic routing module to obtain the module's response. The response is then input into the system back-end module to finally obtain the overall routing result.

[0040] In step S210, the reference dynamic routing module corresponding to the reference dynamic routing model obtained in step S207, which is fixed as the inference state rather than the training state, is installed into the reference target scene system, and the reference overall result is obtained using the same process as in step S209.

[0041] Step S211: Based on the overall result of step S209, the reference overall result of step S210, and the true answer, the sum of the result correctness reward function and the format correctness reward function is used as the value of the reward function. The result correctness reward function determines the routing result based on the target scenario system, classifying it as correct (1 point) or incorrect (0 points). The format correctness reward function is defined as follows: the output sequentially contains "..." <think> ”、"< / think> " <answer> ”、"< / answer> "This function is worth 1 point, and further points will be awarded based on this." <extra> ”、"< / extra> "This function is scored as 1.25 points; otherwise, it is scored as -1 point."

[0042] Step S212: Based on the reward function value in step S211, the dynamic routing model in step S209 is trained using the GRPO (Group Relative Preference Optimization) algorithm to obtain the optimized dynamic routing model.

[0043] Step S213, repeat steps S209-S212 a fixed number of times using different data in the training iterative optimization dataset obtained in step S208, and then replace the corresponding model in step S209 with the optimized dynamic routing model, and keep the reference dynamic routing model in step S210 unchanged.

[0044] Step S214, repeat step S213 a fixed number of times to obtain an optimized dynamic routing model, and verify the routing accuracy of the optimized dynamic routing model on the verification iterative optimization dataset.

[0045] Step S215, repeat step S214 multiple times, and use the optimized dynamic routing model with the highest routing accuracy to obtain the corresponding optimized dynamic routing model as the final dynamic routing module.

[0046] The downstream module mentioned in this embodiment can refer to a single model, or an agent module composed of multiple models cooperating with each other.

[0047] Embodiment 2 This embodiment is based on embodiment 1, and provides a routing method of NL2SQL according to difficulty.

[0048] NL2SQL (Natural Language to Structured Query Language) aims to convert natural language into SQL statements for a given database. Since NL2SQL can be used to reduce the use threshold of widely used relational databases, it has broad application prospects. However, due to the variety of natural language inputs and the complex database relationships, it is challenging to achieve balanced high accuracy and low response speed for NL2SQL. To this end, one solution is to divide the generated SQL according to the difficulty, and use different model configurations of different sizes to solve different difficulties, so as to reduce the cost and improve the efficiency while ensuring the performance.

[0049] The target scene system is an NL2SQL reference system, which includes a dynamic routing module; the system pre-module of the dynamic routing module is a schema linking module, which will use Qwen2.5-72B-Instruct to require the LLM to generate relevant schema information (table, column, value) according to the given database schema and user question; the system post-module of the dynamic routing module is a SQL error correction module, which will use Qwen2.5-72B-Instruct to require the LLM to correct potential SQL errors according to the given database schema, user question, and dynamic routing module generated SQL, and feed back the correct SQL to the user.

[0050] As Figure 1As shown, for a user question, three downstream models are set up to solve respectively: (1) simple question, denoted as EASY, the final SQL of which does not need JOIN and nested query; (2) non-nested question, denoted as NON-NESTED, the final SQL of which needs JOIN but does not need nested query; (3) nested question, denoted as NESTED, the final SQL of which needs nested query. Here, the nested query refers to a query containing at least one of the keywords such as INTERSECT, UNION, EXCEPT, IN, NOT IN, etc.

[0051] The input information of step S101 will contain the user question, the above-mentioned routing criteria assigned to each downstream model, and the intermediate result of the NL2SQL pre-module (pattern linking module).

[0052] Step S102 is a dynamic routing LLM based on Qwen2.5-3B fine-tuning. After the information in step S101 is formatted through the prompt template preset by Qwen2.5-3B, the information is input into the LLM to obtain the dynamic routing output.

[0053] Step S103 is a dynamic routing dispatcher that parses the output of S102 to identify the dynamic routing result in <answer>< / answer> If the dynamic routing result matches one of EASY, NON-NESTED, and NESTED, the routing target is set to the corresponding downstream module, otherwise the routing default value NESTED is used. The fine-tuned dynamic routing LLM should be able to output a list of sub-questions that are helpful for the NESTED module in <extra>< / extra> .

[0054] The three downstream modules EASY, NON-NESTED, and NESTED in step S104 are all implemented using Qwen2.5-72B-Instruct LLM, with different prompts: EASY requires the LLM to answer directly according to the pattern linking and user question, NON-NESTED requires considering the join table problem on the basis of the EASY prompt, and NESTED requires answering sub-questions one by one and finally obtaining the merged SQL on the basis of NESTED.

[0055] Step S105 will reduce and integrate the results of the three downstream modules and pass them to the post-module (SQL error correction module).

[0056] As shown in Figure 2 , for the aforementioned dynamic routing model, the specific process of training the dynamic routing model using the consistent Qwen2.5-3B as the basic small-parameter LLM is as follows: Step S201: For the original dataset containing (user questions, SQL answers), totaling 6500 entries, label the actual route classification results according to the aforementioned EASY, NON-NESTED, and NESTED difficulty classification criteria based on the SQL answers.

[0057] Step S202: For the dataset (user question, route classification result, SQL answer) generated in step S201, manually design the process of exporting this route classification result: which tables are needed to solve the user question and whether JOIN is needed, whether nested queries using keywords such as INTERSECT, UNION, EXCEPT, IN, NOT IN are needed, and therefore what is the classification of the user question.

[0058] Step S203: For the dataset (user question, intermediate reasoning process, route classification result, SQL answer) generated in step S202, connect the intermediate reasoning process and the route classification result through " <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> "The concatenation process yields routing reasoning examples, resulting in a supervised fine-tuning dataset of (user questions, routing reasoning examples). 100 examples are selected as the validation supervised fine-tuning dataset, and the remainder is used as the training supervised fine-tuning dataset."

[0059] Step S204: Add the following to the q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj modules of the Qwen2.5-3B basic small parameter LLM. =8、 =32, LoRA module with loradropout=0.1. Here, LoRA is a parameter-efficient fine-tuning method for the weight matrix of one layer of a pre-trained model. LoRA introduces two small matrices and , where rank Much smaller than The product of these two small matrices It's a low-rank matrix. Then use... As the weight of this layer, here It is the scaling intensity factor that controls low-rank updates. During fine-tuning, it maintains... Unchanged, only training and The parameters of the two matrices are reduced, thus greatly reducing computational and storage costs. To prevent overfitting, some input elements are randomly zeroed with the probability of `lora_dropout` during training. The LoRA module is initialized using the following method: the matrix... Zero initialization, matrix Kaiming initialization, i.e., each element in the matrix is sampled from a uniform distribution, where is taken. After initialization, the LoRA is integrated with the base small parameter LLM to obtain the LLM to be trained.

[0060] Step S205, set the system prompt word as the combination of the following parts: the routing classification standard of the foregoing step S104, the database schema, the requirement to obtain the module routing result through reasoning, and output in the format of "The answer to the question of <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> Step S203". Take each question in the training supervision fine-tuning data set obtained in step S203 as a user question, format the system prompt word, the user question text based on the corresponding LLM dialogue template, and input the LLM to be trained in step S204 to obtain the reply of the LLM. Train the LoRA module on the training supervision fine-tuning data set through the cross-entropy loss. The cross-entropy loss is , where represents the training supervision fine-tuning data set, is a piece of input corresponding to the data set, is the probability of the large model output when the prefix is . The loss measures the difference between the probability distribution predicted by the model and the probability distribution of the true label.

[0061] Step S206, perform multi-round training on the training supervision fine-tuning data set according to step S205. After each round of training, verify the routing accuracy on the verification supervision fine-tuning data set obtained in step S203, and select the LoRA module checkpoint with the highest routing accuracy as the LoRA module after the cold start fine-tuning stage. The LoRA module checkpoint refers to a storage file containing the LoRA weight matrix after training in step S205.

[0062] Step S207, combine the Qwen2.5-3B base small parameter LLM with the LoRA module fine-tuned in step S206 to obtain the dynamic routing model in the cold start fine-tuning stage.

[0063] Step S208, use the (user question, routing classification result, SQL answer) data set obtained in step S201, combine a small amount of user questions collected in an online environment, manually annotate the SQL answer according to user feedback, about 100, and synthesize the (user question, routing classification result, SQL answer) data set through the process of step S201 to obtain an iterative optimization data set. The first 100 of the iterative optimization data set are used as the verification iterative optimization data set, and the rest are used as the training iterative optimization data set.

[0064] Step S209, install the dynamic routing module obtained in step S207 into the aforementioned NL2SQL reference system, set the system prompt as the combination of the following parts: the routing classification criteria in step S104, the database schema, the requirement to obtain the module routing result through reasoning, the need to output additional sub-questions if it is a NESTED classification, and output in the format of "SELECT <think>Intermediate reasoning process< / think> <answer>Routing classification result< / answer> <extra>Additional sub-problems for NESTED classification< / extra> ". Take one question in the training iteration optimization data set obtained in step S208 as a user question, pass it through the system pre-module to obtain the relevant intermediate result, format the system prompt, user question text, and system pre-module link module intermediate result into the dynamic routing module based on the LLM dialogue template format of Qwen2.5-3B, obtain the reply of the module, input it into the system post-SQL error correction module, and finally generate the predicted SQL.

[0065] Step S210, install the reference dynamic routing module corresponding to the reference dynamic routing model obtained in step S207 fixed as the reasoning state instead of the training state into the reference target scene system, and use the same process as step S209 to obtain the SQL generated by the reference model.

[0066] Step S211, based on the predicted SQL in step S209, the reference SQL in step S210, and the true SQL answer, take the sum of the result correctness reward function and the format correctness reward function as the value of the reward function. Wherein the result correctness reward function is defined as if the execution result of the predicted SQL is consistent with the execution result of the true SQL answer, and otherwise; the format correctness reward function is defined as if "SELECT <think> ”、"< / think> " and "FROM <answer> ”、"< / answer> " appear in the output in turn, if "WHERE <extra> ”、"< / extra> " continues to appear on this basis, and otherwise.

[0067] Step S212, according to the reward function value , the dynamic routing model in step S209 is trained using a GRPO (Group Relative Preference Optimization) algorithm to obtain an optimized dynamic routing model. The GRPO algorithm is a reinforcement learning algorithm that aims to optimize the strategy by maximizing the advantage of a group of strategies relative to other strategies to improve the generalization ability and robustness of the model in different groups or scenarios.

[0068] Step S213: Repeat steps S209-S212 using different data in the training iteration optimization dataset obtained in step S208, repeat 3 times, then replace the corresponding model in step S209 with the optimized dynamic routing model, and keep the reference dynamic routing model in step S210 unchanged.

[0069] Step S214: Repeat step S213 200 times to obtain an optimized dynamic routing model checkpoint, and verify the routing accuracy of the optimized dynamic routing model checkpoint on the verification iteration optimization dataset.

[0070] Step S215: Repeat step S214 multiple times, which can be iterated, and use the routing accuracy of the optimized dynamic routing model checkpoint with the highest routing accuracy to obtain the corresponding optimized dynamic routing model as the final optimized dynamic routing module.

[0071] Figure 3 and Figure 4 To analyze the changes in routing classification accuracy, routing average time, and full-process average time with the number of iterations on the verification iteration optimization dataset based on the aforementioned training routing model process, three methods are used: the "result correctness reward" method uses the cold start fine-tuning stage model as the starting point and only uses the result correctness reward function as the reward function to train the model using GRPO; the "result correctness reward + format correctness reward" method uses the cold start fine-tuning stage model as the starting point and uses the result correctness reward function and the format correctness reward function as the reward function to train the model using GRPO; the "baseline large model" method is Qwen2.5-3B without fine-tuning, as it does not follow the output format and has poor classification effect, it is the starting point of the cold start fine-tuning stage. Figure 3 The classification accuracy of the routing module output is counted, Figure 4The average time consumed by the routing module and the average time consumed by the entire process is counted, and it is shown that step S215 repeats step S214 for 17 times, and each dynamic routing model checkpoint verifies the verification result on the verification iteration optimization dataset.

[0072] By Figure 3 And Figure 4 As can be seen, experiments show that by combining supervised fine-tuning and reinforcement learning, the LLM in the system can output correct formats and more correct routing classification results compared to the baseline large model with little change in execution accuracy. The "result correctness reward" method improves the starting classification accuracy of the cold start fine-tuning stage model by 5.9%, and the "result correctness reward + format correctness reward" method improves the starting classification accuracy of the cold start fine-tuning stage model by 8.5%. Moreover, it consumes less time. The "result correctness reward" method reduces the average routing time of the cold start fine-tuning stage model by 22.1%, and the "result correctness reward + format correctness reward" method reduces the average routing time of the cold start fine-tuning stage model by 36.9%. The "result correctness reward" method reduces the average time of the entire process of the cold start fine-tuning stage model by 2.7%, and the "result correctness reward + format correctness reward" method reduces the average time of the entire process of the cold start fine-tuning stage model by 6.6%. As can be seen, the LLM trained by the method proposed in the present application can have better routing effect.

[0073] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A dynamic routing training method based on large model self-explanation, characterized in that, The method comprises: a cold start fine-tuning stage: obtain a pre-trained base small parameter large language model and an initialized LoRA module, perform supervised fine-tuning training on the base small parameter large language model based on a synthetic routing dataset, and obtain a fine-tuned LoRA module; an iterative optimization stage: combine the base small parameter large language model and the fine-tuned LoRA module to obtain a dynamic routing model; generate an iterative optimization dataset based on the synthetic routing dataset; based on the iterative optimization dataset, optimize the dynamic routing model by using a group relative strategy optimization algorithm, in the optimization process, match a downstream module according to an output result of the dynamic routing model, calculate a reward function of the group relative strategy optimization algorithm according to an output result of the downstream module, and the reward function comprises a result correctness reward function and a format correctness reward function.

2. The dynamic routing training method based on large model self-explanation according to claim 1, characterized in that, The output result of the dynamic routing model includes an intermediate inference process and a routing classification result, and the output format is: <think>intermediate reasoning process< / think> <answer>routing classification result< / answer> According to the routing classification result therein, a downstream module is matched.

3. The dynamic routing training method based on large model self-explanation of claim 2, wherein, The output result of the dynamic routing model further comprises routing additional information, and the output format is: <think>intermediate reasoning process< / think> <answer>routing classification result< / answer> <extra>routing additional information< / extra> .

4. The dynamic routing training method based on large model self-explanation of claim 3, wherein, if the matched downstream module supports routing additional information, the routing additional information is transmitted to the corresponding downstream module, otherwise, the routing additional information is not transmitted.

5. The dynamic routing training method based on large model self-explanation according to claim 1, characterized in that, if the corresponding downstream module is matched, the routing target is set as the downstream module; if the corresponding downstream module is not matched, a routing default value is used.

6. The dynamic routing training method based on large model self-explanation according to claim 1, characterized in that, the result correctness reward function determines the routing result according to the target scene corresponding to the downstream module, if the determination is correct, the result correctness reward function is 1, and if the determination is incorrect, the result correctness reward function is 0.

7. The dynamic routing training method based on large model self-explanation of claim 3, wherein, The correctness reward function is defined as: if the output results of the dynamic routing model successively appear <think> 、< / think> , <answer> 、< / answer> , the correctness reward function is 1, if on this basis successively appear <extra> 、< / extra> , the correctness reward function is recorded as 1.25 points, and in other cases, the correctness reward function is recorded as-1 points.

8. The dynamic routing training method based on large model self-explanation of claim 1, wherein, the acquisition process of the synthetic routing dataset is as follows: obtain an original dataset comprising user questions and answer pairs, determine the routing classification true result of each question according to the functional boundaries of the downstream modules of the dynamic routing model and the answer conditions, or synthesize the classification true result by using a large parameter LLM, obtain a first intermediate dataset comprising user questions, routing classification results and answer pairs; for a specific routing scene, the following two methods are used to generate an intermediate reasoning process that guides correct routing classification results, or a combination of the two methods is used to generate an intermediate reasoning process that guides correct routing classification results: manually design a fixed reasoning process solution, and the intermediate reasoning process is obtained by a deterministic method; use a large parameter reasoning model to synthesize the intermediate reasoning process; based on the second intermediate dataset, the intermediate reasoning process and the routing classification result are spliced to obtain a routing reasoning example, and finally a synthetic routing dataset comprising user questions and routing reasoning examples is obtained.

9. A dynamic routing method based on large model self-explanation, characterized in that, The method comprises the following steps: obtain user input, downstream module capability description, and format the input according to the preset prompt word template, and input the dynamic routing model trained by the method of any one of claims 1-8; match the corresponding downstream module according to the output result of the dynamic routing model; output the dynamic routing result by using the downstream module.

10. The dynamic routing method based on large model self-explanation of claim 9, wherein, The obtained information also includes other constraint information, and the other constraint information includes time constraint conditions.

Citation Information

Patent Citations

  • Intelligent routing model selection method, device and equipment and computer readable medium

    CN119167051A