Tool usability injection method based on component importance

By analyzing the hidden vector state and component importance of the model, combining hybrid LoRA adapter and full-parameter fine-tuning, the balance between tool usage capabilities and general performance of large language models is solved, and the tool usage capabilities of the model are enhanced and general performance maintained.

CN120373414APending Publication Date: 2025-07-25INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476625.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Large language models have limitations in tool usage capabilities, especially when maintaining common capabilities while enhancing tool usage capabilities, existing methods are prone to degradation of general performance and catastrophic forgetting of models.

Method used

By analyzing the incremental changes in the hidden vector state of the model and the importance of gradient-based components, a combined approach of hybrid LoRA adapter MOLoRA and full-parameter fine-tuning is used to enhance the tool usage capabilities of the model while maintaining universal performance.

Benefits of technology

Effectively enhance the tool usage capabilities of the model, while maintaining the general performance of the model, avoiding general capability decline and catastrophic forgetting caused by excessive fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373414A_ABST
    Figure CN120373414A_ABST
Patent Text Reader

Abstract

The invention provides a tool usability injection method based on component importance, and belongs to the technical field of natural language processing. Comprising the following steps: analyzing increment change of a model hidden vector state before and after training to determine change represented by a model hidden state before and after tool learning; sorting the component importance of the model by a gradient-based importance analysis method; according to the incremental change analysis result and the component importance sorting result, a strategy for enhancing the model tool capability and maintaining the general capability is adopted, and the strategy specifically comprises the steps that for important components, a mixed LoRA adapter MOLoRA is added; and carrying out all-parameter fine tuning on non-important components. According to the method, the tool use capability of the model can be effectively enhanced without excessively sacrificing the general performance. Experimental results show that the method has excellent performance on a series of evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method for injecting tool usage capabilities based on component importance. Background Art

[0002] Large language models (LLMs) have demonstrated remarkable capabilities in understanding language and generating human-like text. Despite their impressive performance, LLMs still face limitations, such as a tendency to generate hallucinated information and an inability to interact with the real physical world. To address these limitations and expand the scope of the model's capabilities beyond traditional natural language tasks, there has been an increasing interest in enabling LLMs to interact with the external environment by equipping them with various tools.

[0003] There have been many benchmark tests evaluating different aspects of tool usage, as well as a series of proposed methods for enhancing the tool usage capabilities of large models. These methods mainly utilize in-context learning or fine-tuning techniques to endow the model with the ability to call tools. Compared with in-context learning, fine-tuning is popular because it can deeply integrate task-specific knowledge into the model parameters, enhancing its adaptability to specific tasks. However, as Figure 1 shown, it has been found that excessive fine-tuning on tool learning datasets can severely weaken the broader cognitive capabilities of the model, resulting in a decline in the general capabilities of the model and catastrophic forgetting of parameter knowledge.

[0004] Tools serve the model, not the other way around. Therefore, for enhanced models, it is important to enhance tool usage capabilities while maintaining the original generality. Summary of the Invention

[0005] In the present invention, this problem is analyzed from two aspects: hidden representations and components. Specifically, from the perspective of hidden vector representations, for different tasks, when subtracting the original hidden vector representation of the model from the vector representation of the model after fine-tuning on the tool learning dataset, a significant phenomenon is observed: the vector increment obtained by subtraction shows a "co-directional displacement" phenomenon in the hidden state space, which means that there is a strong correlation between the increment directions of different tool-agnostic tasks (such as mathematics, code generation) and tool-related tasks. From the perspective of components, for different capabilities, the gradient-based importance score rankings of the linear modules of the model (referred to as linear components, simply called components in the present invention) are calculated respectively based on the corresponding datasets. It is found that the important components in the rankings play a more important role in the expression of the corresponding capabilities, while the unimportant components contribute less to this capability. At the same time, it is found that the component importance rankings of different capabilities are highly consistent. This indicates that certain components of the model are crucial in multiple general tasks, while some components have limited impact on most tasks. Further explore the impact of fine-tuning different components based on the relevant importance rankings of tool learning tasks. It is found that optimizing important components will lead to a significant decline in general performance, and at the same time, the model may not be able to fully learn the knowledge of tool invocation only by optimizing unimportant components.

[0006] The present invention proposes a method for injecting tool usage ability based on component importance, including the following steps:

[0007] Analyze the incremental changes in the hidden vector states of the model before and after training to determine the changes in the hidden state representations of the model before and after tool learning;

[0008] Sort the component importance of the model using the gradient-based importance analysis method;

[0009] According to the results of the incremental change analysis and the component importance ranking results, adopt a strategy of enhancing the tool ability of the model while maintaining the general ability. The strategy specifically includes:

[0010] For important components, add a mixed LoRA adapter MOLoRA;

[0011] For unimportant components, perform full-parameter fine-tuning.

[0012] The present invention has the following beneficial effects:

[0013] The present invention proposes a method for injecting tool usage ability based on component importance (CITI), which applies a mixed LoRA adapter to important components and adopts full-parameter fine-tuning on unimportant components to learn the knowledge of tool invocation.

[0014] In terms of the model structure, first, CITI determines the gradient-based importance scores of all linear components in the model. Second, for important linear components, a set of Mixed LoRA (MOLoRA) adapters are integrated to learn the knowledge of tool invocation. To handle the "co-directional displacement" phenomenon, a router network is designed on the MOLoRA adapter to separate tool-related and tool-unrelated inputs, and different routing weights are adopted for different types of inputs to reduce the impact on the LLM backbone network. Third, for unimportant linear components, full-parameter fine-tuning is used to make full use of more parameters.

[0015] In terms of the training method, specifically, the training process adopts a three-stage approach. In the initial training stage (router pre-training stage), the router network in MOLoRA is pre-trained to teach it to distinguish tool-related and tool-unrelated inputs. In the second stage (MOLoRA optimization stage), focus on fine-tuning the MOLoRA adapter while freezing the backbone network of the LLM. In the third stage (unimportant component optimization stage), CITI fine-tunes a small number of unimportant components in the backbone network to improve the model performance while maintaining its general capabilities. The method achieves competitive tool usage results on two tool learning datasets, API-Bank and ToolAlpaca, and preserves the general performance of the model. Experiments prove that it effectively enhances the tool usage ability while maintaining the general performance of LLMs. Description of the Drawings

[0016] Figure 1 are the results of the full-scale fine-tuning model performance change on the tool dataset;

[0017] Figure 2 are the results of the hidden state analysis experiment;

[0018] Figure 3 are the results of the component importance analysis experiment;

[0019] Figure 4 is the schematic diagram of the model structure. Detailed Implementation Manner

[0020] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, the present invention adopts the following technical solutions. The present invention has carried out the following analyses: 1. Hidden state analysis: By verifying the incremental changes in the hidden states of the model before and after training, it is found that fine-tuning the model on the tool learning dataset will result in a strong correlation between the incremental directions of different tool-agnostic tasks (such as mathematics, coding) and tool-related tasks. 2. Component perspective analysis: According to the gradient-based importance analysis method, the importance of the components of the model is ranked. By operating on components of different importance levels, the relationship between the model's capabilities and component importance is discovered. Subsequently, a method for injecting tool usage ability CITI is proposed. By adding MOLoRA to important components and fine-tuning unimportant components in full, the tool capabilities of the model are enhanced while maintaining its general capabilities.

[0021] The present invention proposes a method for injecting tool usage ability based on component importance. According to the gradient importance scores of the components, different training strategies are applied to different components to alleviate the ability conflicts caused during the fine-tuning process. For important components, a hybrid LoRA expert structure is applied. At the same time, for some unimportant components, the parameters in the backbone network of the large language model are fine-tuned while keeping other parameters unchanged, which specifically includes the following steps:

[0022] Step 1: Analyze the changes in the hidden vector state representation of the model before and after tool learning by examining the incremental changes in the hidden vector state of the model before and after training;

[0023] Step 2: Rank the importance of the components of the model according to the gradient-based importance analysis method. By operating on components of different importance levels, discover the relationship between the model's capabilities and component importance;

[0024] Step 3: Based on the analysis results of Steps 1 and 2, a CITI method is proposed. By adding MOLoRA to important components and fine-tuning unimportant components in full, enhance the tool capabilities of the model while maintaining its general capabilities.

[0025] The following will elaborate on each step of the present invention in detail.

[0026] In Step 1, the concept of the incremental change in ability (abbreviated as ) is introduced to quantify the change in the hidden vector state representation during the fine-tuning process. Specifically, for ability and instruction in the preprocessed dataset in is the original hidden vector state representation of the last word in the untrained model, while layer instruction represents the hidden vector state representation of the same word after fine-tuning. It can be expressed as:

[0027] ,

[0028] The average of the increments calculated through the dataset is considered the ability to change in the hidden state space. Then, the cosine similarity between these increments is calculated to evaluate the relationship between the ability and vector changes in the model's hidden state space:

[0029] ,

[0030] Analysis experiment: To obtain the representation, a part of the standard answer is randomly intercepted and appended to the instruction to create a "final input". For a specific task, 1000 "final inputs" are sampled from the dataset to construct the dataset and input it into the model. The vectors corresponding to the "final inputs" of different layers are considered the hidden state representation vectors . Then, the similarity between different is calculated.

[0031] In Figure 2 , the solid line represents the similarity between the tool-related increment and the general task increment , where the LLM is fine-tuned on the tool learning dataset. For better comparison, a model is trained on the code generation dataset and its increment is calculated, represented by the dashed line. Here represents a specific general ability.

[0032] It is found that there is a "co-directional change" phenomenon in the model's hidden state space. Compared with the dashed line, the change direction of the general task is positively correlated with the direction in the model's hidden state space . Assume that after fine-tuning through the tool learning dataset, the model cannot correctly distinguish tool-related and tool-unrelated inputs in the hidden state space, resulting in more similar activation states for different types of instructions, which affects the general performance of the model.

[0033] ​In step 2, the large language model is generally composed of the decoder architecture of the transformer, usually containing multiple layers, and uses gradient-based importance scoring to identify important components.

[0034] In the autoregressive language model, for a given dataset the input and the corresponding label the loss function of the optimized model is calculated as follows:

[0035] ,

[0036] where represents the complete set of model parameters. For a specific parameter set the importance score is calculated as follows:

[0037] ,

[0038] To speed up the calculation, the importance score is estimated by calculating the first derivative of the Taylor expansion in the above formula.

[0039] Analysis experiment: To examine the changes in components in terms of various general task capabilities of the model during the tool learning process, a portion of the instructions is selected from the training set as representatives of specific general capabilities . Using the above formula, the importance scores of the linear components of these data in the original model (Meta-Llama-3-8B-Instruct) are determined. To explore the impact of component importance, the following three experiments are conducted:

[0040] Experiment 1: Sort the components according to their importance scores and replace the parameter weights of the corresponding components in the model fine-tuned on the tool learning dataset API-Bank with those of these components in the original model, while keeping all other parameters unchanged. Then evaluate the performance of the replaced model.

[0041] Table 1

[0042] In Table 1, it shows the performance of the model on corresponding tasks after component parameter replacement. T-x% means replacing the top x% components in the model, D-x% means replacing the bottom x% components in the model, and Vanilla represents the original model. It is observed that tool learning affects the expression of other abilities of the model under most tasks. After the replacement operation, the performance of the model (GSM8K, HumanEval, and MT-Bench) significantly decreases compared with the original model, indicating that fine-tuning on the tool learning dataset stimulates the tool invocation ability in these components while suppressing their original inherent abilities. When high-importance components are changed, the performance impact is more significant, that is, after replacing more important weights, the performance of the model significantly decreases. In addition, with the increase of the replacement ratio, this phenomenon becomes more obvious.

[0043] Experiment 2: In addition, to deeply study the cross-task relationship of component importance rankings, the Jaccard index is applied to analyze the correlation between the importance rankings of model components. The Jaccard index can be expressed as: , used to compare two samples and for differences and similarities. Specifically, based on the importance rankings of components related to the ability , the components are evenly divided into three groups: high, medium, and low. Then the Jaccard index between different ability groups is calculated respectively. For example, represents the Jaccard index between the high-importance component groups calculated by the ability and respectively, where h represents components with high importance.

[0044] As Figure 3 shown, there is a high correlation between the importance rankings of components in different tasks. Some components play a key role in the broad expression of model abilities, while some components have limited impact on most tasks.

[0045] Experiment 3: Try to inject tool invocation knowledge into less important components because they store less tool-related abilities. Considering the component set that contains all linear modules of the network, for the tool learning dataset , calculate the component-level gradient-based importance score of component as follows:

[0046] ,

[0047] where represents the importance score of parameter , Represents a set of parameters. To better combine the importance ranking of components and the training process, according to the component importance Select 20% (top, bottom, random) of the components in the model and train them using full-parameter fine-tuning and the LoRA method while keeping other parameters frozen. Then, analyze the fine-tuning results in detail.

[0048] Table 2

[0049] As shown in Table 2, it is found that fine-tuning components with lower importance scores has the least impact on general performance, which confirms a high correlation between different capabilities of the model because it is calculated from the tool learning dataset. is calculated from the tool learning dataset.

[0050] At the same time, it is observed that fine-tuning components with lower importance may not perform well in terms of tool invocation ability. If only unimportant parameters are fine-tuned, the model cannot thoroughly understand the format of tool invocation. However, "Top" and "Random" do not have such problems. It can be seen that sometimes only fine-tuning unimportant components may not be optimal.

[0051] Based on this, the present invention proposes a component importance-based tool utilization ability injection method (CITI) to inject tool utilization ability while maintaining the general ability of the model. Based on the analysis experiment, to fully utilize the capabilities of the model, a hybrid LoRA adapter is applied to important components and some unimportant components are fine-tuned. As Figure 4 shown, first, the parameter importance of different components of the model is obtained through the component importance calculation method and data of different tasks, and then according to the importance scores and important and unimportant parameters in the model are determined. For the important parameters determined according to a hybrid LoRA adapter (MOLoRA) is added. This method mainly includes three training stages: Router Pre-training (RP), for pre-training the router network (gating function) of the hybrid LoRA adapter; MOLoRA Improvement (MI), for training the hybrid LoRA adapter and the gating function; Unimportant Components Optimization (UCO), for training the identified unimportant parameters of the model.

[0052] The design of MOLoRA and the method of adding MOLoRA include:

[0053] Based on the analysis of component importance, it is desired to fine-tune and optimize both important and unimportant components simultaneously to comprehensively master the knowledge of tool invocation. However, due to the "co-variation" phenomenon existing in the hidden state vector space, fine-tuning to optimize important components may sometimes have a negative impact on the overall performance. To mitigate this impact, a router network, i.e., a gating function is used to distinguish tool-related and unrelated inputs, and then different adapters are assigned to process them separately.

[0054] To alleviate this problem, importance scores are calculated for all components and sorted, and hybrid LoRA adapters are added to the key components with higher . A gating function with additional neurons is applied in the hybrid LoRA adapter to separate the inputs and distinguish the types of input data. Tool-related inputs are jointly processed by the LLM backbone network and the hybrid LoRA adapter. Tool-unrelated data is mainly processed autonomously by the LLM backbone network itself. Thus, the adverse effects brought by "co-variation" are effectively alleviated, and the overall adaptation performance is improved.

[0055] Meanwhile, mixing data during downstream task fine-tuning helps to mitigate the model's catastrophic forgetting and helps to avoid overfitting. To train the router network and avoid the catastrophic forgetting problem, data from other domains is mixed during the tool learning process.

[0056] For the router network, i.e., the gating function , a linear module is used to implement weight allocation. The input features are represented as , and the formula is as follows:

[0057] ,

[0058] where is the trainable parameter matrix, represents the dimension of the input, is the number of LoRA adapters, 1 represents the additional neuron (i.e., dimension) used to represent the input is the probability of tool-unrelated data. The output of the hybrid LoRA adapter is:

[0059] ,

[0060] where are the parameters in the model backbone network, is the th element of the router output, is not used in the forward process. Matrices and It is a trainable linear module of the hybrid LoRA adapter, where is the rank of the hybrid LoRA adapter.

[0061] To achieve the separation of inputs, the data is divided into different task types related and unrelated to the tool. A local balance constraint is used to ensure that the linear module focuses on its respective tasks. This constraint assigns an importance score to each neuron to control the weights of different hybrid LoRA adapters, and the importance matrix has the following rule:

[0062] ,

[0063] Here, controls the degree of imbalance between the router outputs, is the th input.

[0064] Weight the output of the importance matrix and the gating function , denoted as . The routing loss of the router network is:

[0065] ,

[0066] where, and are the variance and mean of respectively. Ensure that for data unrelated to the tool , is relatively large, and vice versa, is relatively small. Adjust the weights of the LoRA adapter according to the data type.

[0067] The non-important component identification process specifically includes:

[0068] Select components with unimportant general capabilities by averaging the importance scores of different tasks , defined as the general capability importance . For the general capability set , the formula is as follows:

[0069] ,

[0070] Rank the components of the model according to the importance , and select The lower components form a set of unimportant parameters, and these components are not frozen during the third-stage UCO training process. For unimportant components, since updating these parameters has a relatively small impact on the general ability of the model, full-parameter fine-tuning is used to stimulate the tool invocation ability of the model.

[0071] Model training strategy:

[0072] To sum up, the overall loss function of the CITI method is:

[0073] ,

[0074] where, are the parameters of the LLM backbone network, represents the expert parameters in MOLoRA, are the router network parameters, are predefined hyperparameters. represents a mixture of the tool learning dataset sampled from the general ability dataset and other instructions.

[0075] To effectively optimize the model, a three-stage training strategy is designed as described below:

[0076] The first stage RP: Importance based on component gradients Add MOLoRA. In this stage, the core LLM backbone network model parameters and MOLoRA remain unchanged, and only are trained.

[0077] The second stage MI: and are trained. Fine-tune the parameters in MOLoRA, and other parameters remain frozen.

[0078] The third stage UCO: Here, fix the parameters of the MOLoRA adapter, including and , as well as the important parameters of the LLM backbone network model in the backbone identified by importance ranking . Then continue to train the selected subset of unimportant parameters among the backbone network parameters selected by .

[0079] Experiments:

[0080] Experiments are conducted on two tool learning benchmark tests: API-Bank and ToolAlpaca. In addition, the general ability is also evaluated through experiments on the datasets GSM8K (GSM), HumanEval (HE), TriviaQA (TQA), MT-Bench (MT).

[0081] For the evaluation metrics, for the API-Bank dataset, follow the evaluation metrics proposed in the original text, including the correctness (C) of API calls and ROUGE-L (R) for testing the response quality. For the ToolAlpaca dataset, use GPT-4 to evaluate the tool invocation process in the real-world test subset and follow the original metrics: Program (P): the correctness of tool utilization by the program, Response (R): the quality of the final response, Overall (O): whether both the program and the response are correct.

[0082] Experiments were conducted using the following models: Meta-Llama-3-8B-Instruct, Phi-3-mini-128k-instruct, and Mistral-7B-Instruct-v0.2. Compare the method of the present invention with the following baselines: (1) FT: full parameter fine-tuning of the model on the tool learning dataset; (2) LoRA: LoRA adapter fine-tuning of the model on the tool learning dataset. The baseline training data was not mixed.

[0083] As shown in Tables 3 and 4, CITI demonstrated effective performance in maintaining general capabilities in both datasets.

[0084] In terms of tool utilization ability, CITI showed competitive results compared to the baselines in both datasets. It is worth noting that in the API-Bank dataset, CITI's tool invocation ability is worse than the baselines (e.g., the ROUGE-L score in Mistral). There are probably two reasons: (1) MOLoRA may not be optimal, and insufficient training data may lead to underfitting or overfitting of the LoRA adapter. (2) The evaluation metrics are not comprehensive enough. API-Bank requires the model to generate tool invocations or responses based on the conversation history, but the path for the model to obtain the final correct answer is not unique.

[0085] In addition, it is noted that in most cases, LoRA is superior to full parameter fine-tuning in terms of tool utilization and general capabilities. This may be because full parameter fine-tuning is prone to overfitting.

[0086] Table 3

[0087] Table 4

[0088] To prove the effectiveness of the method, an ablation study was conducted on the model Meta-Llama-3-8B-Instruct, and the results are shown in Tables 5 and 6. Among them, w / o MOLoRA means replacing MOLoRA with LoRA adapters, and w / o means not using additional neurons and importance matrices in the router for the normal MOLoRA. DM represents the data mixture in the training data. The experimental results show that CITI still shows advantages in general capabilities compared with FT and LoRA that perform data mixture. In addition, the contributions of each module in CITI were analyzed. Compared with the baseline, the MI and UCO modules showed obvious advantages in maintaining the general capabilities of the model respectively.

[0089] It can be found that: (1) The UCO module shows that selectively fine-tuning unimportant components can enhance the model's ability to use new tools while maintaining its original performance, especially on GSM8K in the two datasets. (2) Comparing RP + MI with w / o MOLoRA, it is found that by integrating MOLoRA instead of simply adding LoRA adapters, the decline in general capabilities can be alleviated. (3) The results of only w / o are better than w / o MOLoRA in most cases, further indicating that the MOLoRA structure is effective. In addition, it is found that some results in w / o are better than RP + MI. It is inferred that this is because the adapters absorb the general knowledge in the general instructions in the mixed training set, but they have little impact in MI because the router separates them into the frozen LLM backbone during the training phase. (4) The results of w / o RP show that router pre-training provides benefits, especially in tool utilization in ToolAlpaca and general performance in API-Bank, indicating that pre-training the router network may accelerate model convergence, especially when the training data is limited. (5) By further fine-tuning the model by applying UCF after MI, CITI can maintain or even improve its general capabilities on the two datasets compared with RP + MI. Moreover, the tool utilization ability on the ToolAlpaca dataset has been further significantly improved.

[0090] Table 5

[0091] Table 6

[0092] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for injecting tool usage capabilities based on component importance, characterized in that, It includes the following steps: Analyze the incremental changes in the hidden vector states of the model before and after training to determine the changes in the model's hidden state representations before and after tool learning; Rank the importance of the model's components using a gradient-based importance analysis method; According to the results of the incremental change analysis and the component importance ranking results, adopt a method to enhance the model's tool capabilities while maintaining general capabilities. The method specifically includes: For important components, add a mixed LoRA adapter MOLoRA; For unimportant components, perform full-parameter fine-tuning.

2. The method for injecting tool usage ability based on component importance according to claim 1, wherein In the step of analyzing the incremental changes in the hidden vector states of the model before and after training, by calculating the difference between the hidden vector state representation of the model after fine-tuning on the tool learning dataset and the original hidden vector state representation without fine-tuning, a vector increment is obtained, and the cosine similarity between the vector increments of different tasks is calculated to evaluate the relationship between the capabilities and vector changes in the model's hidden state space.

3. The method for injecting tool usage ability based on component importance according to claim 1, characterized in that In the step of ranking the importance of the model's components using a gradient-based importance analysis method, for the input and corresponding labels in a given dataset, calculate the importance score of a specific parameter set through the gradient importance formula, and this importance score is estimated by calculating the first derivative of the Taylor expansion of the loss function.

4. The method for injecting tool usage ability based on component importance according to claim 1, wherein In the step of adding a mixed LoRA adapter, design a router network to separate tool-related and tool-unrelated inputs, and use different routing weights for different types of inputs to reduce the impact on the LLM backbone network.

5. The method for injecting tool usage ability based on component importance according to claim 4, wherein The router network realizes weight allocation through a linear module, and its output is used to control the weights of the LoRA experts to achieve the separation of inputs.

6. The method for injecting tool usage ability based on component importance according to claim 1, wherein In the step of fine-tuning the unimportant components with full parameters, select the components with unimportant general capabilities by averaging the importance scores of different tasks as the set of unimportant parameters, and fine-tune these unimportant parameters during training to stimulate the model's tool invocation capabilities.

7. The method for injecting tool usage capabilities based on component importance according to claim 1, wherein The strategy of enhancing the model's tool capabilities while maintaining general capabilities also includes adopting a three-stage training method: Router pre-training stage: Pre-train the router network in MOLoRA so that it can distinguish tool-related and tool-unrelated inputs; MOLoRA optimization stage: Focus on fine-tuning the MOLoRA adapter while freezing the backbone network of the LLM; Unimportant component optimization stage: Fine-tune the unimportant components in the backbone network to improve the model performance while maintaining its general capabilities.

8. The method for injecting tool usage capabilities based on component importance according to claim 5, characterized in that For the router network, i.e., the gating function , a linear module is used to implement weight allocation, and the input features are represented as , and the formula is as follows: , Among them, is a trainable parameter matrix, represents the dimension of the input, is the number of LoRA adapters, 1 represents an additional neuron, that is dimensional, used to represent the input is the probability of tool-independent data, and the output of the mixed LoRA adapter is: , Among them, are the parameters in the model backbone network, is the th element output by the router, which is not used in the forward process. The matrices and are the trainable linear modules of the hybrid LoRA adapter, where is the rank of the hybrid LoRA adapter.

9. The method for injecting tool usage ability based on component importance according to claim 1, characterized in that, The loss function of the method for enhancing the model's tool capabilities while maintaining general capabilities is: , Among them, are the parameters of the LLM backbone network, represent the expert parameters in MOLoRA, are the router network parameters, are predefined hyperparameters, represents a mixture of the tool learning dataset sampled from the general ability dataset and other instructions.

10. The method for injecting tool usage capabilities based on component importance according to claim 1, wherein The method is applicable to a variety of large language models, including but not limited to Meta-Llama-3-8B-Instruct, Phi-3-mini-128k-instruct, and Mistral-7B-Instruct-v0.2.