Gradient approximation-based large language model fusion method and system
By evaluating parameter importance through gradient approximation, pruning and fusing large language models, the problem of model fusion without the need for extensive data training is solved, improving computational efficiency and data processing accuracy.
Patent Information
- Application Number
- CN202511079823.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-02
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to effectively fuse large language models without requiring extensive training with high-quality data, resulting in high computational overhead and insufficient data processing accuracy.
The importance of parameters is evaluated using gradient approximation. With a small number of samples, the activation function and pruning rate are constructed by comparing the model parameters under clean and corrupted running modes. The model is then pruned and fused to fine-tune it, resulting in the final model.
It significantly improves the ability to locate parameters most relevant to specific tasks, reduces computational overhead, and maintains the reliability and accuracy of data processing results.
Smart Images

Figure CN120974411A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and more particularly to a large language model fusion method and system based on gradient approximation. BACKGROUND
[0002] At present, in recent years, large language models (LLM) have made significant breakthroughs in various NLP tasks, and many checkpoints for specific task fine-tuning have been publicly provided. These fine-tuned models integrate high-quality and valuable specific task information on the basis of pre-trained models. However, due to security and privacy issues, obtaining supervised specific task data is still a difficult task, and training LLM on large-scale datasets is also very costly. Therefore, it is of great practical significance to propose a method that can directly utilize the capabilities of these fine-tuned models with minimal or no additional training requirements.
[0003] However, the Linear Model Connectivity theory shows that fine-tuned models derived from the same pre-trained model, even with different hyperparameter settings, are usually located in the same low-error region. By combining these models in the parameter space, a model with better generalization performance can be identified. Based on this theory, many studies have begun to explore the model fusion problem: combining multiple homologous model parameters to obtain better generalization performance without the need for complete fine-tuning of LLM on large amounts of high-quality data, while improving data processing accuracy.
[0004] Therefore, how to provide a large language model fusion method that can solve the above problems is a problem that those skilled in the art need to solve. SUMMARY
[0005] Therefore, the present application provides a large language model fusion method and system based on gradient approximation, which significantly improves the ability to locate the most relevant parameters for specific tasks using a small amount of sample data, significantly reduces the computational overhead, and maintains the reliability of the data processing results.
[0006] To achieve the above purpose, the present application adopts the following technical solutions:
[0007] A large language model fusion method based on gradient approximation, comprising the following steps:
[0008] Obtain language text data samples;
[0009] Construct a fine-tuned model set and a pre-trained model, wherein the fine-tuned model set includes a plurality of fine-tuned models, and define the difference value between the fine-tuned model and the pre-trained model as a task vector;
[0010] The fine-tuning model set and the pre-training model are respectively run in multiple manners, and parameter importance evaluation results of the task vector are calculated based on gradients in different running manners;
[0011] An activation function is constructed, and a pruning rate is determined in combination with the activation function and the parameter importance evaluation results;
[0012] Pruning and fusion of the fine-tuning model set are completed according to the pruning rate, and a final fine-tuning model is obtained;
[0013] The fine-tuning model is used to process the language text data sample, and a corresponding processing result is obtained.
[0014] Preferably, the specific processing process of calculating the parameter importance evaluation results of the task vector based on the gradients in different running manners includes:
[0015] An operation sample is obtained;
[0016] Multiple running modes are constructed, and the operation sample is input into the fine-tuning model for prediction in multiple running modes, and multiple corresponding prediction output results are obtained;
[0017] Differences corresponding to the multiple prediction output results and gradients corresponding to the task vector are calculated;
[0018] The corresponding parameter importance evaluation results are determined according to the operation sample, the differences and the gradients.
[0019] Preferably, the specific processing process of obtaining the multiple corresponding prediction output results includes:
[0020] The running modes include a clean running mode and a damaged running mode;
[0021] In the clean running mode, the operation sample is input into the fine-tuning model for prediction, and a corresponding true probability is obtained;
[0022] In the damaged running mode, a certain partition of the fine-tuning model is replaced with a noise part in the pre-training model, and the operation sample is input into the modified fine-tuning model for prediction, and a corresponding prediction output is obtained.
[0023] Preferably, the specific processing process of determining the corresponding parameter importance evaluation results according to the operation sample, the differences and the gradients includes:
[0024] The inner product of the operation sample and the gradient is calculated;
[0025] The parameter importance evaluation result corresponding to the inner product, the difference and the gradient is determined.
[0026] Preferably, the specific process of determining the pruning rate in combination with the activation function and the parameter importance evaluation result comprises:
[0027] The specific expression of the pruning rate is constructed by the activation function, and the specific expression is:
[0028]
[0029] In the formula, Pruning rate, λ represents the initial pruning rate, tanh represents the activation function, τ1 represents the temperature coefficient of tanh, and ∈ represents the threshold.
[0030] Preferably, the specific process of obtaining the final fine-tuning model comprises:
[0031] The task vector of the fine-tuning model is pruned by the pruning rate, and the model is fused by the parameter importance evaluation result, and the specific expression is:
[0032]
[0033] In the formula, τ2 is the temperature parameter of the softmax activation function, is the merging weight of the target task t.
[0034] The application also provides a large language model fusion system based on gradient approximation, comprising:
[0035] An acquisition module is configured to acquire language text data samples.
[0036] A model construction module is configured to construct a fine-tuning model set and a pre-training model, wherein the fine-tuning model set comprises a plurality of fine-tuning models, and a difference value between the fine-tuning model and the pre-training model is defined as a task vector.
[0037] A simulation running module is configured to run the fine-tuning model set and the pre-training model in multiple ways respectively, and calculate a parameter importance evaluation result of the task vector based on the gradient under different running modes.
[0038] A determination module is configured to construct an activation function, and determine a pruning rate in combination with the activation function and the parameter importance evaluation result.
[0039] A fusion module is configured to complete pruning and fusion of the fine-tuning model set according to the pruning rate, and obtain a final fine-tuning model.
[0040] A processing module is configured to process the language text data samples by using the fine-tuning model, and obtain a corresponding processing result.
[0041] Through the above technical solutions, compared with the prior art, the application discloses a large language model fusion method and system based on gradient approximation, which uses few-shot data of each task to significantly improve the ability to locate the most relevant parameters of specific tasks, and performs two runs: one is clean run, that is, the fine-tuned model normally performs prediction; the other is corrupted run, that is, some fine-tuned parameters are replaced with corresponding parameters in the pre-training model, and then prediction output is performed. The prediction difference generated by the two runs reflects the difficulty of the damaged model in reproducing the clean model, thereby revealing the information specific to the task and the importance of the damaged parameters. By dividing the parameters of the model at different levels in different granularities, such as model level, layer level and hidden level, the importance of each parameter partition can be evaluated and the activated parameters can be located.
[0042] The parameter importance is helpful in the following two key aspects in model fusion: (1) dropout ratio calibration: the parameter partition importance score can guide pruning, and the parameter partitions of different levels or hidden levels are calibrated in the dropout ratio. Higher importance score indicates that the partition contains more task-specific information, so more neurons should be retained in the pruning process; on the contrary, for the partitions with lower importance, more aggressive pruning can be performed; (2) merged weight adjustment: the model-level importance reflects the total amount of task-specific information in the fine-tuned model and the generalization ability of the pre-trained model, therefore, the fine-tuned model with higher model-level importance should occupy a larger weight in the merging process.
[0043] Based on the above two adjustment methods, by calibrating the dropout ratio and fusion weight of the parameter partition of each level, aggressive pruning can be performed while the conflicts are minimized and the model performance is preserved. In addition, in actual scenarios, for models with a large number of parameter partitions, reducing the computational cost of causal tracing becomes particularly important. Therefore, on the basis of the previous APL framework, this paper further proposes an approximate method based on gradient to efficiently evaluate the importance of parameters.
[0044] Unlike the standard APL method that requires multiple runs, the gradient approximation method approximates the importance of a parameter partition by computing the inner product of the task vector (i.e., the parameters) and the gradient vector of the pre-trained model on the specific parameter partition. Intuitively, this method utilizes the information provided by each parameter partition in the gradient signal to quantify its contribution to the task-specific information through the inner product. This intuition can be verified by a theoretical derivation of the Taylor expansion of the pre-trained model parameters. In addition, experimental results also show that this method can maintain the reliability of importance estimation while significantly reducing the computational overhead. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only a part of the embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.
[0046] Figure 1 The overall flowchart of the gradient approximation-based large language model fusion method provided by the present application and the existing processing method schematic diagram;
[0047] Figure 2 The structural principle block diagram of the gradient approximation-based large language model fusion system provided by the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0049] In the prior art, it is crucial to reduce the redundancy and conflict in the delta parameters of multiple fine-tuned models for efficient model fusion. Existing methods usually remove redundant parameters through random or numerical amplitude-based methods. Although these methods have advantages in efficiency, they often ignore the task-specific information in each fine-tuned model.
[0050] In practical scenarios, it is often feasible to obtain running samples for the target task, and these samples can provide valuable task-specific information through inference or gradient signals. These information can significantly improve the ability to identify redundant parameters, thereby greatly reducing conflicts in the model fusion process. In addition, considering the potential benefits of reducing redundancy, the memory and time costs involved in using running samples are acceptable. In this case, the core challenge in the model fusion parameter pruning process is how to effectively identify redundant parameters through running samples, i.e., locating task-specific activation parameters in the fine-tuned model.
[0051] Inspired by the identification of neuron activations that are crucial to the causal tracing of model factual predictions, embodiments of the present application propose a large language model fusion method based on gradient approximation, which uses running samples of specific tasks to locate activation parameters. Specifically, given the running samples of the target task, the parameter importance of different parameter partitions is first estimated through causal tracing, then the dropout ratio of each parameter partition is calibrated and the merging weight is adjusted for different importance. By fusing task-specific information in the running samples, redundant parameters can be more accurately identified, thereby alleviating parameter conflicts and improving the effectiveness of model fusion, and improving the accuracy of subsequent data processing.
[0052] To solve the above technical problems, referring to Figure 1 The embodiments of the present application disclose a large language model fusion method based on gradient approximation, comprising the following steps:
[0053] Obtain language text data samples;
[0054] Construct a set of fine-tuned models and a pre-trained model Θ b , wherein the set of fine-tuned models includes T fine-tuned models Define the difference value between the fine-tuned model and the pre-trained model as the task vector δ, and specifically the δ parameter applied to task t is defined as:
[0055] Δ t =Θ t -Θ b (1)
[0056] In the formula, Θ t represents the model parameters of task t obtained by fine-tuning Θ b , and Δt represents the δ parameter of task t;
[0057] Run the set of fine-tuned models and the pre-trained model in multiple ways respectively, and calculate the parameter importance evaluation results of the task vector based on the gradient under different running modes;
[0058] Construct an activation function, and determine the pruning rate based on the activation function and the parameter importance evaluation results;
[0059] According to the pruning rate, pruning and fusion of the fine-tuning model set are completed to obtain a final fine-tuning model.
[0060] The fine-tuning model is used to process the language text data sample to obtain a corresponding processing result.
[0061] In a specific embodiment, the specific processing process of calculating the parameter importance evaluation result of the task vector based on the gradient in different running modes includes:
[0062] Obtain a running sample D t , wherein the running sample D t is a running sample in the total sample;
[0063] Construct multiple running modes, and input the running sample D t into the fine-tuning model for prediction to obtain a corresponding plurality of prediction output results;
[0064] Calculate the difference corresponding to the plurality of prediction output results, and the gradient corresponding to the task vector;
[0065] Determine the corresponding parameter importance evaluation result according to the running sample, the difference, and the gradient.
[0066] In a specific embodiment, the specific processing process of obtaining the corresponding plurality of prediction output results includes:
[0067] The running mode includes a clean running mode and a damaged running mode;
[0068] In the clean running mode, the running sample D t is input into the fine-tuning model for prediction to obtain a corresponding true probability;
[0069] In the damaged running mode, a certain partition of the fine-tuning model is replaced with a noise part in the pre-training model, and the running sample is input into the modified fine-tuning model for prediction to obtain a corresponding prediction output.
[0070] In a specific embodiment, the specific processing process of determining the corresponding parameter importance evaluation result according to the running sample, the difference, and the gradient includes:
[0071] Calculate the inner product of the running sample and the gradient;
[0072] Determine the corresponding parameter importance evaluation result according to the inner product, the difference, and the gradient.
[0073] Specifically, the specific implementation process of the above process mainly includes the following steps:
[0074] Analyzing the information flow in the causal trace can help determine the relative importance of certain parameters. To quantify the contribution of parameters to the running samples in a specific task, inspired by the parameter causal localization method in model editing, two activation analysis runs are mainly performed, the specific process is as follows:
[0075] 1. In the clean run, the running sample D t is input into the corresponding fine-tuned model Θ t , and the true label probability P t is obtained, and the information flow is represented by→, and the process is represented as follows:
[0076] D t →Θ t →P t (2) 2. In the damaged run, by replacing a certain parameter partition (denoted as ) of the fine-tuned model Θ t with the corresponding "noise" parameter (denoted as ) in the pre-trained model, the fine-tuned model Θ t is modified, and then the running sample D t is input into the modified model to obtain the predicted output, denoted as , and the information flow of the damaged run is represented as:
[0077]
[0078] where represents the splicing operation, and represents the part of the fine-tuned model Θ t except .
[0079] The clean run evaluates the impact of all δ parameters on a specific task, while the damaged run only evaluates a subset of these parameters, and the difference between P t and quantifies the difficulty of recovering from the damaged run to the clean run, and by evaluating this difficulty, the importance of the damaged δ parameter partition for a specific task can be estimated.
[0080] The parameter importance (ParameterImportance), denoted as , is defined as the difference between the predicted output in the damaged run and the true label probability P t , and the specific expression is:
[0081]
[0082] It should be noted that is usually smaller than P t , because the fine-tuned model Θt Replacing part of the parameters usually does not improve the prediction performance. In this case, a smaller value (i.e. a larger gap between P t and P t ) indicates that it is more difficult to accurately reconstruct the prediction, which means that the damaged parameter partition has a stronger causal effect on the prediction, i.e., the damaged parameter partition is crucial to the task t and therefore should not be removed in the pruning process. By tracking the causal effect, the importance of the parameters can be determined, and the active parameters for a specific task can be identified.
[0083] Ideally, causal intervention can obtain the importance of each neural activation unit or parameter, however, in practical applications, evaluating the importance at the parameter level is usually computationally expensive. To solve this problem, the parameters are divided into model level, layer level and hidden level according to the network structure, where the hidden level refers to the smallest parameter vector that plays a role in the network structure, such as the attention module.
[0084] For the model level parameter importance, the importance of the corresponding model is denoted as which represents the complete task-specific information contained in the fine-tuned model, which is very effective in determining the relevant weight for each model when fusing multiple models, the specific process is as follows:
[0085] For the layer level parameter importance and the hidden state level parameter importance, they are denoted as and respectively, and the corresponding damaged run is defined as:
[0086] D t →Θ b →P b (5)
[0087]
[0088] where l represents the l-th layer of the total L layers in the fine-tuned model Θ t , and h represents the h-th hidden level parameter vector of the total H layers.
[0089] After the above analysis, l∈[1,L] (or h∈[1,H]) can be traversed to calculate (or the corresponding ), and the parameters are sorted according to the importance of the parameter partition. However, when L or H is large, even if the step-by-step replacement is only performed at the layer level or the hidden level, the computational complexity will be significantly increased.
[0090] To solve this problem, the embodiments of the present application introduce a gradient approximation method for causal tracking, which can effectively reduce the computational cost when L or H is large, the specific process is as follows:
[0091] a、Gradient approximation of parameter importance
[0092] Let L(Θ, D) denote the loss of the pre-trained model Θ on the run sample D t The fine-tuned model Θ t The validation loss of the fine-tuned model Θ t on the run sample D b The gradient of the fine-tuned model Θ indicates the direction in which the loss L(Θ b ,D t ) decreases.
[0093] If the parameter partition in which the δ parameter in the fine-tuned model Θ t is divided is orthogonal to the corresponding partition in the pre-trained model Θ , then this partition does not capture task-specific information and can be deleted most of the time.
[0094] Based on this, the embodiments of the present application propose to use the size of the inner product of D t and to represent the importance of each parameter partition. For example, for the hierarchical partition l, let and be the δ parameter and its corresponding gradient of the l-th layer, respectively, and the importance of the partition l is represented as:
[0095]
[0096] The loss L(Θ t ,D b ) of the fine-tuned model Θ t at the pre-trained model Θ t is first-order Taylor expanded, and its pruned version Θ' t =(1-M t )⊙Δ t +Θ b is also Taylor expanded at the pre-trained model Θ b , and the specific expression is:
[0097]
[0098] The goal is to find a mask matrix M t such that the performance of the model Θ' t is similar to that of the original fine-tuned model Θ t , that is, we hope L(Θ' t ,D t ) and L(Θ t ,D t ) are as close as possible, and according to the above formula, the loss difference can be derived as follows:
[0099]
[0100] When M t is an all-zero vector, it can be deduced that |L(Θ t ,D t )-L(Θ' t ,D t )|=0. However, M t =0 indicates that no parameters are pruned, which contradicts the goal of removing as many parameters as possible to reduce redundancy.
[0101] In fact, the right side of the equality of formula (11) can be restructured as:
[0102]
[0103] In the formula, d represents the total number of parameters. This indicates that if for a certain dimension i, is small (i.e. close to zero), setting M t [i]=1 does not significantly increase |L(Θ t ,D t )-L(Θ' t ,D t )|, so the Δ t parameters of this dimension can be removed. Conversely, if is large, M t [i] should be set to 0, indicating that the parameters of this dimension should be retained.
[0104] Extending the above steps to hierarchical or hidden level parameter partitioning, the inner product between the gradient vector of the fine-tuned parameters and the pre-trained model is used to quantify the importance of the parameters in each partition in the model, i.e. It should be noted that introducing higher order terms (such as the Hessian matrix) in the Taylor expansion can obtain a more accurate approximation, however, the complexity involved in calculating these high order terms is very high. Therefore, the embodiments of the present application only consider the first order approximation, which only requires one backpropagation operation on the pre-trained model, and this way is more efficient when dealing with a large number of partitions compared to the method of replacing parameter partitions step by step.
[0105] In a specific embodiment, the specific processing process of determining the pruning rate in combination with the activation function and the parameter importance evaluation result includes:
[0106] The parameter partition with a higher contains more task-related information, which indicates that more δ parameters need to be retained, i.e. the corresponding pruning rate is relatively low. In order to maintain the relative importance relationship and further utilize the absolute importance value of each partition, the present application introduces a tanh activation function to determine the pruning ratio of each parameter based on calibrating the initial pruning ratio λ;
[0107] The specific expression of the activation function of the pruning ratio is:
[0108]
[0109] wherein, represents the pruning ratio, λ represents the initial pruning ratio, tanh represents the activation function, τ1 represents the temperature coefficient of tanh, and ∈ represents the threshold value.
[0110] In a specific embodiment, the specific processing procedure for obtaining the final fine-tuned model includes:
[0111] By using the pruning ratio guided by the parameter importance, the redundant parameters are more accurately reduced, so that the parameter conflicts are more effectively alleviated in the fusion process. In addition, the model-level importance measures all the task-specific information learned by the fine-tuned model compared with the pre-trained model. Generally speaking, a smaller model importance indicates that the pre-trained model has been well generalized to the specific task, while a larger value indicates that the task contains more specific knowledge that needs to be learned by the pre-trained model. Therefore, when fusing multiple models, the fine-tuned model with higher importance should occupy a higher proportion to retain more task-specific information.
[0112] Accordingly, first, the parameters of the fine-tuned model are pruned, and then the multiple models are fused using the model importance guided weight, the task vector of the fine-tuned model is pruned using the pruning ratio, and the model fusion is performed using the parameter importance evaluation result. The specific expression is:
[0113]
[0114] wherein τ2 is the temperature parameter of the softmax activation function, is the merged weight of the target task t.
[0115] Referring to Figure 2 The embodiment of the present application also provides a system using the gradient approximation-based large language model fusion method, which comprises:
[0116] The acquisition module is configured to acquire the language text data sample.
[0117] The model construction module is configured to construct a fine-tuned model set and a pre-trained model, wherein the fine-tuned model set comprises a plurality of fine-tuned models, and the difference value between the fine-tuned model and the pre-trained model is defined as a task vector.
[0118] The simulation running module is configured to run the fine-tuning model set and the pre-training model in multiple ways respectively, and to calculate the parameter importance evaluation results of the task vectors based on gradients in different running modes;
[0119] The determination module is configured to construct an activation function, and to determine a pruning rate in combination with the activation function and the parameter importance evaluation results;
[0120] The fusion module is configured to complete pruning and fusion of the fine-tuning model set according to the pruning rate, and to obtain a final fine-tuning model.
[0121] The processing module is configured to process language text data samples by using the fine-tuning model, and to obtain corresponding processing results.
[0122] The method provided by the embodiment of the application is verified, LLaMA-2-chat-7B is used as a pre-training model, and 6 representative tasks are selected to fine-tune the pre-training model, including AG News, Hellaswag, MNLI, MRPC, SST2 and Winogrande, the tasks are selected from FLAN, and Accuracy is used as an evaluation index to evaluate, the data set statistics are shown in Table 1, and the specific experimental results are shown in Table 2.
[0123] Table 1 data set statistics
[0124]
[0125] Table 2 ID model fusion results
[0126]
[0127] From the results, it can be observed that MI-TA is obviously better than TA, which shows that the fusion weight adjustment based on model-level importance guidance is effective. In addition, the performance of APL is better than that of Dare, which proves that, compared with Dare, the method provided by the application can more effectively alleviate parameter conflicts.
[0128] In order to evaluate the generalization ability of APL in OOD tasks, in addition to the three basic models of pre-training LLaMA, context-based learning (ICL) LLaMA and multi-task learning (Multi-task-learning), this section also compares with the recent SOTA method LM-Cocktail. Since LM-Cocktail has calibrated the weights by few-shot, this section only applies Dare and APL on LM-Cocktail for comparison. The results are shown in Table 3.
[0129] Table 3 model fusion results
[0130]
[0131] As shown in Table 3, parameter pruning can reduce conflicts and improve the effect of model fusion. The effect of APL is greater than that of Dare, which proves that APL can better locate the active parameters, thereby achieving a better pruning effect.
[0132] The various embodiments described in the specification are progressive in nature, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be mutually referred to. For the apparatus disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0133] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for fusing large language models based on gradient approximation, characterized in that, Includes the following steps: Obtain language text data samples; Construct a fine-tuned model set and a pre-trained model, wherein the fine-tuned model set includes multiple fine-tuned models, and the difference value between the fine-tuned model and the pre-trained model is defined as the task vector; The fine-tuned model set and the pre-trained model are run in various ways, and the parameter importance evaluation results of the task vector are calculated based on gradients under different running modes. Construct an activation function, and determine the pruning rate by combining the activation function and the parameter importance evaluation results; The fine-tuning model set is pruned and fused according to the pruning rate to obtain the final fine-tuning model; The language text data samples are processed using a fine-tuning model to obtain the corresponding processing results.
2. The method for fusing large language models based on gradient approximation according to claim 1, characterized in that, The specific processing steps for calculating the parameter importance evaluation results of the task vector based on gradients under different operating modes include: Obtain the running sample; Multiple operating modes are constructed, and the operating samples are input into the fine-tuning model for prediction under multiple operating modes to obtain multiple corresponding prediction output results; Calculate the differences between multiple prediction outputs and the gradient corresponding to the task vector; The corresponding parameter importance assessment results are determined based on the running samples, the differences, and the gradient.
3. The method for fusion of large language models based on gradient approximation according to claim 2, characterized in that, The specific processing steps to obtain multiple corresponding prediction outputs include: The operating modes include: clean operating mode and damaged operating mode; In the clean running mode, the running samples are input into the fine-tuning model for prediction to obtain the corresponding true probabilities; In the damaged operation mode, a certain partition of the fine-tuning model is replaced with the noise part in the pre-trained model, and the running sample is input into the modified fine-tuning model for prediction to obtain the corresponding prediction output.
4. The method for fusing large language models based on gradient approximation according to claim 3, characterized in that, The specific processing steps for determining the corresponding parameter importance evaluation results based on the running samples, the differences, and the gradient include: Calculate the inner product of the running sample and the gradient; The corresponding parameter importance evaluation result is determined based on the inner product, the difference, and the gradient.
5. The method for fusing large language models based on gradient approximation according to claim 1, characterized in that, The specific process for determining the pruning rate by combining the activation function and parameter importance evaluation results includes: The specific expression for the pruning rate in constructing the activation function is as follows: In the formula, Let λ represent the pruning rate, λ represent the initial pruning rate, tanh represent the activation function, τ1 represent the temperature coefficient of tanh, and ∈ represent the threshold.
6. The method for fusing large language models based on gradient approximation according to claim 1, characterized in that, The specific processing steps to obtain the final fine-tuned model include: The task vector of the fine-tuned model is pruned using a pruning rate, and the model is fused using the parameter importance evaluation results. The specific expression is as follows: In the formula, τ2 is the temperature parameter of the softmax activation function. It is the combined weight of the target task t.
7. A system utilizing the large language model fusion method based on gradient approximation as described in any one of claims 1-6, characterized in that, include: The acquisition module is used to acquire language text data samples; The model building module is used to build a fine-tuned model set and a pre-trained model, wherein the fine-tuned model set includes multiple fine-tuned models, and the difference value between the fine-tuned model and the pre-trained model is defined as the task vector. The simulation execution module is used to run the fine-tuned model set and the pre-trained model in various ways, and calculate the parameter importance evaluation results of the task vector based on gradients under different execution modes; The module is defined to construct the activation function, and the pruning rate is determined by combining the activation function with the parameter importance evaluation results. The fusion module is used to prune and fuse the fine-tuned model set according to the pruning rate to obtain the final fine-tuned model; The processing module is used to process the language text data samples using a fine-tuning model to obtain the corresponding processing results.
Citation Information
Cited By
Explanatability fusion and recovery method after large language model training based on interpretability
CN121936572A