Incremental model merging method and system for large language models
By using gradients and the direction of task vectors in large language model merging to calculate confusing parameters, and perform parameter importance sampling and scaling, the task conflict problem is solved, the model merge performance is improved, and the model is closer to the expert model.
Patent Information
- Application Number
- CN202411803376.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The existing large language model merging technology is difficult to effectively solve task conflict problems, resulting in large performance losses of the merged model on specific tasks.
By introducing the direction of gradients and task vectors, the confusing parameters and their parameter mask matrix are calculated, and the parameter importance is further calculated and probability sampling and scaling is performed for model merging to avoid parameter conflicts and task conflicts.
The model merging performance is improved, making the merged model closer to the expert model before the merger, effectively solving the problem of parameter conflicts and task conflicts, and improving the universality and flexibility of the model.
Smart Images

Figure CN119294465B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and in particular relates to an incremental model merging method and system for a large language model. Background Art
[0002] With the development of modern machine learning and artificial intelligence, large language model (LLM) technology has received more and more attention, and many companies have also joined the development of large language model technology. At present, the mainstream large language model mainly achieves its powerful performance based on model pre-training and downstream task fine-tuning. Institutions with limited computing resources, such as micro-enterprises, universities, and hospitals, use the open source pre-trained model base of large companies, combined with domain-specific data sets, to fine-tune large models suitable for specific scenarios for specific tasks. In order to meet the multi-faceted needs of applications, institutions require the model to have the ability to solve different tasks when fine-tuning the pre-trained model, so it is necessary to fine-tune multiple dedicated large models for solving different tasks. However, training these models requires sufficient computing resources, and maintaining these models will also incur huge storage costs, and independent tasks are isolated from each other, which greatly limits the versatility of the model. In addition, some tasks may involve privacy data sets that are difficult to collect, which greatly limits the innovative application of large model technology.
[0003] The large language model merging (Model Merge, MM) technology provides a simple and efficient solution, allowing users to use multiple domain-specific expert models to fuse and generate complex models that have multi-task solving capabilities. This method not only inherits the high performance of each expert model in its own exclusive field, but also integrates their knowledge so that the merged model can handle cross-domain or multi-task problems, thereby improving the versatility and flexibility of the model. The biggest advantage of model merging is that it is simple and efficient, and there is no need to train large models from scratch, which greatly reduces computing resources and time. In the scenario of adding new task capabilities to large language models, model merging technology does not require the use of additional task data sets for fine-tuning, so it has obvious advantages when dealing with new tasks.
[0004] At present, large language model merging technology can be divided into weight-based model merging method, subspace-based model merging method, routing-based model merging method and post-processing-based model merging method. The simplest and most common method is the weight-based model merging method, which calculates the weighted average of the expert model parameters as the parameters of the merged model. The subspace-based model merging method selects some model parameters for model merging through the parameter importance algorithm to minimize the parameter conflict problem in the model merging process. The routing-based model merging method abstracts the expert model into an expert module in the model, constructs a mixture of experts (MOE), and trains the parameters of the expert model routing selector through a calibration data set. The post-processing-based model merging method adds an additional calibration module after the model is merged, and uses the calibration data set to edit the model output features, so that the features of the merged model are as similar as possible to the output features of the expert model before the merge, so that the merged model is close to the distribution of the expert model.
[0005] However, the current mainstream model merging methods cannot effectively solve the problem of task conflict, that is, the ability of the merged model to solve specific tasks still suffers from a significant performance loss compared to the expert model before merging. Therefore, it is of great practical significance to design a model merging technology that can alleviate task conflict and parameter conflict by using the available information such as model parameters, gradients, and update directions. Summary of the invention
[0006] In view of the above, the purpose of the present invention is to provide an incremental model merging method and system for large language models. A model merging method that alleviates task conflicts is designed by using information such as the update direction and gradient direction of model parameters. By introducing the direction of gradient and task vectors, a new parameter importance calculation standard is constructed. After probability sampling of parameter importance, the task vector is scaled for model merging. By merging expert models of different specific tasks in a sequential and incremental manner, parameter conflicts and task conflicts are avoided, thereby further improving the model merging performance.
[0007] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention is as follows:
[0008] In a first aspect, an embodiment of the present invention provides an incremental model merging method for a large language model, comprising the following steps:
[0009] During the first incremental model merging, the perplexity parameters of the expert model are calculated using the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration dataset, and the perplexity parameters are converted into a parameter mask matrix.
[0010] Calculate the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, sample the parameter importance to generate a sampling mask matrix, and use the sampling mask matrix to scale the task vector to generate an incremental task vector;
[0011] Add the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtain the merged model, and complete the first incremental model merge;
[0012] In subsequent incremental model merging, the merged model and the new expert model obtained from the previous merge will replace the above-mentioned pre-trained model and expert model respectively each time, and the new expert model will be merged into the merged model obtained from the previous merge using the same method as the first incremental model merging, until the incremental model merging of all expert models is completed in sequence.
[0013] Specifically, the task vector updated by the expert model in the incremental model merging process and the gradient of the pre-trained model on the calibration data set are used to calculate the perplexity parameters of the expert model, and the perplexity parameters are converted into a parameter mask matrix, including:
[0014] Calculate the The task vector of the expert model , expressed as:
[0015] ,
[0016] in, represents the parameters of the pre-trained model, Indicates Parameters of the expert model;
[0017] Compute the pre-trained model in Calibration dataset used for fine-tuning expert models The gradient on , expressed as:
[0018] ,
[0019] in, represents the gradient calculation, Representation is based on loss function;
[0020] When the gradient The direction and mission vector When the directions are different, the result is the same as the task vector The relevant parameters are used as perplexity parameters in the The parameter mask matrix of the expert model Different element values are used in to indicate that the gradient and the task vector have the same or different directions.
[0021] Specifically, the parameter importance calculation formula of the expert model is:
[0022] ,
[0023] in, Indicates The importance of parameters of the expert model, Indicates The task vector of the expert model, Indicates that the pre-trained model is The gradient on the calibration dataset used for fine-tuning the expert model, Indicates The parameter mask matrix of the expert model, represents dot product calculation, and They represent scaling function and normalization function respectively.
[0024] Specifically, the calculation formula of the sampling mask matrix is:
[0025] ,
[0026] in, Indicates The sampling mask matrix of the expert model, Indicates The importance of parameters of the expert model, represents the binomial distribution.
[0027] Specifically, the calculation formula of the incremental task vector is:
[0028] ,
[0029] in, Indicates The incremental task vector of the expert model, Indicates The sampling mask matrix of the expert model, Indicates The task vector of the expert model, Represents dot product calculation.
[0030] Specifically, adding the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model includes:
[0031] ,
[0032] in, represents the parameters of the merged model obtained by the first incremental model merge, represents the parameters of the pre-trained model, Indicates The task-specific merging coefficient of the expert models, Indicates The incremental task vector of the expert model.
[0033] In a second aspect, to achieve the above-mentioned purpose of the invention, an embodiment of the present invention further provides an incremental model merging system for a large language model, which is implemented by the above-mentioned incremental model merging method for a large language model, and includes: a confusion parameter positioning module, a parameter importance calculation and sampling scaling module, an incremental model merging module, and a hierarchical merging module;
[0034] The perplexity parameter positioning module is used to calculate the perplexity parameters of the expert model during the first incremental model merging by using the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration data set, and convert the perplexity parameters into a parameter mask matrix;
[0035] The parameter importance calculation and sampling scaling module is used to calculate the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, sample the parameter importance to generate a sampling mask matrix, and use the sampling mask matrix to scale the task vector to generate an incremental task vector;
[0036] The incremental model merging module is used to add the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtain the merged model, and complete the first incremental model merging;
[0037] The hierarchical merging module is used to replace the pre-trained model and the expert model mentioned above with the merged model and the new expert model obtained from the previous merge each time when merging subsequent incremental models, and merge the new expert model into the merged model obtained from the previous merge using the same method as the first incremental model merging, until the incremental model merging of all expert models is completed in sequence.
[0038] In the third aspect, in order to achieve the above-mentioned purpose of the invention, an embodiment of the present invention also provides an electronic device, including a memory and one or more processors, wherein the memory is used to store a computer program, and the processor is used to implement the above-mentioned incremental model merging method for large language models when executing the computer program.
[0039] In a fourth aspect, in order to achieve the above-mentioned purpose of the invention, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the above-mentioned incremental model merging method for large language models is implemented.
[0040] In a fifth aspect, in order to achieve the above-mentioned purpose of the invention, an embodiment of the present invention further provides a computer product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned incremental model merging method for a large language model is implemented.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention proposes a method for calculating the model confusion parameters based on task vectors and gradients, which can accurately locate the confused parameters in the model parameters according to the calibration data set, and further improve the merging performance of the model by analyzing the confusion parameters, so that the merged model after merging is closer to the expert model before merging.
[0043] (2) The parameter importance calculation and sampling scaling calculation proposed in the invention can effectively calculate the importance of parameters in the task vector, and generate a sampling mask matrix based on the parameter importance to scale the task vector, thereby retaining important parameter features. The support of scaling technology can further amplify the features of important parts and improve the model merging performance.
[0044] (3) The novel and practical incremental model merging technology proposed in this invention uses one expert to merge each time, and merges other expert models in turn based on the merged model, which effectively solves parameter conflicts and task conflicts. It is suitable for model merging in different scenarios and is compatible with the basic large language models commonly used on the market.
[0045] (4) The present invention can effectively reduce the usage of video memory by using a hierarchical merging strategy of sequentially merging other expert models based on the merged model, and supports merging models with a large number of parameters on a single graphics card. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 It is a flowchart of an incremental model merging method for a large language model provided by an embodiment of the present invention;
[0048] Figure 2 It is a schematic diagram of the overall framework of incremental model merging provided by an embodiment of the present invention;
[0049] Figure 3 It is a structural diagram of an incremental model merging system for a large language model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0051] The inventive concept of the present invention is: in view of the fact that the model merging method in the prior art cannot effectively solve the problem of task conflict, the embodiment of the present invention provides an incremental model merging method and system for a large language model, which calculates the perplexity parameter and its parameter mask matrix by introducing the direction of the gradient and the task vector, further calculates the parameter importance and performs probability sampling on it, and then scales the task vector for model merging. In the incremental model merging process, expert models of different specific tasks are merged in sequence and incrementally and superimposedly, thereby avoiding parameter conflicts and task conflicts, and further improving the model merging performance.
[0052] Based on the above-mentioned incremental model merging method and system for large language models, in the embodiment, in specific application, it is first necessary to obtain a pre-trained model pre-trained on a large-scale general corpus, and based on the pre-trained model, different expert models (including expert models for text generation tasks, expert models for question and answer consultation tasks, expert models for text classification tasks, expert models for sentiment analysis tasks, etc.) are fine-tuned on different calibration data sets for different specific tasks (including text generation tasks, question and answer consultation tasks, text classification tasks, and sentiment analysis tasks, etc.). The ultimate goal of the incremental model merging is to merge these expert models into the pre-trained model in turn to obtain a merged model with multi-task processing capabilities.
[0053] Figure 1 FIG. 1 is a flow chart of an incremental model merging method for a large language model provided by an embodiment of the present invention. Figure 1 As shown, the embodiment provides an incremental model merging method for a large language model, comprising the following steps:
[0054] S1, when the incremental model is merged for the first time, the perplexity parameters of the expert model are calculated through the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration dataset, and the perplexity parameters are converted into a parameter mask matrix.
[0055] In the embodiment, the model merging hyperparameter configuration needs to be performed first, including but not limited to:
[0056] (1) Model base (pre-trained model) used by the expert model;
[0057] (2) the number of expert models used in the merging process;
[0058] (3) Specific expert model parameters used in the model merging process;
[0059] (4) The task-specific merging coefficient used in the model merging process;
[0060] (5) Standardized methods used to calculate sampling probabilities;
[0061] (6) The probability distribution function used to calculate the sampling probability;
[0062] (7) Parameter scaling function used when calculating scaling;
[0063] (8) The number of times the expert models are merged during incremental merging;
[0064] (9) Calibration data set used in the gradient calculation process;
[0065] (10) The proportion of the calibration dataset used in the gradient calculation process.
[0066] In the embodiment, the model base and expert model used can be flexibly customized according to the user's needs for the merged model. Usually, the existing open source model base and the model fine-tuned on the base can be downloaded from the open source community hugging face. In addition, the model of the model service provider can also be used. The present invention supports the current mainstream large language model structure, such as the Llama series, Gemma series, Qwen series and GLM series.
[0067] In the embodiment, when merging models incrementally, the number of expert models for incremental merging is set to 3-10, and a total of 30-100 incremental merging is performed, and the specific task merging coefficient used for each merging is between 0.1 and 1. For different expert models, the user can set the corresponding specific task merging coefficient to 0.1-1 according to the weight of the task in the final merged model, set the normalization function to the Norm function, the scaling function to the Softmax function, and the probability distribution function to the binomial distribution.
[0068] When merging incremental models for the first time, the expert model needs to be merged into the pre-trained model, and the perplexity of the expert model parameters is calculated through the task vector updated by the expert model and the gradient of the pre-trained model on the calibration data set to obtain the perplexity parameters, and the parameters that cause confusion are deleted to ensure the performance of the merged model. The task vector of the expert model The calculation formula is as follows:
[0069] ,
[0070] in, represents the parameters of the pre-trained model, Indicates The parameters of the expert model.
[0071] The pre-trained model is Calibration dataset used for fine-tuning expert models The gradient on The calculation formula is as follows:
[0072] ,
[0073] in, represents the gradient calculation, Representation is based on the loss function.
[0074] When the gradient The direction and mission vector If the directions are different, it means that this part of the parameters does not know in which direction the parameters should be updated, then this part of the parameters is confused parameters and will be deleted. Here we use a parameter mask matrix Indicated in The confusion parameter of the task vector of the expert model, where the element value is 1 when the gradient is in the same direction as the task vector, and 0 otherwise.
[0075] S2, calculates the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, samples the parameter importance to generate a sampling mask matrix, and uses the sampling mask matrix to scale the task vector to generate an incremental task vector.
[0076] In an embodiment, the importance of each parameter is calculated based on the dot product between the task vector, the gradient, and the parameter mask matrix, and the task vector is sampled and scaled based on the importance of the parameter.
[0077] First, The parameter importance calculation formula of an expert model is as follows:
[0078] ,
[0079] in, Indicates The importance of parameters of the expert model, Indicates The parameter mask matrix of the expert model, represents dot product calculation, and They represent scaling function and normalization function respectively.
[0080] Then, the importance of parameters Perform sampling and generate a sampling mask matrix , the calculation formula is as follows:
[0081] ,
[0082] in, Indicates The sampling mask matrix of the expert model, Indicates The importance of parameters of an expert model. Represents a binomial distribution, where the element value is either 1 or 0, The larger the value, the greater the probability of being 1.
[0083] Afterwards, the task vector is scaled to generate an incremental task vector for model merging , the calculation formula is as follows:
[0084] ,
[0085] in, Indicates The incremental task vector of the expert model.
[0086] S3, adds the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtains the merged model, and completes the first incremental model merge.
[0087] In the embodiment, by incrementally generating a merge model, based on the calculated incremental task vector , added to the parameters of the pre-trained model, generate the parameters of the merged model, and then obtain the merged model. The merging formula for the first incremental model merging is as follows:
[0088] ,
[0089] in, represents the parameters of the merged model obtained by the first incremental model merge, represents the parameters of the pre-trained model, Indicates The specific task merging coefficient of the expert model is used to adjust the weight of the specific tasks possessed by the expert model.
[0090] S4, when the subsequent incremental models are merged, the merged model and the new expert model obtained from the previous merge will replace the above-mentioned pre-trained model and expert model respectively each time, and the new expert model will be merged into the merged model obtained from the previous merge using the same method as the first incremental model merge, until the incremental model merge of all expert models is completed in sequence.
[0091] Repeat steps S1 to S3 for multiple rounds until the configuration stop condition is reached. NThe merging of expert models eventually results in a merged model that inherits multiple specific task capabilities. N .
[0092] The formula for incremental merging of multiple expert models is as follows:
[0093] ,
[0094] in, and Respectively represent Second and The parameters of the merged model are obtained by merging the incremental models. When is 1, These are the parameters of the pre-trained model in the first incremental model merge .
[0095] Based on the incremental model merging method for a large language model provided above, in an embodiment, Figure 2 The figure shows the overall framework of incremental model merging. Users first collect expert models (expert model-1) and model bases (pre-trained models) for specific tasks that need to be merged from open source communities or model providers, and collect calibration datasets (calibration datasets) according to the specific tasks that need to be merged. ). Afterwards, the task vector is calculated based on the pre-trained model and the expert model-1 , and the specific gradient corresponding to the pre-trained model is calculated based on the calibration dataset-1 Next, the calculated task vector and gradient Perform perplexity parameter calculations, parameter importance calculations, sampling and scaling calculations to obtain the incremental task vector Finally, the incremental task vector Merge with the parameters of the current pre-trained model to generate the merged model -1 for the next incremental calculation. Perform the above process multiple times, traversing all expert models that need to be merged, until the set number of incremental updates is reached, and finally the merged model - N .
[0096] In summary, the incremental model merging method for a large language model provided by the embodiment of the present invention can effectively avoid parameter conflicts and task conflicts during the model merging process, effectively improve the performance of the model, and is suitable for model merging in different scenarios.
[0097] Based on the same inventive concept, Figure 3As shown, an embodiment of the present invention further provides an incremental model merging system 300 for a large language model, including: a perplexity parameter positioning module 310, a parameter importance calculation and sampling scaling module 320, an incremental model merging module 330, and a hierarchical merging module 340.
[0098] The perplexity parameter localization module 310 is used to calculate the perplexity parameters of the expert model during the first incremental model merging by using the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration data set, and convert the perplexity parameters into a parameter mask matrix.
[0099] The parameter importance calculation and sampling scaling module 320 is used to calculate the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, sample the parameter importance to generate a sampling mask matrix, and use the sampling mask matrix to scale the task vector to generate an incremental task vector.
[0100] The incremental model merging module 330 is used to add the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtain the merged model, and complete the first incremental model merging.
[0101] The hierarchical merging module 340 is used to replace the above-mentioned pre-trained model and expert model respectively with the merged model and the new expert model obtained from the previous merge during each subsequent incremental model merging, and to merge the new expert model into the merged model obtained from the previous merge using the same method as the first incremental model merging, until the incremental model merging of all expert models is completed in sequence.
[0102] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device, including a memory and one or more processors, the memory is used to store a computer program, and the processor is used to implement the above-mentioned incremental model merging method for a large language model when executing the computer program.
[0103] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the above-mentioned incremental model merging method for a large language model is implemented.
[0104] Based on the same inventive concept, an embodiment of the present invention further provides a computer product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned incremental model merging method for a large language model is implemented.
[0105] It should be noted that the incremental model merging system for a large language model, electronic device, computer-readable storage medium and computer product provided in the embodiments of the present invention all belong to the same inventive concept as the incremental model merging method for a large language model. The specific implementation process is detailed in the embodiment of the incremental model merging method for a large language model, which will not be repeated here.
[0106] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An incremental model merging method for a large language model, characterized in that: The following steps are involved: First, a pre-trained model pre-trained on a large-scale general corpus is obtained. Based on the pre-trained model, different expert models are fine-tuned on different calibration datasets for different specific tasks, including expert models for text generation tasks, expert models for question-answering consultation tasks, expert models for text classification tasks, and expert models for sentiment analysis tasks. When the incremental model is merged for the first time, the perplexity parameters of the expert model are calculated through the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration dataset, and the perplexity parameters are converted into a parameter mask matrix, including: Calculate the The task vector of the expert model , expressed as: , in, represents the parameters of the pre-trained model, Indicates Parameters of the expert model; Compute the pre-trained model in Calibration dataset used for fine-tuning expert models The gradient on , expressed as: , in, represents the gradient calculation, Representation is based on loss function; When the gradient The direction and mission vector When the directions are different, the result is the same as the task vector The relevant parameters are used as perplexity parameters in the The parameter mask matrix of the expert model Different element values are used in to indicate that the gradient and the task vector have the same or different directions; Calculate the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, sample the parameter importance to generate a sampling mask matrix, and use the sampling mask matrix to scale the task vector to generate an incremental task vector; Add the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtain the merged model, and complete the first incremental model merge; In subsequent incremental model merging, the merged model and the new expert model obtained from the previous merge will replace the above-mentioned pre-trained model and expert model respectively each time, and the new expert model will be merged into the merged model obtained from the previous merge using the same method as the first incremental model merging, until the incremental model merging of all expert models is completed in sequence, thereby avoiding parameter conflicts and task conflicts, further improving the model merging performance, and effectively reducing the occupancy of video memory.
2. The incremental model merging method for a large language model according to claim 1, characterized in that: The formula for calculating the importance of parameters of the expert model is: , in, Indicates The importance of parameters of the expert model, Indicates The task vector of the expert model, Indicates that the pre-trained model is The gradient on the calibration dataset used for fine-tuning the expert model, Indicates The parameter mask matrix of the expert model, Represents dot product calculation, and They represent scaling function and normalization function respectively.
3. The incremental model merging method for a large language model according to claim 1, characterized in that: The calculation formula of the sampling mask matrix is: , in, Indicates The sampling mask matrix of the expert model, Indicates The importance of parameters of the expert model, represents the binomial distribution.
4. The incremental model merging method for a large language model according to claim 1, characterized in that: The calculation formula of the incremental task vector is: , in, Indicates The incremental task vector of the expert model, Indicates The sampling mask matrix of the expert model, Indicates The task vector of the expert model, Represents dot product calculation.
5. The incremental model merging method for a large language model according to claim 1, characterized in that: The step of adding the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model includes: , in, represents the parameters of the merged model obtained by the first incremental model merge, represents the parameters of the pre-trained model, Indicates The task-specific merging coefficient of the expert models, Indicates The incremental task vector of the expert model.
6. An incremental model merging system for a large language model, implemented by the incremental model merging method for a large language model according to any one of claims 1 to 5, characterized in that: include: Confusion parameter location module, parameter importance calculation and sampling scaling module, incremental model merging module, and hierarchical merging module; The perplexity parameter positioning module is used to calculate the perplexity parameters of the expert model during the first incremental model merging by using the task vector updated by the expert model during the incremental model merging process and the gradient of the pre-trained model on the calibration data set, and convert the perplexity parameters into a parameter mask matrix; The parameter importance calculation and sampling scaling module is used to calculate the parameter importance of the expert model according to the task vector, gradient and parameter mask matrix, sample the parameter importance to generate a sampling mask matrix, and use the sampling mask matrix to scale the task vector to generate an incremental task vector; The incremental model merging module is used to add the incremental task vector to the parameters of the pre-trained model to generate the parameters of the merged model, obtain the merged model, and complete the first incremental model merging; The hierarchical merging module is used to replace the pre-trained model and the expert model mentioned above with the merged model and the new expert model obtained from the previous merge each time when merging subsequent incremental models, and merge the new expert model into the merged model obtained from the previous merge using the same method as the first incremental model merging, until the incremental model merging of all expert models is completed in sequence.
7. An electronic device comprising a memory and one or more processors, wherein the memory is used to store a computer program, characterized in that: The processor is used to implement the incremental model merging method for a large language model as described in any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer, the incremental model merging method for a large language model described in any one of claims 1 to 5 is implemented.
9. A computer product comprising a computer program, characterized in that When the computer program is executed by a processor, the incremental model merging method for a large language model as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Content processing model integration method and related equipment
CN118094233A
Multi-task large language model training method and device
CN118261225A