Large model capability expansion method and system based on module fusion
Through module fusion technology, the parameters of the LoRA module are obtained and trained, combined with fixed parameters and trainable fusion modules, the performance balance and efficiency improvement of large models in multi-task learning are achieved, solving the problems of catastrophic forgetting and computing overhead.
Patent Information
- Application Number
- CN202510016701.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
AI Technical Summary
Existing continuous learning methods are prone to catastrophic forgetting when facing new tasks, and it is difficult to balance performance and efficiency in multi-task learning, resulting in increased computing and storage overhead.
Using a large model capability expansion method based on module fusion, by obtaining the parameters of the LoRA module, inserting the trainable LoRA module and training on new tasks, combining the fixed parameter LoRA module and the trainable LoRA fusion module, the fusion module is trained using the equalized sampling data set, and the parameters are merged to obtain the new LoRA module to achieve the integration of multi-task capabilities.
Without sacrificing original capabilities, quickly introduce new capabilities to large models, reduce catastrophic forgetting, reduce computing overhead, improve training and reasoning efficiency, and adapt to changes in different tasks sequences.
Smart Images

Figure CN120069059A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method and system for expanding the capabilities of large models based on module fusion. Background Art
[0002] The rapid development of artificial intelligence technology is profoundly changing the production and lifestyle. In particular, deep learning technology represented by pre-trained models has been widely applied in many fields such as military, communication, industry, and household life. Large-scale pre-trained models have become one of the core technologies in the development of artificial intelligence due to their excellent performance in various natural language processing tasks. However, in order to further improve the accuracy of the model, enable it to handle a wider range of tasks, and achieve a more general artificial intelligence system, the design of continuous learning or incremental learning algorithms becomes crucial. By continuously expanding the capabilities of the model from new task data, the model will be able to continuously adapt to and handle the changing and growing task requirements, thus better supporting practical applications in multiple fields and scenarios.
[0003] One of the main challenges in continuous learning is catastrophic forgetting, that is, the performance of the model trained on the original task will significantly decline after continuing to learn new tasks. To alleviate catastrophic forgetting, there are currently three main strategies: replay-based methods, regularization-based methods, and parameter isolation-based methods.
[0004] Replay-based methods reduce the forgetting of existing capabilities by accessing data from previous tasks. The core lies in determining which samples should be retained and how to use these samples to train the model. For example, iCaRL proposed by Rebuffi et al. (Rebuffi, Sylvestre-Alvise, et al. icarl: Incremental classifier and representation learning. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2017.) sets a fixed-size storage space for previous classification tasks, selects representative samples close to the center of data features, and mixes stored samples and new task data for training and recomputes the representative samples for each class when encountering a new classification task. Similarly, the GEM method proposed by Lopez-Paz D et al. (Lopez-Paz D, Ranzato M A. Gradient episodic memory for continual learning[J]. Advances in neural information processing systems, 2017, 30.) calculates the gradients of the model on old tasks during new task training to ensure that the model does not cause significant damage to the performance of old tasks while learning new tasks.
[0005] Regularization-based methods do not require saving the training data of the old task, but instead use regularization to constrain the direction of model parameter optimization. These methods believe that for different tasks, the effective performance space of the model is different, and the purpose of fine-tuning is to make the model parameters as close as possible to this effective subspace. The EWC method proposed by Kirkpatrick J et al. (Kirkpatrick J, Pascanu R, Rabinowitz N, et al. Overcoming catastrophic forgetting in neural networks[J]. Proceedings of the national academy of sciences, 2017, 114(13): 3521-3526.) measures which parameters are more important for the original task through the Fisher information matrix, and then reduces the impact of new task training on the performance of the old task by freezing these parameters. The LwF method proposed by Li Z et al. (Li Z, Hoiem D. Learning without forgetting[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 40(12): 2935-2947.) does not require using the data of the original task. By using the prediction results of the model on the new task as pseudo-labels, it constrains the model to not deviate too much from the optimization direction of the original task when optimizing the new task.
[0006] The parameter isolation method avoids catastrophic forgetting by adjusting independent model parameters for each downstream task. The PackNet method proposed by Mallya et al. (Mallya, Arun, and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2018.) selects and adjusts some model parameters through pruning when facing new tasks. Specifically, first, all parameters are used to train the current task. After completion, unimportant parameters are removed through pruning, and then training is restarted in the remaining parameter space. The parameters removed by pruning are added back for training subsequent tasks. The Expert Gate method proposed by Aljundi et al. (Aljundi, Rahaf, Punarjay Chakravarty, and Tinne Tuytelaars. Expert gate: Lifelong learning with a network of experts. Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.) fine-tunes different copies of the model for different tasks respectively to ensure the independence of parameters between models. An automatic selection of models for different tasks is achieved through a gating module, thereby aggregating the capabilities of different models. Finally, the method of ensemble learning is used to select the optimal path as the model prediction result.
[0007] However, there are still some challenges in the existing forgetting mitigation methods during the continuous learning process. The data replay method adds representative samples of old tasks to the training data of new tasks. Although it can alleviate catastrophic forgetting, it is difficult to balance the weights between old tasks and new tasks, and the model performance is sensitive to the task order. Therefore, this method usually relies on large-scale storage and computing resources and has poor stability. Although the regularization method has good theoretical effects, its implementation is complex and it is difficult to operate in practical applications. The parameter isolation method trains independent modules for each new task. Although it can effectively avoid forgetting, as the number of tasks increases, the storage overhead and computational complexity of the model will increase significantly, resulting in a decline in training and inference efficiency. Therefore, how to integrate the advantages of these methods and balance performance and efficiency during continuous fine-tuning has become the focus of current research. Summary of the Invention
[0008] The present invention discloses a large model capability expansion method and system based on module fusion, which can efficiently expand new task capabilities for large models without sacrificing original capabilities, reduce the overhead of large models in the continuous learning process, and improve training and reasoning efficiency.
[0009] To achieve the above objectives, the technical solution of the present invention includes the following contents.
[0010] A large model capability expansion method based on module fusion, the method comprising:
[0011] Get LoRA module A' t Parameters Wherein, the LoRA module A′ is included t The large model has the ability to solve t tasks at the same time;
[0012] Insert a trainable LoRA module into the large model M, and train the LoRA module on the training data set of the new task t+1 to obtain the LoRA module A t+1 Parameters Among them, the LoRA module A is included t+1 The large model has the ability to solve the new task t+1;
[0013] Insert the LoRA module A′ with fixed parameters into the large model M t , LoRA module with fixed parameters A t+1 And a trainable LoRA fusion module, and train the LoRA fusion module on a balanced sampling data set to obtain the parameters of the LoRA fusion module The balanced sampling data set is constructed based on the training data set of t tasks and the new task t+1;
[0014] Merge Parameters parameter and parameters Get LoRA module A' t+1 Parameters Wherein, the LoRA module A′ is included t+1 The large model has the ability to solve t tasks and new task t+1 at the same time.
[0015] Further, when t=1, obtain the LoRA module A′ t Parameters include:
[0016] Fixed parameter M of the large model M θ ;
[0017] Insert a trainable LoRA module into the large model M, and train the LoRA module on the training dataset of task t to obtain the LoRA module A suitable for task t t parameters where the training objective is to minimize the loss function of this task t
[0018] Take the said LoRA module A t and the parameters of this LoRA module A t parameters as the parameters of LoRA module A' t and the parameters of this LoRA module A' t parameters
[0019] Furthermore, the balanced sampling dataset is constructed based on the training datasets of t tasks and the new task t+1, including:
[0020] Obtain the maximum data volume L required for the balanced sampling dataset;
[0021] According to the maximum data volume L, calculate the data volume k allocated to each task t and the new task t+1;
[0022] Based on the data volume k, perform data sampling in the training datasets of each task t and the new task t+1 to construct the balanced sampling dataset.
[0023] Furthermore, insert the LoRA module A' with fixed parameters t , the LoRA module A with fixed parameters t+1 and a trainable LoRA fusion module into the large model, and train the LoRA fusion module on the balanced sampling dataset to obtain the parameters of the LoRA fusion module including:
[0024] Insert the said LoRA module A' t , the said LoRA module A t+1 and the trainable LoRA fusion module into the large model;
[0025] Fix the parameters M of the large model M θ ;
[0026] Based on the said parameters and the said parameters Fix the parameters of LoRA module A' t and LoRA module A t+1 respectively;
[0027] Train the LoRA fusion module on the balanced sampling dataset to obtain the parameters of the LoRA fusion module
[0028] Furthermore, the merging parameter parameter and parameter to obtain the parameters of LoRA module A′ t+1 parameters include:
[0029] Perform a dot product operation on parameter parameter and parameter to obtain parameter and use the said parameter as the parameters of LoRA module A′ t+1 parameters.
[0030] A large model capability extension system based on module fusion, the system includes:
[0031] The first parameter acquisition module is used to acquire the parameters of LoRA module A′ t parameters wherein, the large model containing the LoRA module A′ t has the ability to solve t tasks simultaneously;
[0032] The second parameter acquisition module is used to insert a trainable LoRA module into the large model M, and train the LoRA module on the training dataset of the new task t + 1 to obtain the parameters of LoRA module A t+1 parameters wherein, the large model containing the LoRA module A t+1 has the ability to solve the new task t + 1;
[0033] The third parameter acquisition module is used to insert a LoRA module A′ with fixed parameters t , a LoRA module A with fixed parameters t+1 and a trainable LoRA fusion module into the large model M, and train the LoRA fusion module on the balanced sampling dataset to obtain the parameters of the LoRA fusion module wherein, the balanced sampling dataset is constructed based on the training datasets of t tasks and the new task t + 1;
[0034] The parameter merging and acquisition module is used to merge parameter parameter and parameter to obtain the parameters of LoRA module A′ t+1 parameters wherein, the large model containing the LoRA module A′ t+1 has the ability to solve t tasks and the new task t + 1 simultaneously.
[0035] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for expanding the capabilities of the large model based on module fusion described in any one of the above is implemented.
[0036] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method for expanding the capabilities of the large model based on module fusion described in any one of claims 1-5 is implemented.
[0037] A computer program product, when the computer program product runs on a computer device, enables the computer device to execute the method for expanding the capabilities of the large model based on module fusion described in any one of the above.
[0038] Compared with the prior art, the present application has at least the following beneficial effects.
[0039] 1. Improve continuous learning ability: The present application focuses on improving the performance of large pre-trained models in continuous task learning. Different from conventional techniques that usually only focus on one aspect of efficiency or performance, this method can quickly introduce new capabilities to the large model under the premise of low computational overhead, and effectively alleviate the catastrophic forgetting phenomenon of the large model when learning new tasks.
[0040] 2. Innovative method combining parameter isolation and module fusion: The present application proposes an innovative method combining parameter isolation and module fusion. Specifically, first, the LoRA (Low-Rank Adaptation) activation module is independently fine-tuned on different tasks, and then the LoRA fusion module is adjusted through sampling and replay of task data, so as to achieve performance balance of the model in multi-task learning. Compared with conventional data replay strategies, this method not only performs better in terms of performance, but also has higher robustness and can adapt to different task order changes.
[0041] 3. Module fusion improves the flexibility and efficiency of ability expansion: Compared with previous parameter isolation technologies, this method integrates the capabilities of multiple tasks into a unified module through module fusion, thereby improving the flexibility and efficiency of ability expansion and avoiding the increase in storage overhead caused by the increase in tasks.
[0042] 4. Experimental verification: Experimental results show that the ability expansion method based on module fusion proposed in this study can effectively prevent the occurrence of catastrophic forgetting, and without compromising the performance of the original tasks, new capabilities such as code generation and session emotion recognition are added to the large model, significantly expanding the application scope of the model. Description of the Drawings
[0043] Figure 1 Schematic diagram of the large model capacity expansion strategy process proposed in this application.
[0044] Figure 2 Schematic diagram of the model structure adopted in this application.
[0045] Figure 3 Schematic diagram of the principle of the LoRA module used in this application. Detailed implementation
[0046] This application aims to clearly elaborate its purpose, technical solution and advantages, and will be described in detail in combination with the accompanying drawings.
[0047] The present invention adopts the idea of the parameter isolation method. By independently fine-tuning the LoRA (Low-Rank Adaptation) modules of each task and adjusting the LoRA fusion module through the sampling replay of task data, the ability balance in multi-task learning and the effective mitigation of catastrophic forgetting are achieved, and the ability to introduce new tasks for the fine-tuned large model can be efficiently realized.
[0048] The large model capacity expansion method based on module fusion of the present invention, as Figure 1 shown, includes the following steps 1 to 4.
[0049] Step 1: Obtain the parameters of LoRA module A' t wherein, the large model containing the LoRA module A' has the ability to solve t tasks simultaneously. t The LoRA module A' with the ability to solve t tasks described in this embodiment
[0050] is extended from being able to solve only one task t. When training the LoRA module for this task t, first, fix the parameters M of the base model (large model M) t , and insert a trainable LoRA module into the large model M. Then, train the LoRA module parameters separately on task t θ , and the optimization objective is to minimize the loss function of this task and optimize the LoRA module through the dataset D of this task . By minimizing the loss function, the model can perform better on task t. The goal of this step is to obtain the LoRA module parameters suitable for task t t . The optimization objective formula is as follows:
[0051]
[0052] wherein, is the loss function of task t, and D tis the dataset for task t. At this time, the present invention believes that the LoRA module parameters set with parameter are the LoRA module A' that can solve a specific task t , and the LoRA module A' t configured parameters
[0053] Step 2: Insert a trainable LoRA module into the large model M, and train the LoRA module on the training dataset of the new task t+1 to obtain the parameters of the LoRA module A t+1 wherein, the large model containing the LoRA module A has the ability to solve the new task t+1. t+1
[0054] When training the LoRA module for the new task t+1, the parameters of the base model M are still fixed θ , and a new trainable LoRA module is inserted into the large model M. Then, train the LoRA module parameters on the new task t+1 to obtain a LoRA module suitable for the new task. The optimization goal is to minimize the loss function of the new task t+1 and optimize the LoRA module parameters through the dataset D of this task t+1 . The optimization goal formula is as follows:
[0055]
[0056] wherein, is the loss function of the new task t+1, and D t+1 is the dataset of the new task t+1.
[0057] Step 3: Insert the LoRA module A' with fixed parameters t , the LoRA module A with fixed parameters t+1 and a trainable LoRA fusion module into the large model M, and train the LoRA fusion module on the balanced sampling dataset to obtain the parameters of the LoRA fusion module wherein, the balanced sampling dataset is constructed based on the training datasets of t tasks and the new task t+1.
[0058] Train the LoRA fusion module. In Steps 1 and 2, the LoRA module A' t and the LoRA module A t+1 have been obtained respectively. Next, fix the base model M θ and the LoRA module A' t and the LoRA module A t+1After that, a trainable LoRA fusion module is inserted into the model, and the parameters of the fusion module are trained on the balanced sampling dataset of t tasks and the new task t+1. By minimizing the loss function of the fusion module, a fusion module adapted to multi-task learning is obtained.
[0059] Specifically, the training data uses uniform sampling of all accessed tasks. For the accessed n different training task data (D 1 , D 2 ,..., D n ), the training data of each task is evenly distributed. Where L is the maximum amount of data that can be used under given conditions:
[0060]
[0061] This process can sample the training data of all tasks to maintain data balance between tasks. Thus, the optimization target formula of the fusion module is as follows:
[0062]
[0063] Step 4: Merge parameters Parameters and parameters are combined to obtain the parameters t+1 of the LoRA module A′ Among them, the large model containing the LoRA module A′ t+1 has the ability to solve t tasks and the new task t+1 simultaneously.
[0064] Merge the LoRA module and the fusion module. Combine the parameters of the LoRA module for task t trained in step 1, the LoRA module for task t+1 trained in step 2, and the LoRA fusion module trained in step 3 into a new LoRA module, so as to obtain an ability expansion module that can handle task t and task t+1 simultaneously, preparing for the introduction of new tasks in the future. The merged module can be expressed as
[0065]
[0066] In summary, this application proposes an efficient method for expanding the capabilities of large models, aiming to address the performance and efficiency challenges in continual learning. Without forgetting the capabilities of original tasks, this application optimizes the capability expansion of large-scale pre-trained models when introducing new tasks by introducing LoRA modules and fusion modules, maintaining high fine-tuning and inference efficiency. Through the sampling and replay of task data, this application effectively alleviates the problems caused by differences in the quality of different task data and reduces the difficulty of adapting to new tasks. Since the model structure designed in this solution does not contain non-linear layers, the combination of multiple LoRAs is equivalent to the product operation of parameter matrices, and this process can complete capability expansion without adding a large amount of computational overhead.
[0067] Next, a specific experiment is used to illustrate the method for expanding the capabilities of the large model provided by this application.
[0068] This application relates to a method for expanding the capabilities of large-scale pre-trained models in continual learning, aiming to alleviate the catastrophic forgetting problem through efficient module fusion techniques and introduce new task capabilities for large models. The model capability expansion algorithm of this application can be decomposed into the following four main steps: task preprocessing, independent module training, fusion module training, and module combination.
[0069] 1. Task preprocessing stage.
[0070] The purpose of the task preprocessing stage is to ensure the independence of data between different tasks, thus avoiding data interference in subsequent training. Specifically, this application first divides the training data into a total of N non-overlapping subsets (D 1 , D 2 , …, D N ), where each subset Di contains data samples related to task Ti. The training data set for each task should be divided according to the task characteristics to ensure the distinctiveness and independence between tasks. In addition, the task preprocessing stage also needs to define a loss function for each task The choice of the loss function is based on the nature of the task. For example, the cross-entropy loss is used for classification tasks, and the mean squared error loss is used for regression tasks. These loss functions are used to evaluate the performance of the model on their respective tasks and guide the optimization of model parameters.
[0071] 2. LoRA-based independent module training.
[0072] In this stage, this application uses the LoRA (Low-Rank Adaptation) method to fine-tune each task module of the large-scale pre-trained model (such as LLaMA). The design of the LoRA module reduces computational overhead and storage requirements by introducing low-rank matrices, ensuring that only a small part of the parameters are adjusted during fine-tuning, rather than the weights of the entire model. This process is as Figure 2 shown. Specifically, each task Ti For the corresponding independent LoRA module, during the fine-tuning process, by using the gradient descent optimization algorithm, the parameters of the LoRA module are adjusted to minimize the loss function on task T i and thus optimize the performance of the model on this task. The LoRA modules for each task are independently trained and do not have a direct impact on other task modules, as shown in the overall figure Figure 3 as follows
[0073] 3. Training of the fusion module
[0074] To balance the multi-task ability, this application introduces a LoRA fusion module to merge the LoRA modules of different tasks. The fusion module accepts the LoRA modules of multiple tasks and aggregates them into a unified representation through a fully connected layer. The rank value of the fusion module is usually set to the sum of the rank values of all input LoRA modules
[0075] When introducing a new task T i , this application inserts the independently trained LoRA modules and the previously fused LoRA modules into the corresponding layers of the base model M together and fixes their parameters. The fusion module is inserted to receive the concatenated output of multiple task modules and adjusts the parameters of the fusion module. During the training process of the fusion module, to ensure fairness of the model on each task, this application adopts an equalized data sampling method to ensure that the weights of the data for each task are the same during training
[0076] To verify the present invention, five common task ability dimensions are selected in this experiment: general question answering, mathematical reasoning, knowledge question answering, code generation, and conversation emotion recognition, as the evaluation criteria for the capabilities of the large model. For each ability, representative tasks are selected to verify the continuous learning performance of the ability expansion method designed in this application on multiple tasks. The specific datasets selected are shown in Table 1
[0077] Table 1 Details of the multi-task continuous learning dataset
[0078]
[0079] To verify the effectiveness of the module fusion method in continuous learning, this application conducts experiments on the LLaMA-7B model, and uses the above datasets to sequentially introduce AlpacaGPT4, Camel Math, MMLU, Code Contests, and MELD training data for continuous parameter-efficient fine-tuning. In all experiments, the Batch-size is fixed at 512, the maximum learning rate is set to 5e-4, the LoRA rank assigned to each task is 8, LoRA Alpha is set to 32, and LoRA Dropout is set to 0.05
[0080] In continuous learning, 1K representative samples are selected for each task category as the validation set to determine whether the model converges. Experimental verification shows that the module fusion method proposed in this application can effectively alleviate the forgetting phenomenon and efficiently expand new capabilities. The results show that the continuous learning performance of this application on tasks involving 5 different ability dimensions exceeds previous methods.
[0081] This application also compares the efficiency and performance of different continuous learning methods, evaluates three baseline methods including continuous fine-tuning and data mixing fine-tuning, and three forgetting mitigation methods: iCaRL based on data replay, GEM based on regular gradient constraint, and Adapter Fusion based on isolated parameter fusion. Trained on the aforementioned multi-task dataset, the overall training time is used to compare the efficiency of different methods, and the specific results are shown in Table 2.
[0082] Table 2 Cumulative training time of different continuous learning methods in LLaMA-7B
[0083]
[0084] Compared with other continuous learning methods, the training time of this application is relatively low. At the same time, since the LoRA module obtained after training can be merged into the weights of the base model, additional inference time is avoided, fully demonstrating the efficiency of the design method of this application.
[0085] Table 3 Average performance of different continuous learning methods in LLaMA-7B
[0086]
[0087] After parameter-efficient fine-tuning on multiple task datasets, the model performance of the method proposed in this application is compared with other methods. The experimental results are shown in Table 3, which shows that the overall performance of the method of this application is close to the data mixing method, but the performance loss on each task is small. This indicates that the module fusion method proposed in this application can effectively expand new capabilities for the model continuously and maintain good performance during the expansion process, avoiding the common catastrophic forgetting problem in conventional methods.
[0088] The above description only elaborates on a specific example of this application and does not impose any restrictions on this application. Obviously, for those with professional knowledge in this field, once they understand the content and principle of this application, it is possible to make various modifications and changes in form and details without violating the original principle and structure of this application. However, these corrections and changes based on the idea of this application are still considered to be within the protection scope of the claims of this application.
Claims
1. A large model capability expansion method based on module fusion, characterized in that: The method comprises: Get LoRA module A' t Parameters Wherein, the LoRA module A′ is included t The large model has the ability to solve t tasks at the same time; Insert a trainable LoRA module into the large model M, and train the LoRA module on the training data set of the new task t+1 to obtain the LoRA module A t+1 Parameters Among them, the LoRA module A is included t+1 The large model has the ability to solve the new task t+1; Insert the LoRA module A′ with fixed parameters into the large model M t , LoRA module with fixed parameters A t+1 And a trainable LoRA fusion module, and train the LoRA fusion module on a balanced sampling data set to obtain the parameters of the LoRA fusion module The balanced sampling data set is constructed based on the training data set of t tasks and the new task t+1; Merge Parameters parameter and parameters Get LoRA module A' t+1 Parameters Wherein, the LoRA module A' is included t+1 The large model has the ability to solve t tasks and new task t+1 at the same time.
2. The method according to claim 1, characterized in that At t=1, obtain LoRA module A' t Parameters include: Fixed parameter M of the large model M θ ; Insert a trainable LoRA module into the large model M, and train the LoRA module on the training data set of task t to obtain a LoRA module A suitable for task t t Parameters Among them, the training goal is to minimize the loss function of the task t The LoRA module A t And the LoRA module A t Parameters As LoRA module A' t And the LoRA module A′ t Parameters 3. The method according to claim 1, characterized in that The balanced sampling data set is constructed based on the training data set of t tasks and the new task t+1, including: Get the maximum amount of data L required for a balanced sampling data set; According to the maximum data volume L, calculate the data volume k allocated to each task t and the new task t+1; Based on the data volume k, data sampling is performed in the training data sets of each task t and the new task t+1 to construct the balanced sampling data set.
4. The method according to claim 1, characterized in that: The LoRA module A′ with fixed parameters is inserted into the large model M. t , LoRA module with fixed parameters A t+1 And a trainable LoRA fusion module, and train the LoRA fusion module on a balanced sampling data set to obtain the parameters of the LoRA fusion module include: The LoRA module A' t 、The LoRA module A t+1 and a trainable LoRA fusion module inserted into the large model; Fixed parameter M of the large model M θ ; Based on the parameters and the parameters Fix the LoRA module A' separately t And LoRA module A t+1 Parameters; Train the LoRA fusion module on the balanced sampling data set to obtain the parameters of the LoRA fusion module 5. The method according to claim 1, characterized in that The merge parameters parameter and parameters Get LoRA module A' t+1 Parameters include: Parameters parameter and parameters Perform the dot multiplication operation to get the parameters And the parameters As LoRA module A' t+1 Parameters.
6. A large model capability expansion system based on module fusion, characterized in that: The system comprises: The first parameter acquisition module is used to obtain the LoRA module A' t Parameters Wherein, the LoRA module A' is included t The large model has the ability to solve t tasks at the same time; The second parameter acquisition module is used to insert a trainable LoRA module into the large model M and train the LoRA module on the training data set of the new task t+1 to obtain the LoRA module A. t+1 Parameters Among them, the LoRA module A is included t+1 The large model has the ability to solve the new task t+1; The third parameter acquisition module is used to insert the LoRA module A' with fixed parameters into the large model M t , LoRA module with fixed parameters A t+1 And a trainable LoRA fusion module, and train the LoRA fusion module on a balanced sampling data set to obtain the parameters of the LoRA fusion module The balanced sampling data set is constructed based on the training data set of t tasks and the new task t+1; Parameter merging acquisition module, used to merge parameters parameter and parameters Get LoRA module A' t+1 Parameters Wherein, the LoRA module A' is included t+1 The large model has the ability to solve t tasks and new task t+1 at the same time.
7. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the large model capability expansion method based on module fusion as described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the large model capability expansion method based on module fusion as described in any one of claims 1 to 5.
9. A computer program product, characterized in that When the computer program product runs on a computer device, the computer device executes the large model capability expansion method based on module fusion as described in any one of claims 1 to 5.