Fine-tuning Method for Multi-task Mixture-of-Experts Model Based on General LoRA and Domain-specific LoRA
By adopting the multi-tasking hybrid expert model BENLoRA based on general LoRA and domain-specific LoRA in natural language processing, the problems of stability, scalability and high computing resource consumption during multi-task and domain adaptation in the prior art are solved, and efficient multi-task learning and new domain adaptation are achieved.
Patent Information
- Application Number
- CN202411646618.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The prior art has problems such as stability when dealing with multitasking and domain adaptation, lack of domain-specific adapters, limited scalability, large computing resource consumption and general capability loss.
BENLoRA, a multi-task hybrid expert model based on general LoRA and domain-specific LoRA, is adopted to achieve efficient parameter multi-task learning and new domain adaptation by integrating general experts and domain-specific experts and introducing residual connections and dynamic routing.
The training efficiency and parameter efficiency of the model in multi-task learning are improved, the universality and adaptability of the base model are enhanced, and efficient model training and flexible task and domain adaptation are achieved.
Smart Images

Figure CN119398122B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a multi-task hybrid expert model fine-tuning method based on general LoRA and domain-specific LoRA. Background Art
[0002] In recent years, significant progress has been made in the field of natural language processing. Open-source large language models such as LLaMA, Qwen, and Yi have emerged continuously and achieved certain success in natural language processing tasks. However, the complexity of these models and the rapid growth of parameter scales have made efficient training under limited computing resources a challenge. To address this issue, parameter-efficient training (PEFT) methods such as LoRA (Low-Rank Adaptation) have been proposed. LoRA improves performance by using low-rank matrices, but there are stability problems when dealing with large and mixed datasets. MoLoRA improves this problem by extending LoRA and integrating the Mixture-of-Expert (MoE) architecture, but there are still some limitations.
[0003] For example, one limitation of MoLoRA is the "lack of domain-specific LoRA experts". In the MoLoRA architecture, all experts are updated simultaneously during training. This consistency may limit its performance ceiling, especially for tasks with significant differences such as mathematics and code. Conversely, adding domain-specific LoRA experts may potentially improve performance. Additionally, another problem with this architecture is the "limited pluggability": when inserting new domain experts into MoLoRA, all expert parameters need to be retrained with a dataset containing new domain data, which may result in low efficiency and high time consumption.
[0004] However, MoLoRA has the following disadvantages:
[0005] (1) Lack of domain-specific adapters. Using the same experts for all tasks may limit the performance ceiling. Especially for tasks with large differences, such as mathematics and code tasks, the lack of domain-specific experts may prevent the full performance from being achieved.
[0006] (2) Limited scalability. The ability to add new tasks requires retraining all parameters, which is not only inefficient and time-consuming but also increases the training cost, limiting the flexibility and scalability of the model in practical applications.
[0007] The main disadvantages of the prior art include:
[0008] 1. Dataset fusion problem: It is difficult to fuse multiple datasets to construct a multi-functional dataset. There may be contradictions between different datasets, and it is difficult to evaluate data quality.
[0009] 2. Performance degradation: When training on a fused dataset, the performance of large language models may degrade, especially when dealing with tasks in specific domains.
[0010] 3. High computational resource consumption: Training large language models completely requires a large amount of computational resources and time, which limits the rapid adaptation and deployment of the models.
[0011] 4. Loss of general capabilities: When performing domain adaptation, the model may sacrifice its general language understanding capabilities, resulting in degraded performance on tasks in non-target domains.
[0012] 5. Lack of scalability: Existing methods are difficult to flexibly add new tasks or domains and usually require retraining the entire model.
[0013] 6. Lack of effective domain experts: In the MoLoRA architecture, all experts update their parameters simultaneously during training. This consistency may limit its performance ceiling.
[0014] Developing an efficient multi-task learning and domain adaptation method is crucial for the application and development of large language models for the following reasons:
[0015] a) Improving model generalization ability: By training on multiple tasks and domains, the model can learn a wider range of knowledge and skills, improving its performance on unseen tasks.
[0016] b) Resource utilization efficiency: Compared with training separate models for each task or domain, multi-task learning can utilize computational resources and storage space more effectively.
[0017] c) Quick adaptation to new domain tasks: A model with good multi-task learning ability can adapt to new domains or tasks more quickly, reducing the development time for new application scenarios.
[0018] d) Maintaining general capabilities: When performing domain adaptation, it is important to maintain the model's general language understanding capabilities so that it can perform well in specific domains without losing the ability to handle general tasks.
[0019] e) Industrial application requirements: In practical applications, it is often necessary for the model to be able to handle multiple types of tasks simultaneously, such as question answering, dialogue, summarization, etc., which requires the model to have strong multi-task learning capabilities.
[0020] Therefore, developing a method that can effectively perform multi-task learning and domain adaptation while maintaining the model's general capabilities has important theoretical significance and practical application value. Summary of the Invention
[0021] To overcome the deficiencies of the prior art, the present invention provides a fine-tuning method for a multi-task hybrid expert model based on general LoRA and domain-specific LoRA, which integrates general experts and domain-specific experts and introduces residual connections to balance the model's processing capabilities for general tasks and domain-specific tasks, improving performance and stability. The present invention realizes parameter-efficient multi-task learning and new domain adaptation.
[0022] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0023] Step 1: Construct a multi-task hybrid expert model BENLoRA (Blended Enhanced Network with LoRA for Multiple Domain Robust Adaptation) based on general LoRA and domain-specific LoRA. The model architecture includes the following components:
[0024] a) Base language model: The large language model after pre-training is used as the base model;
[0025] b) General LoRA expert: Used to capture general knowledge and capabilities across domains;
[0026] c) Domain-specific LoRA expert: A specialized adapter for a specific domain or task;
[0027] d) Residual connection: Connects the general LoRA expert and the domain-specific LoRA expert to maintain the general capabilities of the model;
[0028] e) Routing: Used to dynamically allocate the weights of different domain-specific LoRA experts;
[0029] Step 2: Adopt a three-stage training process;
[0030] a) The first stage: Only train the general LoRA expert, while the domain-specific LoRA expert and the routing are deactivated;
[0031] b) The second stage: Train each domain-specific LoRA expert separately, and keep the general expert parameters frozen;
[0032] c) The third stage: The parameters of all LoRA experts are frozen, and only the routing is trained to learn the best combination strategy for different tasks; By training the routing, the large model can dynamically allocate the weights of different experts according to the needs of different tasks;
[0033] Step 3: Introduce residual connections;
[0034] The input of the domain-specific LoRA expert includes not only the original input but also the output of the general LoRA expert, expressed as:
[0035]
[0036] where h m is the output of the general LoRA expert; the superscript r on the letter represents BENLoRA-Res; W0 represents the frozen weights of the base large language model, which is a matrix; x m represents the hidden state vector input to the BENLoRA-Res module; represents the weights assigned by the routing module in the BENLoRA-Res module to each expert. B represents the B matrix in each LoRA expert in BENLoRA-Res; A represents the A matrix in each LoRA expert in BENLoRA-Res; a LoRA expert contains two matrices A and B. During the operation, the vector first passes through A and then through B; BENLoRA contains multiple experts. i is the subscript of any one expert, which is universal; the summation symbol represents the unified summation of the operation results of these experts without distinction;
[0037] Step 4: Dynamic routing;
[0038] Use a trainable router to dynamically allocate the weights of different LoRAs. The calculation of the router weights is as follows:
[0039]
[0040] where is the weight assigned by the routing module to the i-th expert, represents the weight matrix of the routing module in the BENLoRA-Res module. The capital R represents Router, represents the process by which the routing module obtains the weights of each LoRA expert based on the input hidden state vector;
[0041] The routing module contains the weight matrix and the softmax function. The hidden state vector x m first passes through and then passes through the softmax function. The output of the softmax is a vector. The subscript i in the parentheses represents the i-th element in this vector, that is, the weight of each LoRA expert
[0042] Preferably, the large language model is LLaMA-2.
[0043] A computer program that causes a computer to execute the above multi-task mixture of experts model fine-tuning method.
[0044] An electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above multi-task hybrid expert model fine-tuning method.
[0045] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above multi-task hybrid expert model fine-tuning method is implemented.
[0046] A chip, comprising: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the above multi-task hybrid expert model fine-tuning method.
[0047] A computer program product, the computer program product includes a computer storage medium, the computer storage medium stores a computer program, the computer program includes instructions that can be executed by at least one processor, and when the instructions are executed by the at least one processor, the above multi-task hybrid expert model fine-tuning method is implemented.
[0048] The beneficial effects of the present invention are as follows:
[0049] 1. The present invention provides a multi-task hybrid expert model based on general LoRA and domain-specific LoRA, which realizes parameter-efficient multi-task learning and new domain adaptation.
[0050] 2. The present invention avoids the problem of dataset fusion by separately training general LoRA and domain-specific LoRA, and improves the performance of the model in each domain.
[0051] 3. The present invention uses LoRA technology to reduce the consumption of computing resources and realizes efficient model training.
[0052] 4. The present invention designs a residual connection structure to maintain the original general ability of the base model while performing domain adaptation.
[0053] 5. The present invention provides a flexible training paradigm, which can conveniently add new tasks or domain experts without retraining the parameters of all experts. Description of the Drawings
[0054] Figure 1 For the structural comparison of MoLoRA and the BENLoRA-Flan and BENLoRA-Res of the present invention;
[0055] Figure 2 For the detailed schematic diagram of the three-stage fine-tuning paradigm of BENLoRA of the present invention. Detailed Embodiments
[0056] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0057] In natural language processing, as the scale of large language models (LLMs) continues to expand and the application scenarios become increasingly diverse, how to achieve efficient training with limited computing resources to meet the needs of multi-task learning has become an important challenge currently faced. Traditional training methods often exhibit insufficient stability in multi-task learning and have a large demand for training resources. In addition, some existing methods such as MoLoRA and SiRA have the problem of lacking domain-specific adapters, that is, when dealing with different tasks, especially those with large differences (such as math and code tasks), the performance may be limited. At the same time, MoLoRA needs to retrain all parameters when adding new task capabilities, which leads to problems of low efficiency and limited scalability.
[0058] The object of the present invention is to provide a multi-task hybrid expert model based on general LoRA and domain-specific LoRA to improve the training efficiency and parameter efficiency of large language models in multi-task learning, enhance the generality and adaptability of the base model, and enable it to better handle various complex natural language processing tasks.
[0059] The present invention proposes a multi-task hybrid expert model BENLoRA based on general LoRA and domain-specific LoRA, which mainly includes the following steps:
[0060] 1. Model architecture;
[0061] The model architecture proposed by the present invention includes the following components:
[0062] a) Base language model: A pre-trained large language model (such as LLaMA-2) is used as the base model.
[0063] b) General LoRA expert: Used to capture general knowledge and capabilities across domains.
[0064] c) Domain-specific LoRA expert: A specialized adapter for a specific domain or task.
[0065] d) Residual connection: Connects the general LoRA and the domain-specific LoRA to maintain the general capabilities of the model.
[0066] e) Routing: Used to dynamically allocate the weights of different domain-specific LoRA experts.
[0067] 2. Training process;
[0068] The present invention adopts a three-stage training process:
[0069] a) First stage: Only train the general expert to quickly adapt to general tasks while deactivating the domain-specific experts and the router. This enables the general expert to learn task-agnostic representations and provides a basic understanding ability for subsequent tasks.
[0070] b) Second stage: Train each domain-specific expert individually to focus on its corresponding task while keeping the general expert's parameters frozen. At this time, the domain-specific experts can focus on learning the specific knowledge of their respective domains and improve the model's performance in specific domains.
[0071] c) Third stage: Freeze the parameters of all experts and only train the router to learn the optimal combination strategy for different tasks. By training the router, the model can dynamically allocate the weights of different experts according to the requirements of different tasks, thereby achieving optimal performance.
[0072] BENLoRA-Flan is based on the BENLoRA paradigm and separates the experts in MoLoRA into general experts and domain-specific experts. During training, in the first stage, the general expert is trained on a general dataset, in the second stage, the domain-specific experts are trained on their respective domain-specific datasets, and in the third stage, all expert parameters are frozen and only the router is trained.
[0073] BENLoRA-Res is a more advanced method that integrates general experts and domain-specific experts and introduces residual connections to balance the model's processing capabilities for general tasks and domain-specific tasks, improving performance and stability. In BENLoRA-Res, first, the general expert is used to calculate the hidden state vector, then the hidden vector is refined by the domain-specific expert, and the output of the general expert is directly incorporated into the final result through a residual connection to ensure that key information is retained and enhance the model's robustness.
[0074] 3. Residual connection;
[0075] To maintain the general capabilities of the model, the present invention introduces a residual connection. Specifically, the input of the domain-specific LoRA includes not only the original input but also the output of the general LoRA, expressed as:
[0076]
[0077] where h m is the output of the general LoRA, and E r (·) represents the operation of the domain-specific LoRA expert.
[0078] 4. Dynamic routing;
[0079] The present invention uses a trainable router to dynamically allocate the weights of different LoRAs. The calculation of the router weights is as follows:
[0080]
[0081] where is the weight of the i-th expert.
Claims
1. A fine-tuning method based on general LoRA and domain-specific LoRA multi-task hybrid expert model, characterized in that: The steps include: Step 1: Build a multi-task hybrid expert model BENLoRA based on general LoRA and domain-specific LoRA. The model architecture includes the following components: a) Base language model: a pre-trained large language model is used as the base model; b) General LoRA Expert: used to capture common knowledge and capabilities across domains; c) Domain-specific LoRA Experts: Specialized adapters for specific domains or tasks; d) Residual connection: connects general LoRA experts and domain-specific LoRA experts to maintain the general capabilities of the model; e) Routing: used to dynamically assign weights to dedicated LoRA experts in different fields; Step 2: Use a three-stage training process; a) Phase 1: Only general LoRA experts are trained, while domain-specific LoRA experts and routers are deactivated; b) Phase 2: Train each domain-specific LoRA expert separately, and keep the general expert parameters frozen; c) Phase 3: The parameters of all LoRA experts are frozen, and only the routing is trained to learn the best combination strategy for different tasks; by training the routing, the large model can dynamically assign weights to different experts according to the needs of different tasks; Step 3: Introduce residual connection; The input of the domain-specific LoRA expert includes not only the original input but also the output of the general LoRA expert, expressed as: where h m is the output of the general LoRA expert; the letter r in the upper right corner represents BENLoRA-Res; W0 represents the frozen weight of the base large language model, which is a matrix; x m Represents the hidden state vector of the input BENLoRA-Res module; Represents the weight assigned to each expert by the routing module in the BENLoRA-Res module, B represents the B matrix in each LoRA expert in BENLoRA-Res; A represents the A matrix in each LoRA expert in R in BENLoRA-Res; a LoRA expert contains two matrices A and B, and the vector passes through A first and then B during operation; BENLoRA contains multiple experts, i is the subscript of any expert, and is universal; the summation symbol indicates that the operation results of these experts are uniformly summed without distinction; Step 4: Dynamic routing; Use a trainable router to dynamically assign weights to different LoRAs. The routing weight is calculated as follows: in is the weight assigned to the i-th expert by the routing module, Represents the weight matrix of the routing module in the BENLoRA-Res module. The capital R stands for Router. It represents the process of the routing module obtaining the weight of each LoRA expert according to the input latent vector; The routing module contains the weight matrix And the softmax function, the hidden vector x m First pass After that, the output of the softmax function is a vector. The i in the lower right corner of the bracket represents the i-th element in the vector, which is the weight of each LoRA expert.
2. According to claim 1, a method for fine-tuning a multi-task hybrid expert model based on general LoRA and domain-specific LoRA is characterized in that: The large language model is LLaMA-2.
3. A computer program, characterized in that The computer program enables a computer to execute the method according to any one of claims 1 to 2.
4. An electronic device, characterized in that: include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method as claimed in any one of claims 1 to 2.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.
6. A chip, characterized in that: include: A processor, configured to call and run a computer program from a memory, so that a device equipped with the chip executes a method as claimed in any one of claims 1 to 2.
7. A computer program product, characterized in that The computer program product comprises a computer storage medium storing a computer program, wherein the computer program comprises instructions executable by at least one processor, and when the instructions are executed by the at least one processor, the method according to any one of claims 1 to 2 is implemented.
Citation Information
Patent Citations
Super-relation knowledge extraction method and device based on fine-tuning large model
CN118171732A
Method and system for improving structure of language model based on hybrid expert model
CN118194917A