A multi-task learning method for large language models based on adaptive activation scaling

By introducing LoRA module and multi-task fusion scaling network into the large language model, adaptive activation of scaling adaptation is achieved, and the problems of high computing costs and conflicts between tasks in multi-task learning of large language models are solved, and the efficiency and performance of multi-task learning are improved.

CN119539004BActive Publication Date: 2025-05-13BEIJING SCI & TECH PATENT OFFICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411591529.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-05-13
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

In multi-task learning of large language models, there are high computational costs and seesaw problems between tasks, resulting in uneven performance of the model on multiple tasks.

Method used

A large language model multi-task learning method based on adaptive activation scaling adaptation is adopted. By introducing LoRA module and multi-task fusion scaling network, knowledge is shared and adaptive scaling is adaptive to alleviate conflicts between tasks and improve resource utilization efficiency.

Benefits of technology

It realizes efficient multi-task learning and optimization under limited resources, alleviates the seesaw problem between tasks and improves the overall performance of the model on multiple tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119539004B_ABST
    Figure CN119539004B_ABST
Patent Text Reader

Abstract

The present invention discloses a large language model multi-task learning method based on adaptive activation scaling adaptation, belonging to the technical field of large language models. The method comprises: constructing a model; initializing a learnable activation scaling adaptation vector for each task k; constructing a multi-task joint fine-tuning training data set; optimizing LoRA module parameters, multi-task fusion scaling network parameters and learnable activation scaling adaptation vectors using the multi-task joint fine-tuning training data set to generate a trained model. The present invention alleviates the seesaw problem between different tasks and realizes efficient multi-task learning and optimization using limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to a large language model multi-task learning method based on adaptive activation scaling adaptation. Background Art

[0002] As large language models (LLMs) have grown significantly in size and complexity in the field of natural language processing (NLP), it has become increasingly critical to fine-tune them to perform well in specific applications. Fine-tuning is the process of taking a pre-trained LLM and further training it on a smaller task-specific dataset. This enables the model to adapt its understanding of general language to the nuances and requirements of a specific task. But as the number of parameters of LLMs increases, full parameter fine-tuning also becomes computationally expensive. PEFT methods, such as LoRA (Low-rank Adaptation), reduce the computational cost by training only a small number of additional parameters while maintaining the adaptability of the model.

[0003] In the fine-tuning of large language models, the advantages of multi-task learning are gradually reflected, mainly in the following aspects.

[0004] Shared knowledge: By training on multiple tasks, these models can leverage shared knowledge and commonalities across tasks, improving overall performance and efficiency.

[0005] Resource efficiency: Compared with training a separate model for each task, a single multi-task LLM can handle a variety of tasks, reducing the computational resources and time required for training.

[0006] Improved generalization: Multi-task learning helps the model generalize better across different tasks and domains because it learns to handle diverse inputs and outputs during pre-training.

[0007] Multi-task fine-tuning aims to adapt pre-trained LLMs to various tasks in one or more specific fields, such as code interpretation, data analysis, medical and product recommendations, etc. In this research direction, there are two main challenges: high computational cost and the seesaw problem between tasks (i.e., the improvement of one task's performance may lead to the decline of another task's performance). Specifically, they include:

[0008] In multi-task learning, we try to share knowledge between different tasks through hard parameter sharing, Cross-stitich Networks, Sluice Networks and other methods to improve the generalization ability of the model. However, these methods may face the problem of negative transfer, that is, the learning of one task may have a negative impact on another task.

[0009] The hybrid expert model combines task-general and task-specific experts through a gating mechanism to adapt to the needs of different tasks. This approach has shown effectiveness in multi-task learning, but it also suffers from the problem of negative transfer caused by parameter sharing.

[0010] Based on the above, current fine-tuning of LLMs poses several challenges:

[0011] Inter-task goal conflict: Fine-tuning a model on multiple tasks can lead to inter-task interference, where improvements in one task degrade performance on another. This is particularly problematic in multi-task learning settings, where the model needs to balance various objectives.

[0012] Parameter complexity: LLMs typically have billions of parameters, and comprehensive parameter fine-tuning is computationally expensive. This requires a lot of hardware resources, such as high-end GPUs or TPUs, and a lot of training time.

[0013] Overfitting: Fine-tuning on a small dataset increases the risk of overfitting, where the model performs well on the training data but poorly on unseen data. This limits the model’s ability to generalize.

[0014] Generalization: Strike a balance between specialization (performing well on a specific task) and generalization (maintaining general language understanding).

[0015] Data distribution changes: When fine-tuning a task in a specific domain, the data distribution may be significantly different from the pre-training data. This requires effective domain adaptation strategies to ensure that the model adapts to the new environment.

[0016] Task-specific features: Identify and leverage task-specific features without losing the general language understanding gained during pre-training. Summary of the invention

[0017] The present invention discloses a large language model multi-task learning method based on adaptive activation scaling adaptation, which can alleviate the seesaw problem between different tasks and realize efficient multi-task learning and optimization using limited resources.

[0018] In order to achieve the above-mentioned object of the invention, the technical solution of the present invention includes the following contents.

[0019] A large language model multi-task learning method based on adaptive activation scaling adaptation, the method comprising:

[0020] Constructing a model, the model comprising: a large language base model, a LoRA module and a multi-task fusion scaling network, the large language base model comprising an attention module and a feedforward neural network, the LoRA module being introduced into a query linear transformation matrix, a key linear transformation matrix and a value linear transformation matrix of the attention module, and the multi-task fusion scaling network being used to calculate a scaling factor according to an input of the model;

[0021] Initialize the learnable activation scaling adaptation vector l for each task k k ; Wherein, the learnable activation scaling adaptation vector l k Includes: Learnable activation scaling adaptation vectors for feed-forward networks Learnable activation scaling adaptation vector for attention values and learnable activation scaling adaptation vectors for attention keys

[0022] Construct a multi-task joint fine-tuning training dataset;

[0023] The multi-task joint fine-tuning training data set is used to adjust the LoRA module parameters, the multi-task fusion scaling network parameters and the learnable activation scaling adaptation vector l k to generate a trained model; wherein, during the optimization process, based on the scaling factor and the learnable activation scaling adaptation vector l k The intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module are scaled and adapted.

[0024] Furthermore, the structure of the multi-task fusion scaling network includes: N layers of linear transformation, nonlinear activation and Dropout.

[0025] Furthermore, the initialization of the learnable activation scaling adaptation vector l for each task k k ,include:

[0026] Construct a training data set train for each task k k ;

[0027] Freeze the parameters of the large language base model;

[0028] On top of the large language base model with frozen parameters, the learnable activations for each task k are scaled to an adaptation vector l k Train and update respectively to obtain the learnable activation scaling adaptation vector l k The initial value of .

[0029] Furthermore, the scaling factor includes: a scaling factor component Scaling of the feedforward neural network FFN , Scaling factor component of attention valueV Scaling factor component of the attention key K .

[0030] Further, based on the scaling factor and the learnable activation scaling adaptation vector l k Scaling and adapting the intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module, including:

[0031] Scale the learnable activations of each task k to the adaptation vector Learnable activation scaling adaptation vector and learnable activation scaling adaptation vector Connect them separately to obtain the learnable activation scaling adaptation vector l FFN , the learnable activation scaling adaptation vector l V and the learnable activation scaling adaptation vector l K ;

[0032] According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l FFN , generate the fused activation scaled adaptation vector L FFN ;

[0033] According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l V , generate the fused activation scaled adaptation vector L V ;

[0034] According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l K , generate the fused activation scaled adaptation vector L K ;

[0035] Perform the fusion activation scaling adaptation vector L FFN Bitwise multiplication with the intermediate hidden layer output P of the feedforward neural network to obtain the scaled intermediate hidden layer output P′;

[0036] Perform the fusion activation scaling adaptation vector L V Bitwise multiplication with the attention value V in the attention module to obtain a scaled attention value V′;

[0037] Perform the fusion activation scaling adaptation vector L K The bitwise multiplication with the attention key K in the attention module obtains the scaled attention key K′.

[0038] A large language model multi-task learning system based on adaptive activation scaling adaptation, the system comprising:

[0039] A model construction module is used to construct a model, the model comprising: a large language base model, a LoRA module and a multi-task fusion scaling network, the large language base model comprises an attention module and a feedforward neural network, the LoRA module is introduced into the query linear transformation matrix, the key linear transformation matrix and the value linear transformation matrix of the attention module, and the multi-task fusion scaling network is used to calculate a scaling factor according to the input of the model;

[0040] Vector initialization module, used to initialize the learnable activation scaling adaptation vector l for each task k k ; Wherein, the learnable activation scaling adaptation vector l k Includes: Learnable activation scaling adaptation vectors for feed-forward networks Learnable activation scaling adaptation vector for attention values and learnable activation scaling adaptation vectors for attention keys

[0041] Dataset construction module, used to construct multi-task joint fine-tuning training dataset;

[0042] A model training module is used to use the multi-task joint fine-tuning training data set to adjust the LoRA module parameters, the multi-task fusion scaling network parameters and the learnable activation scaling adaptation vector l k to generate a trained model; wherein, during the optimization process, based on the scaling factor and the learnable activation scaling adaptation vector l k The intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module are scaled and adapted.

[0043] An electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for multi-task learning of a large language model based on adaptive activation scaling adaptation described in any one of the above items is implemented.

[0044] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, a large language model multi-task learning method based on adaptive activation scaling adaptation as described in any of the above items is implemented.

[0045] Compared with the prior art, the present invention has at least the following beneficial effects.

[0046] 1) Combining the LoRA method and the "Multi-task Adaptive Activation Scaling Adapter", a large-model multi-task learning model architecture is designed and proposed. By sharing LoRA, the model can learn the common knowledge between different tasks and use cross-task shared knowledge. Through the multi-task adaptive activation scaling adapter, the characteristics of different tasks are fully learned, and the seesaw problem between different tasks can be alleviated, so that the performance of a model on multiple tasks can be improved together.

[0047] 2) A multi-task joint fine-tuning method for a large language model based on adaptive activation scaling adaptation is proposed. Through the optimization of extremely low parameters, efficient multi-task learning and optimization using limited resources is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 Flowchart of the multi-task learning method for large language models based on adaptive activation scaling.

[0049] Figure 2 Diagram of a scaling adapter activated for adaptiveness.

[0050] Figure 3 Learning the overall network structure for multi-task.

[0051] Figure 4 Activates scaled adaptation vector optimization for specific tasks. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below through specific implementations and in conjunction with the accompanying drawings.

[0053] The present invention provides a large language model multi-task learning method based on adaptive activation scaling adaptation, such as Figure 1 As shown, the following steps are included.

[0054] Step 1: Build the model.

[0055] Since LLMs usually have billions of parameters, fine-tuning all parameters of these models becomes extremely expensive and impractical, and the limited available data may lead to overfitting. Many domains in the real world contain a large number of heterogeneous tasks, which may have different degrees of similarity or even conflict in their objectives. In order to efficiently and fully learn the characteristics of different tasks using limited resources.

[0056] like Figure 2 and Figure 3As shown in FIG. 1 , the model mainly includes three parts: a large language base model, a LoRA module, and a multi-task fusion scaling network. The large language base model includes an attention module and a feedforward neural network, the LoRA module is introduced into the query linear transformation matrix, the key linear transformation matrix, and the value linear transformation matrix of the attention module, and the multi-task fusion scaling network is used to calculate a scaling factor according to the input of the model.

[0057] A. Large language model.

[0058] Figure 2 The Transformer Block on the far left is the core Transformer module in the large language base model, which includes an attention module (below) and a feedforward neural network (above).

[0059] Multi-Head Attention is a core concept in the Transformer model and one of the key factors that enables the Transformer architecture to capture complex relationships between different positions in sequence data. In natural language processing (NLP) tasks, such as translation or text summarization, it is very important to understand the relationships between different words in a sentence. Traditional recurrent neural networks (RNNs) and long short-term memory networks (LSTMs) are able to capture these relationships due to their sequence processing characteristics, but they suffer from low computational efficiency and difficulty in parallel processing. The multi-head attention mechanism solves these problems by extending an attention mechanism to multiple "heads", each of which learns representations of different subspaces in the sequence. In this way, the model can capture information from different representation subspaces at the same time, enhancing the model's ability to capture complex relationships.

[0060] The calculation process of multi-head attention is as follows:

[0061] 1) Linear transformation: input sequence Each element at each position in is first transformed by three linear transformations Get the query Q, key K and value V.

[0062] 2) Scaled dot product attention: Each head calculates the dot product of the query Q and all keys K, and then normalizes it through the softmax function to obtain the attention weight. This weight is weighted and summed with the corresponding value V to obtain the weighted value. In order to reduce overfitting, the dot product result is divided by the square root of the key vector dimension D.

[0063]

[0064] 3) Multi-head mechanism: The above process is carried out in parallel in multiple heads, and each head uses a different linear transformation matrix, so that information can be learned from different representation subspaces.

[0065] 4) Concatenation and linear transformation: The outputs of all heads are concatenated together and then go through a final linear transformation to get the final output.

[0066] B.LoRA module.

[0067] The present invention proposes a multi-task learning efficient parameter network structure, which improves the overall performance and efficiency of the model on multiple tasks by training the model on multiple tasks.

[0068] (1) Figure 3 As shown, in the Q, K, V linear transformation matrices of the attention network, the multi-task shared LoRA modules are introduced respectively:

[0069] Q=W Q x+B Q A Q x

[0070] K=W K x+B k A k x

[0071] V=W V x+B V A V x

[0072] Among them, W Q , W K , W V are the linear transformation matrices of query, key, and value respectively, A Q , B Q Divided into the query low-rank decomposition matrix A and the low-rank decomposition matrix B, A K , B K Divided into a key low-rank decomposition matrix A and a low-rank decomposition matrix B, A V , B V Divided into a low-rank decomposition matrix A of values ​​and a low-rank decomposition matrix B.

[0073] LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning method for adjusting pre-trained large language models. The core idea of ​​this method is to introduce a low-rank structure on the model's parameter matrix to achieve fine-tuning of the model without updating the parameters of the entire model. This enables the model to adapt to new tasks or domains while keeping most of the pre-trained parameters unchanged, thereby reducing computational and storage costs.

[0074] The working principle of LoRA is as follows:

[0075] 1) Determine the model parameters that need to be fine-tuned, usually the model's weight matrix, such as the linear transformation matrix in each layer of the Transformer multi-head attention and F is the input feature dimension size.

[0076] 2) Build a low-rank structure: Build a low-rank structure for these parameters, usually by multiplying two smaller matrices.

[0077] 3) Fine-tuning: During training, only the parameters in this low-rank structure are updated, while other parameters in the original parameter matrix remain unchanged.

[0078] By introducing the multi-task shared LoRA module, the model can learn the common knowledge between different tasks and utilize cross-task shared knowledge.

[0079] C. Multi-task fusion scaling network.

[0080] Figure 2 The rightmost is the multi-task fusion scaling network proposed by the present invention. The network takes the input of the overall model as input and is composed of N layers of linear transformation, nonlinear activation and Dropout. The specific calculation method is as follows:

[0081] Scaling = Softmax(W N *…W2*Dropout(Relu(W1X)))

[0082] Among them, Scaling is the output of the multi-task fusion scaling network, W N is the weight parameter of the Nth linear transformation layer, Scaling = [Scaling FFN ,Scaling V ,Scaling K ], S represents the length of the sequence (such as Figure 2 In Seq Len, T represents the total number of tasks to be learned and integrated ( Figure 2 Task Num in Represents the FFN scaling factor component of the multi-task fusion scaling network output, Represents the attention value scaling factor component of the multi-task fusion scaling network output, Represents the attention key scaling factor component of the multi-task fusion scaling network output.

[0083] Figure 2 The middle part includes "task-specific learnable activation scaling adaptation vectors", which is the combination of multi-task fusion scaling network and large language model.

[0084] 1) in It indicates that the feedforward network corresponding to the kth task can learn the activation scaling adaptation vector, and H is the hidden layer size output by the intermediate hidden layer activation unit of the feedforward neural network.

[0085] 2) in represents the learnable attention Value scaling adaptation vector corresponding to the k-th task, and M is the size of the attention median V.

[0086] 3) in represents the learnable attention key scaling adaptation vector corresponding to the k-th task, and D is the size of the key K in the attention.

[0087] Figure 2 The middle part also includes the "Adaptive Multi-Task Fusion Activation Scaling Adaptation Vector", which is calculated as follows:

[0088] 1)

[0089] 2)

[0090] 3)

[0091] Step 2: Initialize the learnable activation scaling adaptation vector l for each task k k .

[0092] The present invention obtains the initialized learnable activation scaling adaptation vector l by constructing a training data set for each task k k Specifically, the present invention targets multiple tasks respectively, uses the training data of each task, and separately performs Figure 4 The model shown is fine-tuned, freezing the large language base model parameters and scaling the adaptation vectors only for the learnable activations for each specific task. Train and update separately.

[0093] Step 3: Construct a multi-task joint fine-tuning training dataset.

[0094] The multi-task joint fine-tuning training data set constructed by the present invention must not only ensure the diversity of the heavy data of each task, but also ensure the balance of the samples of each task.

[0095] Step 4: Use the multi-task joint fine-tuning training dataset to adjust the LoRA module parameters, multi-task fusion scaling network parameters and learnable activation scaling adaptation vector l k to generate a trained model.

[0096] The present invention scales the adaptive vector of the learnable activation of each task after initialization Load to Figure 2 , Figure 3 In the model proposed by the present invention shown, the training data of multiple tasks are used to jointly optimize the model parameters, and the model parameters to be optimized include: a multi-task fusion scaling network (Scaling Network), parameters of a shared LoRA module, and task-specific learnable activation scaling adaptation vectors.

[0097] Among them, Figure 3 , Figure 4 As shown, the "adaptive multi-task fusion activation scaling adaptation vector" output by the "adaptive activation scaling adapter" proposed in the present invention scales and adapts the activation output, attention Key and attention Value of the middle hidden layer of the feedforward neural network respectively. The calculation formula is as follows:

[0098] P′=L FFN ⊙P

[0099] V′=L V ⊙V

[0100] K′=L K ⊙K

[0101] in, is the output of the middle hidden layer of the feedforward neural network, H is the output dimension of the hidden layer, are the attention key and attention value respectively, and ⊙ represents bitwise multiplication.

[0102] Finally, the degree of convergence of the model training is judged by the change in loss during training and the evaluation results of the validation set. The end of training is controlled according to the model convergence and the set maximum training step size and other parameters.

[0103] In summary, the present invention proposes a new "adaptive activation scaling adapter model structure". The adaptive activation scaling adapter model structure integrates the LoRA method and the "multi-task adaptive activation scaling adapter" to design and propose a large-model multi-task learning model architecture. By sharing LoRA, the model can learn the common knowledge between different tasks and utilize cross-task shared knowledge. Through the multi-task adaptive activation scaling adapter, the characteristics of different tasks can be fully learned, and the seesaw problem between different tasks can be alleviated, so as to achieve a common improvement in the performance of a model on multiple tasks.

[0104] The present invention also proposes a large language model multi-task joint fine-tuning method based on adaptive activation scaling adaptation, which achieves efficient multi-task learning and optimization using limited resources through optimization of extremely low parameter amounts.

[0105] The above is one of the embodiments of the present invention, and the present invention should not be limited to the contents disclosed in the embodiment and the drawings. Any equivalent or modification completed without departing from the spirit disclosed in the present invention shall fall within the scope of protection of the present invention.

Claims

1. A multi-task learning method for a large language model based on adaptive activation scaling, characterized in that: The method is applied to natural language processing tasks, including: Constructing a model, the model comprising: a large language base model, a LoRA module and a multi-task fusion scaling network, the large language base model comprising an attention module and a feedforward neural network, the LoRA module being introduced into a query linear transformation matrix, a key linear transformation matrix and a value linear transformation matrix of the attention module, the multi-task fusion scaling network being used to calculate a scaling factor according to an input of the model, and the structure of the multi-task fusion scaling network comprising: N layers of linear transformation, nonlinear activation and Dropout; Initialize the learnable activation scaling adaptation vector l for each task k k ; Wherein, the learnable activation scaling adaptation vector l k Includes: Learnable activation scaling adaptation vectors for feed-forward networks Learnable activation scaling adaptation vector for attention values and learnable activation scaling adaptation vectors for attention keys Construct a multi-task joint fine-tuning training dataset; The multi-task joint fine-tuning training data set is used to adjust the LoRA module parameters, the multi-task fusion scaling network parameters and the learnable activation scaling adaptation vector l k to generate a trained model; wherein, during the optimization process, based on the scaling factor and the learnable activation scaling adaptation vector l k The intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module are scaled and adapted.

2. The method according to claim 1, characterized in that: The learnable activation scaling adaptation vector l of each task k is initialized k ,include: Construct a training data set train for each task k k ; Freeze the parameters of the large language base model; On top of the large language base model with frozen parameters, the learnable activations for each task k are scaled to an adaptation vector l k Train and update respectively to obtain the learnable activation scaling adaptation vector l k The initial value of .

3. The method according to claim 1, characterized in that The scaling factor includes: Scaling factor component Scaling of the feedforward neural network FFN , Scaling factor component of attention value V Scaling factor component of the attention key K .

4. The method according to claim 3, characterized in that Based on the scaling factor and the learnable activation scaling adaptation vector l k Scaling and adapting the intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module, including: Scale the learnable activations of each task k to the adaptation vector Learnable activation scaling adaptation vector and learnable activation scaling adaptation vector Connect them separately to obtain the learnable activation scaling adaptation vector l FFN , the learnable activation scaling adaptation vector l V and the learnable activation scaling adaptation vector l K ; According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l FFN , generate the fused activation scaled adaptation vector L FFN ; According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l V , generate the fused activation scaled adaptation vector L V ; According to the scaling factor component Scaling FFN and the learnable activation scaling adaptation vector l K , generate the fused activation scaled adaptation vector L K ; Perform the fusion activation scaling adaptation vector L FFN Bitwise multiplication with the intermediate hidden layer output P of the feedforward neural network to obtain the scaled intermediate hidden layer output P′; Perform the fusion activation scaling adaptation vector L V Bitwise multiplication with the attention value V in the attention module to obtain a scaled attention value V′; Perform the fusion activation scaling adaptation vector L K The bitwise multiplication with the attention key K in the attention module obtains the scaled attention key K′.

5. A large language model multi-task learning system based on adaptive activation scaling adaptation, characterized in that: The system is applied to natural language processing tasks, including: A model construction module is used to construct a model, the model comprising: a large language base model, a LoRA module and a multi-task fusion scaling network, the large language base model comprises an attention module and a feedforward neural network, the LoRA module is introduced into the query linear transformation matrix, the key linear transformation matrix and the value linear transformation matrix of the attention module, the multi-task fusion scaling network is used to calculate a scaling factor according to the input of the model, and the structure of the multi-task fusion scaling network comprises: N layers of linear transformation, nonlinear activation and Dropout; Vector initialization module, used to initialize the learnable activation scaling adaptation vector l for each task k k ; Wherein, the learnable activation scaling adaptation vector l k Includes: Learnable activation scaling adaptation vectors for feed-forward networks Learnable activation scaling adaptation vector for attention values and learnable activation scaling adaptation vectors for attention keys Dataset construction module, used to construct multi-task joint fine-tuning training dataset; A model training module is used to use the multi-task joint fine-tuning training data set to adjust the LoRA module parameters, the multi-task fusion scaling network parameters and the learnable activation scaling adaptation vector l k to generate a trained model; wherein, during the optimization process, based on the scaling factor and the learnable activation scaling adaptation vector l k The intermediate hidden layer activation output of the feedforward neural network, the attention value and the attention key in the attention module are scaled and adapted.

6. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements a large language model multi-task learning method based on adaptive activation scaling adaptation as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the large language model multi-task learning method based on adaptive activation scaling adaptation as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Multi-task large model fine tuning method based on adapters and low-rank adaptation

    CN116822611A

  • Systems and methods for finetuning with learned hidden representations of parameter changes

    US20240020486A1