A Parameter-Efficient Fine-Tuning Method, System, and Application Based on Interleaved Memory of Twin Large Language Models
Through the interwoven memory method of the twin language model, the catastrophic forgetting problem in the fine-tuning process is solved, the performance of downstream tasks is improved, and the effect of efficient fine-tuning is achieved.
Patent Information
- Application Number
- CN202410829181.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Existing large language models are prone to catastrophic forgetting problems during fine-tuning, and existing high-efficiency fine-tuning methods for parameter tuning are difficult to improve downstream tasks while maintaining the original model knowledge.
The twin large language model interleaving memory method is adopted, and two LLMs sharing the same structure and pretrained parameters are constructed, one of which remains frozen, the other is efficiently fine-tuned, and the contribution of the next word element is generated by combining the two memories through a simple or gating-based interleaving mechanism.
It effectively alleviates the catastrophic forgetting problem, improves the performance of downstream tasks, and at the same time and space complexity are comparable to existing methods, maintaining the stability and adaptability of the model.
Smart Images

Figure CN119089940B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer science and technology, and relates to a parameter-efficient fine-tuning method, system and application based on the interleaved memory of twin large language models. Background Art
[0002] Parameter-Efficient Fine-Tuning (PEFT) refers to a technique that reduces the amount of parameter updates when fine-tuning a Large Language Model (LLM) to reduce computational costs and resource consumption. This method can be divided into four major categories [1] : additive, selective, reparameterization, and hybrid methods. Additive PEFT, such as prefix-tuning [2] and (IA)3 [3] introduces additional modules into the large language model to adapt to specific downstream tasks. Selective PEFT, such as BitFit [4] focuses on selecting a small subset of parameters from the base model for fine-tuning. Reparameterization PEFT [5] methods aim to address the problem that additive methods increase the model depth, resulting in inference latency. Hybrid PEFT [6] aims to combine different PEFT methods to provide a comprehensive solution. PEFT provides a promising alternative to Full Fine-Tuning (FFT), enabling comparable or even better performance while requiring less time and resources. Current PEFT methods allow the alignment of LLMs with new domains at the parameter level. However, modifying or adding parameters may disrupt the existing world knowledge and reasoning capabilities of the LLMs. This not only affects the performance of LLMs on other tasks but also leads to performance crises when processing inputs in new domains.
[0003] LLMs face the challenge of catastrophic forgetting during both pre-training and fine-tuning, i.e., the model is prone to forgetting the previously acquired knowledge after fine-tuning. Empirical studies have shown that even efficient fine-tuning techniques like LoRA are affected by catastrophic forgetting. One scenario explored to address the catastrophic forgetting problem is the continual learning scenario. To address this challenge, several effective strategies have been proposed, including methods such as constrained parameter updates, data replay, and parameter isolation. However, for LLMs, obtaining extensive pre-training data is an expensive task. Additionally, ensuring that the existing knowledge learned by the LLMs is not overwritten at the parameter level is also challenging. Summary of the Invention
[0004] To address the deficiencies of the existing technologies, the objective of the present invention is to provide a parameter-efficient fine-tuning method, system, and application based on the interweaving memories of siamese large language models, which can enhance the effect of parameter-efficient fine-tuning and alleviate the catastrophic forgetting caused by fine-tuning.
[0005] The present invention discloses a simple and effective parameter-efficient fine-tuning method (Parameter-Efficient Fine-Tuning, PEFT), namely, the Interweaving Memories of Siamese Large Language Model (IMSM). This is a model-agnostic framework that can theoretically be applied to all open-source large language models (Large Language Model, LLM) and existing PEFT methods. Specifically, the siamese LLM can be regarded as two LLMs sharing the same structure and pre-trained parameters, where one remains frozen and the other undergoes PEFT. Given an input, the siamese large language models generate different memories, i.e., the last-layer hidden states, and then IMSM introduces an interweaving mechanism to regulate the contributions of the memories of the frozen LLM and the fine-tuned LLM to generate the next token. Two memory interweaving mechanisms are proposed in this paper: a simple interweaving mechanism and a gate-based interweaving mechanism. Extensive experiments are conducted to evaluate the performance of IMSM compared with state-of-the-art PEFT methods such as LoRA, (IA) 3 、AdaLoRA. Three popular LLMs are used as the backbone large models, including ChatGLM3, Qwen1.5, and Llama3, and fine-tuning is performed on three datasets. The results show that IMSM not only achieves excellent performance on downstream tasks but also effectively alleviates the catastrophic forgetting problem without increasing additional fine-tuning time.
[0006] The present invention provides a parameter-efficient fine-tuning method based on the interweaving memories of siamese large language models, and the method comprises the following steps:
[0007] Step 1: Pre-construct two large language models sharing the same structure and pre-trained parameters; one of the large language models keeps its original parameters frozen, and the other large language model is fine-tuned using the PEFT method;
[0008] Step 2: Use the two large language models constructed in Step 1 to process the given input respectively, and generate different memories; the memory refers to the last-layer hidden state generated by the large language model;
[0009] Step 3: Introduce the interweaving memory mechanism of the siamese large language model to regulate the contributions of the memories of the large language model with frozen original parameters and the large language model fine-tuned using the PEFT method to generate the next token.
[0010] In step one, the parameters are updated by adding an adapter to the original parameters to obtain a fine-tuned large language model;
[0011] The update of the parameters is represented by the following formula:
[0012] W = W o + ΔW = W o + BA,
[0013] where W o is the original parameter and remains frozen, while and represent trainable low-rank matrices, and r' is much smaller than d and k; ΔW represents the adapter, and W represents the parameters of the fine-tuned large language model.
[0014] In step three, the twin large language model interleaved memory mechanism includes a simple memory interleaving mechanism and a gated-based memory interleaving mechanism;
[0015] The simple memory interleaving mechanism directly adds the logical values of the frozen large language model and the fine-tuned large language model to calculate and generate the probability distribution of the next token;
[0016] For each time step t, the logical values output by the frozen large language model and the fine-tuned large language model are respectively accessed. The logical value refers to the unnormalized score of each token in the vocabulary set V as the next token, which is used to reflect the confidence level of the model in predicting the next token;
[0017] Based on the input query x and response y <t , calculate the final logical value:
[0018]
[0019] where x and y belong to the vocabulary set V, and respectively represent and the generated logical values, where represents the frozen large language model with the original parameter W o , while represents the fine-tuned large language model after parameter-efficient fine-tuning with the adapter ΔW added.
[0020] The gated-based memory interleaving mechanism concatenates the last hidden states of the frozen large language model and the fine-tuned large language model, and then uses a low-rank linear layer to generate a gate value to dynamically balance the contributions of the two models to the next generation, thereby optimizing the performance of the model;
[0021] At each time step, access and For the last hidden state, concatenate the hidden states along the feature dimension. Construct a low-rank linear layer at the top of the model to generate a gate that determines the contribution of each large language model (LLM) at the next step in the logical value generation process, which is expressed by the formula as follows:
[0022]
[0023] where ⊕ represents vector concatenation, f represents the low-rank linear layer, and represent the last hidden states of the frozen large language model and the fine-tuned large language model respectively. where d represents the dimension of the hidden layer state. In addition, W A and W B are the weight matrices of the linear layer respectively; where r is a hyperparameter. It should be noted that r is much smaller than d to ensure a limited number of parameters. In addition, the gate is obtained through sigmoid activation.
[0024] The updated last hidden state and the logical value logits can be calculated as follows:
[0025]
[0026]
[0027] where ° represents element-wise multiplication, and W out represents the parameter of the twin large language model linear layer for predicting the next token. In step three, the probability distribution of the next token is expressed as follows:
[0028] p(y t |x,y <t ) = softmax(logits).
[0029] In the present invention, the parameter adapter △W in the fine-tuned large language model, the weight matrices W A and W B in the gated twin large language model memory interleaving mechanism are optimized by the cross-entropy loss function:
[0030]
[0031] where T is the length of the output sequence.
[0032] The present invention also provides an adjustment system for implementing the above parameter-efficient fine-tuning method, and the adjustment system includes: a twin large language model generation module, a memory interleaving module, and a fine-tuning module;
[0033] The twin large language model generation module is used to generate two twin large language models that share the same structure and pre-trained parameters;
[0034] The memory interleaving module is used to adjust the contributions of the memories of the frozen model and the fine-tuned model to generate the next token;
[0035] The fine-tuning module is used to perform parameter-efficient fine-tuning on one of the twin large language models.
[0036] The present invention also provides the application of the above parameter-efficient fine-tuning method or adjustment system in field specialization, machine translation, text summarization, etc.
[0037] The time and space complexity of the present invention: Compared with only using the ordinary PEFT method, IMSM performs equivalently in terms of space and time complexity. Theoretically, although two different memories need to be output, the twin LLM does not require double the storage resources. Although the gated-based memory interleaving mechanism introduces some additional trainable parameters, compared with the original PEFT parameters, these effects are negligible. For example, in ChatGLM with a hidden layer dimension of 4096, when the ranks of LoRA and the gate are both 16, the trainable parameters of the gate only account for 5.04% of LoRA. In addition, the two memories can be generated in parallel. During the fine-tuning and inference processes, the time required to interleave the two memories at each step remains constant. This ensures that IMSM has the same time complexity as the existing PEFT methods.
[0038] The beneficial effects of the present invention include: The method of the present invention generates different memories (i.e., the last hidden state) by keeping one of the two twin large language models that share the same structure and pre-trained parameters frozen and fine-tuning the other. Then, through a simple memory interleaving mechanism or a gated-based memory interleaving mechanism, the memories of the two models are combined to dynamically balance their contributions to generating the next token, thereby improving the performance in downstream tasks while maintaining the original knowledge of the model. In the present invention, IMSM with two interleaving mechanisms is superior to the ordinary fine-tuning method to varying degrees, with average improvements of 1.53% and 2.21% respectively. In a specific embodiment, the two memory interleaving mechanisms of the present invention both show better performance for catastrophic forgetting on non-target datasets. Among them, the simple IMSM reduces LoRA forgetting by 1.90, while the gated-based IMSM reduces it by 3.70. For AdaLoRA, both mechanisms reduce it by 1.86. For (IA) 3, simple IMSM reduces forgetting by 1.11, and gated IMSM reduces forgetting by 1.29. In different tasks, the IMSM method achieves a good balance between stability and adaptability by adjusting the rank of the gate, ensuring optimal performance and minimal knowledge forgetting. At the same time, the method of the present invention shows high adaptability in various tasks and PEFT frameworks and can flexibly adjust parameters to achieve the best results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0040] Figure 1 It is a comparison schematic diagram of the efficient fine-tuning of ordinary parameters of the present invention and the interleaved memory of a simple twin large language model.
[0041] Figure 2 It is a schematic diagram of the memory interleaving mechanism of the gated twin large language model of the present invention.
[0042] Figure 3 It is a schematic diagram of the influence of the rank of the gate of the present invention on three data sets.
[0043] Figure 4 It is a schematic diagram of the result of summing logical values of three data sets in different proportions by the present invention.
[0044] Figure 5 It is a schematic diagram of the knowledge forgetting situation of different ranks of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no special restrictions.
[0046] The present invention proposes a model-agnostic framework called Interweaving Memories of Siamese Large Language Model (IMSM), and develops a parameter-efficient fine-tuning method based on this framework. The Siamese LLM can be regarded as two LLMs with the same size, structure, and pre-trained parameters, such as Llama3. One of them keeps the original parameters frozen, while the other will be fine-tuned using the PEFT method. The frozen LLM and the fine-tuned LLM process the same input token simultaneously, thus generating different memories, i.e., the last hidden state. Finally, two memory interweaving mechanisms are proposed: a simple interweaving mechanism or a gate-based interweaving mechanism. The next token of the response output comes from the interwoven memories. In this framework, the parameters related to PEFT and the parameters related to the gate mechanism will be fine-tuned using the downstream task data, while the remaining parameters will be frozen. Specifically, the PEFT-related parameters include a small part of the parameters adjusted during the fine-tuning process, such as the low-rank matrix introduced in a specific layer; the parameters related to the gate mechanism include the linear layer parameters used in the gating mechanism. The objective of the present invention is to maintain the model's inherent world knowledge and general reasoning ability while improving its performance in downstream tasks. The method of using Interweaving Memories of Siamese Large Language Model (IMSM) proposed by the present invention can be applied to various open-source LLMs and PEFT methods.
[0047] 1. Task Definition:
[0048] The main objective of IMSM proposed by the present invention is to control the output generated by the LLM. Mathematically, a language model aims to model the probability distribution of natural language:
[0049]
[0050] where P(x) represents the joint probability of generating the sequence x, x i represents the i-th token in the sequence, and W represents the parameters of the language model;
[0051] To determine the next token in the generation process, various strategies can be adopted. A popular approach is greedy decoding, i.e., the model selects the token with the highest probability:
[0052]
[0053] Currently, most LLMs are autoregressive models that mainly rely on the decoder framework. During the generation process, the LLM considers the input and the previously generated tokens. The IMSM proposed in this invention can interleave two different memories before each generation step. Then, based on this updated memory, the probability of the next token is predicted.
[0054] 2. Twin Large Language Models
[0055] Twin LLMs can be regarded as two LLMs that share the same Transformer structure and pre-trained parameters. When using the PEFT method, one of them remains frozen to retain the original knowledge, while the other is fine-tuned with downstream data to align with the new target. Taking LoRA as an example, the parameter W is updated by adding an adapter △W to the original parameter :
[0056] W = W o + △W = W o + BA
[0057] where, W o remains frozen, while and represent trainable low-rank matrices. And r' is much smaller than d and k. Such a configuration significantly reduces the number of trainable parameters. W o and W represent the parameters of the frozen LLM and the fine-tuned LLM respectively.
[0058] 3. Memory Interleaving Mechanism
[0059] During fine-tuning and inference, the twin LLMs proposed in this invention generate different hidden states and corresponding logical values based on the input and the tokens that have been generated. The last-layer hidden state can be used as a feature representation, encapsulating the LLM's memory of the previous input. In addition, the logical value represents the model's confidence level in predicting the next token. This invention will use two methods, the simple memory interleaving mechanism and the gate-based memory interleaving mechanism, to interleave different memories.
[0060] Let represent the frozen LLM with parameter W o , while represents the fine-tuned LLM with partial fine-tuning parameters △W. Two memory interleaving mechanisms are proposed: the simple memory interleaving mechanism and the gate-based interleaving mechanism.
[0061] Mode 1: Simple Memory Interleaving Mechanism
[0062] At each time step t, and The logical value of the output, where the logical value refers to the unnormalized score of each token in the vocabulary set V as the next token. Then, based on the input query x and the response y <t , the final logical value can be calculated:
[0063]
[0064] where x and y belong to the vocabulary set V, and respectively represent and the generated logical values.
[0065] The vocabulary set V is the set of all possible tokens in the language model. In the simple memory interleaving mechanism, the frozen large language model and the fine-tuned large language model calculate the logical value of each token in the vocabulary set based on the input and the previously generated tokens respectively. By adding these logical values together, the model can predict and generate the next token.
[0066] Note that for the prediction of the next token, directly adding the two hidden states has the same effect as adding the two logical values. Therefore, for clarity and intuitiveness, this strategy is explicitly described as the addition of two logical values.
[0067] Mode 2: Gated Memory Interleaving Mechanism
[0068] Considering that the logical value is calculated from the last layer hidden state, and thus the last layer hidden state can capture more potential information, the present invention also provides an alternative memory interleaving mechanism for the twin LLM, namely the gated memory interleaving mechanism. At each time step, the last layer hidden states of and can be accessed and these hidden states are concatenated along the feature dimension. A low-rank linear layer is constructed on top of the model to generate a gate that determines the contribution of each LLM in the next step of the logical value generation process:
[0069]
[0070] where ⊕ represents vector concatenation, f represents the low-rank linear layer, and respectively represent the last layer hidden states of the frozen LLM and the fine-tuned LLM. where d represents the dimension of the hidden layer state. Additionally, where r is a hyperparameter. It should be noted that r is much smaller than d to ensure a limited number of parameters. Additionally, the gate is obtained through sigmoid activation.
[0071] The updated last hidden state The sum logits can be calculated as follows:
[0072]
[0073]
[0074] where ○ represents element-wise multiplication, and W out represents the parameters of the twin LLM linear layer for predicting the next token.
[0075] 4. Model Training and Inference
[0076] Finally, the probability distribution of the next token can be calculated as follows:
[0077] p(y t |x, y <t ) = softmax(logits)
[0078] During the fine-tuning stage, the cross-entropy loss function is used to optimize the parameters ΔW, W A and W B :
[0079]
[0080] where T is the length of the output sequence.
[0081] Figure 2 is a schematic diagram of the gated twin large language model memory interleaving mechanism in the present invention. In the figure,
[0082] 1. Input:
[0083] The input text sequence is fed into the frozen large model and the fine-tuned large model simultaneously.
[0084] 2. Twin Large Language Model:
[0085] The figure contains two models: the frozen large model and the fine-tuned large model. Both have the same structure and pre-trained parameters, but the fine-tuned large model has additionally undergone parameter-efficient fine-tuning.
[0086] 3. Last Layer Hidden State:
[0087] The frozen large model and the fine-tuned large model respectively generate their last layer hidden states. These hidden states contain the model's understanding and memory of the input sequence.
[0088] 4. Concatenated Hidden State:
[0089] The last layer hidden states of the frozen large model and the fine-tuned large model are concatenated along the feature dimension to form the concatenated hidden state.
[0090] 5. Linear Layer:
[0091] The concatenated hidden state passes through two linear layers A and B to generate the input for the gating mechanism. Both linear layers A and B are fine-tunable parameters.
[0092] 6. Gating mechanism:
[0093] The gating mechanism generates a gate value through the sigmoid function, which determines the contribution ratio of the frozen model and the fine-tuned model to the final output. The specific formula is as follows:
[0094]
[0095] where σ represents the sigmoid function, ⊕ represents the concatenation operation, and W A and W B are the weight matrices of the linear layers respectively;
[0096] 7. Updated hidden state:
[0097] According to the generated gate value, the updated hidden state is calculated, which combines the memories of both:
[0098]
[0099] where ° represents element-wise multiplication, and W out represents the parameter of the twin large language model linear layer for predicting the next token.
[0100] 8. Output generation:
[0101] Finally, the updated hidden state is used to generate the probability distribution of the next token, and the final output is generated according to the probability distribution.
[0102] In the inference stage, as shown in Algorithm 1, the present invention uses a greedy search strategy to generate the output sequence. The token with the highest probability is selected as the next generated token, which is added to the input sequence, and then this process is repeated until the complete sequence is generated.
[0103]
[0104]
[0105] Effective effects compared with the prior art
[0106] 1. Dataset
[0107] Experiments were conducted on three benchmark datasets to evaluate the alignment ability of the proposed IMSM in downstream tasks and its general reasoning ability in other tasks.
[0108] · The MRPC dataset aims to determine whether two given sentences have the same meaning.
[0109] · CoLA is a dataset on language annotation, aiming to evaluate the model's ability to recognize the grammatical acceptability in sentences.
[0110] · ROPES is a reading comprehension dataset, aiming to evaluate the model's reasoning ability for multi-step reasoning in new situations.
[0111] 2. Evaluation Metrics
[0112] Different metrics are used for different downstream tasks. For MRPC, accuracy (Accuracy, Acc.) and F1 are usually used as metrics. For CoLA, the Matthews Correlation Coefficient (MCC) is used as the main metric. For ROPES, Exact Match (EM) and the F1 token overlap between the answer text and the gold are used as metrics.
[0113] 3. Baseline PEFT Methods
[0114] The following PEFT methods are used as baselines and the proposed IMSM method is applied.
[0115] LoRA: It freezes the pre-trained model weights and incorporates detachable plug-ins (trainable rank decomposition matrices) into the layers of the Transformer architecture.
[0116] (IA) 3 : This method uses the learned vectors to readjust the internal activations and has fewer trainable parameters than LoRA.
[0117] AdaLoRA: It adaptively adjusts the ranks of different modules according to the importance of the weight matrix.
[0118] 4. Performance Comparison
[0119]
[0120]
[0121] The above table shows the results of IMSM and the baseline models on three datasets. Generally speaking, the fine-tuned LLMs perform better than the LLMs enhanced only by prompting. This highlights the importance of fine-tuning in aligning LLMs to unfamiliar tasks compared to in-context learning. The IMSM with two interleaving mechanisms outperforms the ordinary fine-tuning methods to varying degrees, with average improvements of 1.53% and 2.21% respectively. In most cases, the interleaving method based on the gating mechanism has better performance than the simple interleaving method.
[0122] The effects of fine-tuning are influenced by multiple factors, including the complexity of natural language tasks, the reasoning ability of the original LLM, and the adaptability of the fine-tuning methods employed. Specifically, the improvements observed on the MRPC dataset were relatively small. MRPC focuses on determining semantic similarity between two sentences and relies more on general language capabilities rather than highly specialized downstream patterns. The effectiveness of fine-tuning is closely related to the quality and capabilities of the underlying LLM. As one of the state-of-the-art open-source LLMs, the model based on Llama3 almost achieved the best performance. Additionally, LoRA includes low-rank factorization, highlighting its effectiveness in aligning with downstream tasks.
[0123] However, IMSM, as a simple yet powerful model-agnostic framework, excels at combining the stability of the original LLM with the adaptability of the fine-tuned LLM. Moreover, IMSM also introduces a gating mechanism that allows for selective control over whether the understanding comes from the original knowledge or the adjusted knowledge. Therefore, it has achieved continuous improvements in various downstream tasks.
[0124] 5. Catastrophic Forgetting Verification
[0125]
[0126] Fine-tuning inevitably introduces a certain degree of forgetting. The present invention fine-tunes LLMs on one dataset and then evaluates their performance on other datasets. In particular, in a specific implementation of the present invention, ChatGLM3-6B is used as the backbone model, and then it is fine-tuned using existing PEFT methods and IMSM on three datasets. For each target task, two datasets for evaluating forgetting are introduced. The first dataset shared by all tasks is an out-of-context question-answering dataset, WebQ, which is used to evaluate the degree of forgetting of world knowledge. The present invention evaluates the performance on WebQ by checking whether the generated responses contain the gold answers, rather than strictly requiring exact matches. The second dataset varies depending on the target task. For simplicity, CoLA, MRPC, and MultiRC are used respectively.
[0127] As can be seen from the above table, the three common PEFT methods all prioritize improving adaptability while neglecting stability. For example, after fine-tuning MRPC using LoRA, (IA) 3 and AdaLoRA, the performance of CoLA decreased from 43.99 to 29.54, 28.46, and 35.86 respectively. Similarly, the performance on WebQ also decreased from 25.49 to 24.02, 22.19, and 23.52. Overall, LoRA exhibited the highest degree of catastrophic forgetting, followed by AdaLoRA and (IA) 3This can be attributed to the number of parameters in the fine-tuning process. The fewer parameters involved, the less impact on the general capabilities of the model.
[0128] The IMSM method in the present invention can effectively mitigate the performance degradation of other datasets caused by fine-tuning. On average, compared with ordinary PEFT, the two memory interleaving mechanisms of the present invention both exhibit better performance on non-target datasets. Among them, the simple IMSM reduces LoRA forgetting by 1.90, while the gate-based IMSM reduces it by 3.70. For AdaLoRA, both mechanisms reduce it by 1.86. For (IA) 3 , the simple IMSM reduces forgetting by 1.11, and the gate-based IMSM reduces forgetting by 1.29.
[0129] By summing the logical values from the frozen and fine-tuned LLMs, the degree to which the probability of the initially correct tokens on the non-target dataset decreases due to fine-tuning can be mitigated. The last hidden state serves as the memory for the LLM to understand the input sequence. By considering the memories from the original and fine-tuned LLMs, it is ensured that the twin LLMs maintain a balance between their prior knowledge and adaptation to new data. In addition, it shows better performance in alleviating catastrophic forgetting compared with the gate-based interleaving mechanism. By dynamically balancing the knowledge of the frozen model and the fine-tuned model, the gate mechanism can effectively recall previous memories, helping the twin LLMs prevent overfitting to new data while retaining the original knowledge of the model.
[0130] 6. Hyperparameter Experiments
[0131] Figure 3 Shows the impact of the rank of the gate on three datasets. The present invention evaluates the performance of MRPC and CoLA using ranks of 2, 4, 8, and 16 for the gate, and evaluates the performance of ROPES using ranks of 4, 8, 16, and 32 for the gate. First, all selected ranks exceed or match the corresponding ordinary PEFT performance. The best rank varies depending on different PEFT methods and tasks. This emphasizes the effectiveness and adaptability of IMSM in various PEFT frameworks and tasks. For MRPC, the overall performance improves as the rank increases. Specifically, the peak performance of LoRA, (IA)3, and AdaLoRA reaches at ranks 8, 16, and 16 respectively. For CoLA, the performance reaches the peak when the rank is 8. For ROPES, it is obvious that higher ranks actually tend to reduce the performance. This indicates that lower ranks may be more beneficial for improving the performance of these tasks.
[0132] Figure 4Shows the results of summing logical values for three datasets in different proportions. The results indicate that satisfactory performance is obtained when the logical values of the frozen knowledge and fine-tuned knowledge in the twin LLMs are added in equal proportions. For example, when fine-tuning the LLM on the CoLA dataset using the LoRA-based IMSM method, the performance reaches 65.53 when the logical values of the frozen model and the fine-tuned model are combined in a 1:1 ratio. However, when the ratio is changed to 2:3 or 3:2, the performance drops significantly, to 63.81 and 63.70 respectively. This result highlights the potential benefits of leveraging the pre-trained knowledge stored in the frozen LLM and the new knowledge obtained from the fine-tuned LLM. By combining these memories in a balanced manner, IMSM exhibits enhanced performance.
[0133] In addition, the present invention also tested the knowledge forgetting situation of different ranks, such as Figure 5 shown, from left to right represent IMSM based on LoRA, (IA) 3 and AdaLoRA. The present invention uses MRPC as the target dataset and reports the average performance degradation of the other two datasets. When the rank of the gate is 8, the forgetting degree of IMSM based on LoRA and (IA) 3 is the lowest, while when the rank is 16, the forgetting degree of IMSM based on AdaLoRA is the lowest. In addition, considering the experimental results of hyperparameters on the performance of downstream tasks, the method of the present invention achieves a trade-off in terms of stability and adaptability, and even the best performance. In addition, when the rank of the gate is relatively low, such as 2, the forgetting degree will increase. Gates with relatively low ranks cannot capture the contributions of the original knowledge and the new knowledge.
[0134] Example 1
[0135] Such as Figure 1As shown, the user input and the tokens already generated by the model are "Please answer the following question. Question: What family of animals does a sheep belong to? Answer: A sheep belongs to". For the model after parameter-efficient fine-tuning with ordinary parameters, after the model receives the input, it generates corresponding logical values for all the tokens in the vocabulary (including "sheep", "dog", "cow", etc.). These logical values represent the unnormalized scores for each token as the next token. Through the softmax function, they are transformed into a probability distribution, and the model makes a prediction based on the token with the highest probability. The model may get an incorrect predicted token, such as "sheep". For the twin large language model, the input tokens are simultaneously passed into a fine-tuned large model and a frozen large model, and they each generate a set of logical values. Then, the logical values of the fine-tuned large model and the frozen large model are interleaved. As shown in the figure, a simple addition is performed to combine them into a new set of logical values. The combined logical values are transformed into a probability distribution through the softmax function, and the model makes a prediction based on the new probability distribution, thereby generating the correct predicted token "cow". The above steps are iterated repeatedly until the final end token is generated. Through the above method steps, the present invention can effectively improve the prediction accuracy of the model and achieve more accurate token prediction.
[0136] The present invention includes a twin large language model, which can be regarded as two LLMs sharing the same structure and pre-trained parameters. The parameters of one model are kept frozen, while the other model uses the existing parameter-efficient fine-tuning technology to fine-tune a small part of the parameters. The same output passes through the twin LLMs, resulting in two different logical values. The simplest interleaving method is to fuse the logical values of the two models in a 1:1 ratio and then use the softmax function to predict the next token. The model output and the original input are concatenated together, and the above process is iterated repeatedly until the entire sequence is generated. This method ensures the stability and original performance of the model by keeping the parameters of one model frozen, while providing flexibility and adaptability by fine-tuning a small part of the parameters of the other model. The design of this twin large model helps to improve its performance on specific tasks while maintaining the original capabilities of the model.
[0137] References:
[0138] [1] Han, Z., Gao, C., Liu, J., Zhang, S. Q., et al. (2024). Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608.
[0139] [2]Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190.
[0140] [3]Bansal, M., & Raffel, C. A. (2022). Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems, 35, 1950 - 1965.
[0141] [4]Ben Zaken, E., Goldberg, Y., & Ravfogel, S. (2022). BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 1–9).
[0142] [5]Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
[0143] [6]Mao, Y., Mathias, L., Hou, R., Almahairi, A., Ma, H., Han, J., Yih, S., & Khabsa, M. (2022). UniPELT: A unified framework for parameter-efficient language model tuning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 6253-6264).
[0144] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the scope of protection is defined by the appended claims.
Claims
1. Application of a parameter-efficient fine-tuning method in domain specialization, machine translation, and text summarization, characterized in that, It includes the following steps: Step 1: Pre-construct two large language models that share the same structure and pre-trained parameters; one of the large language models keeps the original parameters frozen, and the other large language model is fine-tuned using the parameter-efficient fine-tuning method; Step 2: Use the two large language models constructed in Step 1 to process the given text sequence respectively, and generate different memories; the memory refers to the last hidden state generated by the large language model; Step 3: Introduce the twin large language model interleaved memory mechanism to adjust the contributions of the memories of the large language model with frozen original parameters and the large language model fine-tuned using the parameter-efficient fine-tuning method to generate the next token; In Step 3, the twin large language model interleaved memory mechanism includes a simple memory interleaving mechanism and a gated-based memory interleaving mechanism; The simple memory interleaving mechanism directly adds the logits of the frozen large language model and the fine-tuned large language model to calculate the probability distribution of generating the next token; The gated-based memory interleaving mechanism splices the last hidden states of the frozen large language model and the fine-tuned large language model, and then uses a low-rank linear layer to generate a gate value to dynamically balance the contributions of the two models to the next generation, thereby optimizing the performance of the model.
2. The application according to claim 1, characterized in that In Step 1, the fine-tuned large language model is obtained by updating the parameters by adding an adapter to the original parameters; The update of the parameters is represented by the following formula: W = W o + ΔW = W o + BA, Among them, W o is the original parameter and remains frozen, while and represent trainable low-rank matrices, and r' is less than d and k; ΔW represents the adapter, and W represents the parameters of the fine-tuned large language model.
3. The application according to claim 1, characterized in that In the simple memory interleaving mechanism, for each time step t, the logits output by the frozen large language model and the fine-tuned large language model are accessed respectively. The logits refer to the unnormalized scores of each token in the vocabulary set V as the next token, which are used to reflect the confidence level of the model in predicting the next token; Based on the input query x and the response y <t , calculate the final logical value: where x and y belong to the vocabulary set V, and respectively represent and the generated logical values, where represents the frozen large language model with the original parameter W o while represents the large language model after parameter-efficient fine-tuning with the adapter ΔW added.
4. The application according to claim 3, characterized in that, In Step 3, the probability distribution of the next token is represented as follows: p(y t | x, y <t ) = softmax(logits).
5. The application according to claim 3, wherein, Fine-tune the parameter adapter ΔW in the large language model and the weight matrix W of the linear layer in the gated twin large language model memory interleaving mechanism A and W B Optimize through the cross-entropy loss function: where T is the length of the output sequence.
6. The application according to claim 1, characterized in that, In the gating-based memory interleaving mechanism, at each time step, access and of the last layer hidden states, and concatenate the hidden states along the feature dimension; construct a low-rank linear layer at the top of the model to generate a gate, which determines the contribution of each large language model to the next step in the logical value generation process, expressed by the formula as follows: Among them, denotes vector concatenation, f represents a low-rank linear layer, and respectively denote freezing the last hidden state of the large language model and fine-tuning the last hidden state of the large language model; where d represents the dimension of the hidden layer state; W A and W B are respectively the weight matrices of the linear layer; r is a hyperparameter; Updated final hidden state And the logical values logits can be calculated in the following way: Among them, represents element-wise multiplication, and W out represents the parameters of the twin large language model linear layer for predicting the next token.
7. An adjustment system for implementing the application described in any one of the above claims 1-6, characterized in that, The adjustment system includes: a twin large language model generation module, a memory interleaving module, and a fine-tuning module; The twin large language model generation module is used to generate two twin large language models that share the same structure and pre-trained parameters; The memory interleaving module is used to adjust the contributions of the memories of the frozen model and the fine-tuned model to generate the next token; The fine-tuning module is used to perform parameter-efficient fine-tuning on one of the twin large language models.