A language model adapter's free training transfer method
Patent Information
- Application Number
- CN202511619920.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-11-06
AI Technical Summary
本发明公开了一种语言模型适配器的免训练迁移方法,旨在解决现有技术中跨模型免训练适配器无法直接迁移至其他结构或规模的模型上使用的问题
本发明能够将基于原语言模型训练的适配器,在不做任何参数更新的前提下,直接迁移给目标语言模型应用,增强目标语言模型的能力。通过利用不同语言模型具有相似的阶级注意力这一性质,将阶级注意力作为跨模型相似的注意力特征表示,并作为适配器的输入,使训练出适配器的适配器适用于不同语言模型,进而能够在免训练的情况下实现适配器的跨模型迁移。
Smart Images

Figure CN121435970B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of language model and natural language processing technology, and specifically relates to a training-free transfer method for a language model adapter. Background Technology
[0002] Pre-trained language models, with their powerful general representation capabilities, have demonstrated outstanding performance in numerous natural language processing tasks. However, in many specialized and domain-specific scenarios, directly using general-purpose models often fails to meet practical needs. Therefore, a common approach is to fine-tune a lightweight adapter module on top of the pre-trained model, thereby adapting the model to downstream tasks at a lower cost. This parameter-efficient fine-tuning method has become an important paradigm for applying large language models. However, the adapters trained by existing methods are usually strongly coupled with the original base model, making them unsuitable for direct transfer to models of other structures or scales. Once the base model needs to be upgraded or replaced, the original adapter becomes invalid and must be retrained, resulting in significant waste of computational resources and severely limiting the flexibility of model iteration and deployment. Therefore, achieving cross-model transfer of adapters is expected to significantly improve the reusability of model components and reduce deployment and maintenance costs. Several methods have attempted to alleviate the adapter-model binding problem, such as transferring the capabilities of existing adapters to new models through knowledge distillation. However, these methods still rely on additional training processes and cannot avoid re-consuming computational resources, thus failing to achieve truly training-free transfer.
[0003] To date, how to achieve direct transfer of adapters between pre-trained language models with different structures and parameter scales without retraining remains an open research question with significant practical implications. Summary of the Invention
[0004] (1) Technical problems to be solved This invention discloses a training-free transfer method for language model adapters, aiming to solve the problem in the prior art that cross-model training-free adapters cannot be directly transferred to models of other structures or sizes.
[0005] (2) Technical solution This invention discloses a training-free transfer method for a language model adapter, comprising the following steps: Step 101, Transferable Adapter Training Phase: Based on the original language model, train a transferable adapter that carries knowledge of the downstream task data on the downstream task data; Step 102, Adapter Migration Phase: Apply the transferable adapter to the target language model to enhance its capabilities.
[0006] Further, step 101 includes the following steps: Step 10101: Sample batch data from downstream task data, including input data and real data annotations; Step 10102: Input the input data into the original language model and calculate the hierarchical attention; Step 10103: Concatenate the attention levels from 1 to k to obtain attention features. Use a transferable adapter and calculate the prediction results for the downstream task using the attention features and the parameters of the transferable adapter. Then use the loss function and calculate the loss for the downstream task using the prediction results and real data annotations. Step 10104: Optimize the parameters of the migrated adapter based on the downstream task loss and output the optimized migrated adapter.
[0007] Furthermore, in step 10104, if the number of training steps reaches a preset number, the parameters of the transferable adapter carrying the data knowledge of downstream tasks are output to obtain the optimized transferable adapter; otherwise, the process jumps to step 10101.
[0008] Furthermore, the computation of hierarchical attention in step 10102 includes the following steps: Step 1010201: Input the input data into the original language model and concatenate them to obtain the multi-head attention of each layer; Step 1010202: Calculate the mean for each head of the multi-head attention in each layer to obtain the mean attention for each layer; Step 1010203: Calculate hierarchical attention by using different mean attention aggregation patterns and combining them with aggregation weights.
[0009] Further, the calculation expression for calculating hierarchical attention in step 1010203 is as follows:
[0010] in, This means inputting the number The first, calculated from the input original language model Class attention of the class Indicates the order, This represents the total number of layers in the original language model. The maximum order of hierarchical attention is the same as the total number of layers in the original language model. The original language model is represented by the first... Mean attention of layers, Represented as aggregate weight; Represents permutation and combination symbols; the 0th order class attention is the identity matrix. ;
[0011] This indicates that the original language model will be divided into 1 to 2. All continuous layers Summing the products of the mean attention of the layers Indicates the continuity of the original language model Layer number index.
[0012] Furthermore, step 102 includes the following steps: Step 10201, put the data to be inferred Input target language model And calculate the order attention from 1 to k, expressed as:
[0013] in, This indicates that the data to be inferred is... Input target language model The calculated first Class attention of the class A computational function representing class attention; Step 10202: Input the calculated hierarchical attention from order 1 to k into the optimized transferable adapter to obtain the optimized prediction results for the downstream task. The expression is:
[0014] in, This represents the optimized parameters of the transferable adapter; This indicates an optimized, portable adapter; This represents the attention feature obtained by concatenating the hierarchical attention from level 1 to level k. This represents the downstream task prediction results obtained by the transferable adapter after the target language model is fused and optimized.
[0015] Furthermore, the downstream tasks include a variety of natural language processing tasks, such as named entity recognition and relation extraction.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention enables the direct transfer of an adapter trained on the original language model to the target language model without any parameter updates, thereby enhancing the target language model's capabilities. By leveraging the similarity of hierarchical attention among different language models, hierarchical attention is used as a cross-model similarity attention feature representation and as input to the adapter. This allows the trained adapter to be applied to different language models, thus achieving cross-model transfer of the adapter without training. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0018] Figure 2 This is a schematic diagram of the calculation process for class attention in this invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] like Figure 1 As shown in this embodiment, suppose we fine-tune a language model to obtain an adapter for applying it to a specific task. If a general fine-tuning method is used, the trained adapter cannot be applied to other language models. Therefore, this invention discloses a training-free transfer method for language model adapters. The adapter fine-tuned by this method can be applied to multiple different language models, thereby reducing training costs.
[0021] Specifically, it includes the following steps: Step 101, Transferable Adapter Training Phase: Based on the original language model, train a transferable adapter that carries knowledge of the downstream task data on the downstream task data; This step ensures that the trained transferable adapter carries relatively "pure" task knowledge, rather than "operational instructions" that are tightly coupled with the model, thereby enabling the transferable adapter to be transferred to another base model with a similar structure but different scale or parameters.
[0022] Step 102, Adapter Migration Phase: Apply the transferable adapter to the target language model to enhance its capabilities.
[0023] The target language model can be a different version of the original language model, or a model with a different structure or scale. When the trained transferable adapter is applied to the target language model, the target language model will immediately obtain the specific capabilities carried by the adapter. Even if the underlying basic model technology is iterated and replaced, the data carried in the transferable adapter can be continuously used instead of being eliminated along with the old model.
[0024] Specifically, the training phase of the transferable adapter involves training a transferable adapter carrying downstream data knowledge on the downstream data based on the original language model, and includes the following steps: Step 10101, from downstream task data Batch data sampled from the middle ,in, It is the input data. It is real data annotation; the downstream tasks include a variety of natural language processing tasks: named entity recognition, relation extraction, etc. Step 10102, input data Input original language model To calculate class attention, the expression is:
[0025] in, This means inputting the input data into the original language model. Calculate the first Class attention of the class Represents the original language model. The order of class attention. This represents the total number of layers in the original language model. The maximum order of hierarchical attention is the same as the total number of layers in the original language model. The computational function representing class attention; the role of computational class attention is to "purify" and "abstract" the complex attention mechanisms in the original model.
[0026] Step 10103: Since class attention has cross-model similarity, class attention from level 1 to the kth level is used as cross-model similarity attention features and input into the transferable adapter to obtain predictions for downstream tasks. The downstream task loss is then calculated, expressed as:
[0027] in, This indicates losses in downstream tasks. Indicates input data The attention features are obtained by concatenating the hierarchical attention from level 1 to level k calculated by the original language model. The parameters represent the transferable adapter, used to adapt attention features to downstream tasks; This represents a transferable adapter, which takes the concatenated attention features and the parameters of the transferable adapter as input and outputs the prediction results for downstream tasks. This represents the loss function used to measure the difference between the model's predictions and the actual data labels; It is real data annotation, used as a supervisory signal to calculate the error between the predicted value and the true value, thereby guiding the update of model parameters; It should be noted that, since the 0th-order class attention is usually an identity matrix for any input data, it has little impact on the calculation of loss and parameter optimization, so there is no need to consider the 0th-order class attention here.
[0028] Step 10104, based on downstream task losses Optimize the parameters of the transferable adapter This enables the transferable adapter to enhance performance on downstream tasks; the specific steps are as follows: if the training steps reach a preset number, output the parameters of the transferable adapter that carry the data knowledge of the downstream task. If the optimized transferable adapter is obtained, then proceed to step 10101. Since the attention of the same level in different language models is similar, the optimized transferable adapter can be directly applied to other language models without retraining.
[0029] Furthermore, such as Figure 2 As shown, the computation of hierarchical attention in step 10102 includes the following steps: Step 1010201, input data Input original language model The splicing results in multiple heads of attention for each layer. ;in, The original language model is represented by the first... Multi-head attention in layers Indicates the number of layers in the original language model This represents the total number of layers in the original language model; Step 1010202: Calculate the mean for each head of the multi-head attention for each layer to obtain the mean attention for each layer. ;in Indicates the first Mean attention of a layer; the purpose of calculating mean attention is to aggregate the multi-head attention of the same layer to obtain a comprehensive representation of the attention of each layer; Step 1010203: Calculate hierarchical attention using different mean attention aggregation patterns and aggregation weights. The specific expression is as follows:
[0030] in, This means input data The first result obtained by inputting the original language model Class attention of the class Indicates the order, This represents the total number of layers in the original language model. The maximum order of hierarchical attention is the same as the total number of layers in the original language model. Represented as aggregate weight; Represents permutation and combination symbols; the 0th order class attention is the identity matrix. ;
[0031] This indicates that the original language model will be divided into 1 to 2. All continuous layers Summing the products of the mean attention of the layers Indicates the original language model connection Layer number index; For example, when It is 5. When the value is 2, first find all consecutive two-layer combinations in layers 1 to 5 of the original language model, including (1,2), (2,3), (3,4) and (4,5), then calculate the product of the mean attention of each combination, and finally add the products of the mean attention of all combinations. By calculating hierarchical attention, the mean attention of each layer of the language model is aggregated. Different aggregation modes are used for different layers. For example, the first-order hierarchical attention represents the context aggregation effect of all paths with an aggregation count of 1; the second-order hierarchical attention represents the context aggregation effect of all paths with an aggregation count of 2.
[0032] Further, step 102 includes the following steps: Step 10201, put the numbers to be deduced Input target language model And calculate the order attention from order 1 to order k, expressed as:
[0033] in, This indicates that the number to be inferred is... Input target language model The calculated first Class attention; Step 10202: Input the attention features obtained by concatenating the attention levels from the first to the kth order obtained from the target language model into the optimized transferable adapter to obtain the optimized downstream task prediction results. The expression is:
[0034] in, This represents the optimized parameters of the transferable adapter; This indicates an optimized, portable adapter; This represents the downstream task prediction results obtained by the transferable adapter after the target language model fusion and optimization. This represents the attention feature obtained by concatenating the hierarchical attention from level 1 to level k; this process enables cross-model application of the adapter without requiring additional training or parameter tuning of the target language model, thus significantly reducing the cost and complexity of model transfer.
[0035] In practical applications, we can first train an adapter for a specific downstream task based on an original language model (e.g., Qwen) using the training-free transfer method of the language model adapter, and then directly transfer the adapter to the target model (e.g., Deepseek) to improve the performance of the target model in that downstream task.
[0036] Furthermore, the effectiveness of this method benefits from the similarity of hierarchical attention across different language models. Although language models may differ in parameters and structure, they often focus on similar syntactic structures and semantic information when processing natural language, which is reflected in hierarchical attention. Therefore, by using hierarchical attention as a unified feature representation across models, this method can train adapters applicable to multiple language models, achieving true training-free transfer.
[0037] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0038] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style of the specification is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0039] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A training-free transfer method for a language model adapter, characterized in that, Includes the following steps: Step 101, Transferable Adapter Training Phase: Based on the original language model, train a transferable adapter that carries knowledge of the downstream task data on the downstream task data; Step 102, Adapter Transfer Phase: Apply the transferable adapter to the target language model to enhance its capabilities; Step 101 includes the following steps: Step 10101: Sample batch data from downstream task data, including input data and real data annotations; Step 10102: Input the input data into the original language model and calculate the hierarchical attention; Step 10103, transfer 1 to Attention features are obtained by concatenating the hierarchical attention of each level. A transferable adapter is used to calculate the prediction results of the downstream task using the attention features and the parameters of the transferable adapter. Then, a loss function is used to calculate the loss of the downstream task using the prediction results and real data annotations. Step 10104: Optimize the parameters of the transferable adapter based on the downstream task loss, and output the optimized transferable adapter; The calculation of hierarchical attention in step 10102 includes the following steps: Step 1010201: Input the input data into the original language model and concatenate them to obtain the multi-head attention of each layer; Step 1010202: Calculate the mean for each head of the multi-head attention in each layer to obtain the mean attention for each layer; Step 1010203: Calculate hierarchical attention by using different mean attention aggregation patterns and combining them with aggregation weights; Step 102 includes the following steps: Step 10201, put the data to be inferred Input target language model And calculate 1 to The class attention of a class is expressed as: ; in, This indicates that the data to be inferred is... Input target language model The calculated first Class attention of the class A computational function representing class attention; Step 10202, calculate the first to... The hierarchical attention of each step is input into the optimized transferable adapter to obtain the optimized prediction results for downstream tasks. The expression is: ; in, This represents the optimized parameters of the transferable adapter; This indicates an optimized, portable adapter; Indicates 1 to the number Attention features obtained by splicing together the attention of different social classes; This represents the downstream task prediction results obtained by the transferable adapter after the target language model fusion and optimization. The downstream tasks include a variety of natural language processing tasks, including named entity recognition and relation extraction. When processing natural language, each language model will pay attention to similar syntactic structures and semantic information.
2. The training-free transfer method for a language model adapter according to claim 1, characterized in that, In step 10104, if the number of training steps reaches the preset number of steps, the parameters of the transferable adapter carrying the data knowledge of the downstream task are output to obtain the optimized transferable adapter; otherwise, the process jumps to step 10101.
3. The training-free transfer method for a language model adapter according to claim 1, characterized in that, The calculation expression for calculating class attention in step 1010203 is as follows: ; in, This means input data The first result obtained by inputting the original language model Class attention of the class Indicates the order, This represents the total number of layers in the original language model. The maximum order of hierarchical attention is the same as the total number of layers in the original language model. The original language model is represented by the first... Mean attention of layers, Represented as aggregate weight; Represents permutation and combination symbols; the 0th order class attention is the identity matrix. ; ; This indicates that the first to second parts of the original language model will be used. All continuous layers The products of the mean attention of each layer are summed. Indicates the continuity of the original language model Layer number index.
Citation Information
Patent Citations
Efficient lightweight parameter migration method for large pre-training language model
CN116484832A