Method and device for generating simulator for pre-trained large model
By performing progressive cropping and knowledge distillation of pre-trained large models and only updating the parameters of the PEFT module, the problem of generating emulators with performance meet the requirements is solved, and efficient emulator generation is achieved.
Patent Information
- Application Number
- CN202510100580.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
When fine-tuning the pretrained large models across domains, generating simulators with performance meet the requirements requires a large amount of computing resources, storage resources, and time resources.
The second largest model is gradually cut and knowledge distilled, and only the parameters of the uncropped Parameters High Efficiency Fine Tweak (PEFT) module are updated to generate a simulator with performance compliance.
With relatively few computing resources, storage resources and time resources consumption, an emulator with performance meets the requirements can be obtained.
Smart Images

Figure CN119990216A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of computer technology, and more particularly, to a method and device for generating a simulator for a pre-trained large model. Background Art
[0002] Based on the need for privacy protection, it may be necessary to perform cross-domain fine-tuning on the pre-trained large model. In the process of cross-domain fine-tuning of the pre-trained large model, the model provider can compress the pre-trained large model through lossy compression technology to generate an emulator corresponding to the large model, and the emulator will be passed to the data provider; the data provider can train an adapter corresponding to the large model based on the emulator and its own training data set, and the adapter will be passed to the model provider; the model provider can plug the adapter into the pre-trained large model to obtain the target large model that can be used to perform downstream tasks, completing the cross-domain fine-tuning process.
[0003] In the process of generating the simulator, multiple network layers can be cut from the pre-trained large model, the cut large model is used as the student model, the uncut large model is used as the teacher model, and the knowledge distillation is performed on the cut large model, so as to obtain the simulator that will be passed to the data provider. However, in this embodiment, a large number of parameters included in the cut large model need to be directly updated in a relatively large number of iterations to obtain a simulator with performance that meets the requirements, that is, a large amount of computing resources, storage resources and time resources need to be consumed to obtain a simulator with performance that meets the requirements. Summary of the invention
[0004] The object of the present invention is to provide a method and device for generating a simulator for a pre-trained large model.
[0005] In a first aspect, a method for generating an emulator for a pre-trained large model is provided, wherein the emulator is used to perform cross-domain fine-tuning on a pre-trained first large model, wherein the first large model includes N functional units, and the method includes: performing H rounds of first updates, wherein any h-th round of first updates includes: cutting a number of functional units from a second large model that has been first updated in the h-1th round according to the importance scores of each functional unit in the first large model that has been cut in the h-1th round of first updates, and their corresponding parameter-efficient fine-tuning (Parameter-efficient Fine-tuning); The second largest model after the h-th round of tuning is obtained by a PEFT (Pre-tuning and PEFT) module, wherein the functional units included in the second largest model after the h-1th round of first update are the same as the functional units included in the first large model after the h-1th round of tuning; the first input sample is processed by the second largest model after the h-th round of tuning to obtain a first processing result; the parameters in each PEFT module included in the second largest model after the h-th round of tuning are updated according to the first processing result and the second processing result, wherein the second processing result is obtained by processing the first input sample by the first large model after the h-1th round of tuning; and a simulator corresponding to the first large model is generated according to the second largest model after the h-1th round of first update.
[0006] In a second aspect, a device for generating a simulator for a pre-trained large model is provided, wherein the simulator is used to perform cross-domain fine-tuning on a pre-trained first large model, wherein the first large model includes N functional units, and the device includes: an update processing unit, which is used to perform H rounds of first updates, wherein any h-th round of first updates includes: according to the importance scores of each functional unit in the first large model that has been pruned in the h-1th round in the first update of the h-1th round, a number of functional units and their corresponding PEFT modules are pruned from a second large model that has been pruned in the h-1th round, to obtain a second large model that has been pruned in the h-1th round, The functional units included in the updated second largest model are the same as the functional units included in the first large model that has been pruned for the h-1th round; the first input sample is processed by the second largest model that has been pruned for the hth round to obtain a first processing result; the parameters in each PEFT module included in the second largest model that has been pruned for the hth round are updated based on the first processing result and the second processing result, wherein the second processing result is obtained by processing the first input sample using the first large model that has been pruned for the h-1th round; a simulator generation unit is used to generate a simulator corresponding to the first large model based on the second largest model that has been first updated for the Hth round.
[0007] In a third aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores a computer program / instructions, and when the processor executes the computer program / instructions, the method described in the first aspect is implemented.
[0008] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed in a computing device, the computing device executes the method described in the first aspect.
[0009] In the technical solutions provided in the embodiments of this specification:
[0010] The second largest model, which serves as the student model, is pruned and knowledge distilled by progressive execution. During the knowledge distillation of the second largest model, only the parameters in the PEFT modules that have not been pruned at the current moment are updated. This allows a simulator with performance that meets the requirements to be obtained with relatively less computing resources, storage resources, and time resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0012] Figure 1 A schematic diagram of the relationship between large models is provided as an example in the embodiments of this specification;
[0013] Figure 2 A schematic diagram of a process of performing the hth round of first update on the second largest model provided as an example in an embodiment of this specification;
[0014] Figure 3 A schematic diagram of a process of performing the k-th round of second updating on the second largest model provided as an example in the embodiments of this specification;
[0015] Figure 4 A schematic diagram of a large progressive cropping model exemplarily provided in an embodiment of this specification;
[0016] Figure 5 A schematic diagram of the structure of a device for generating a simulator for a pre-trained large model provided in an embodiment of this specification. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0018] Based on the need for privacy protection, it may be necessary to perform cross-domain fine-tuning on the pre-trained large model (hereinafter referred to as the first large model). In the process of cross-domain fine-tuning of the first large model, the model provider can use lossy compression technology to perform lossy compression on the first large model to generate an emulator corresponding to the source model, and the emulator will be passed to the data provider; the data provider can train an adapter corresponding to the first large model based on the emulator and its own training data set, and the adapter will be passed to the model provider; the model provider can plug the emulator into the first large model to obtain a second large model that can be used to perform downstream tasks.
[0019] During the cross-domain fine-tuning process, on the one hand, the model provider cannot know the training data set held by the data provider, and the data provider cannot know the first large model provided by the model provider. At the same time, the data provider cannot know the second large model that can be used to perform downstream tasks and can meet the needs of privacy protection. On the other hand, the simulator is obtained by lossy compression of the first large model, and the parameter scale is relatively small. The data provider only needs to consume relatively small computing resources, storage resources, and time resources to train an adapter with performance that meets the requirements. On the other hand, the simulator is obtained by compressing the first large model. The adapter trained based on the simulator cannot have the same performance as the second large model obtained through cross-domain fine-tuning. This is also a reflection of the security of the first and second large models.
[0020] The key challenge of cross-domain fine-tuning is how to effectively perform lossy compression of the first large model; that is, the key challenge of cross-domain fine-tuning is how to obtain a simulator corresponding to the first large model.
[0021] In a possible implementation, multiple network layers can be cut from the first large model, the cut first large model is used as the student model, the uncut first large model is used as the teacher model, and the knowledge distillation is performed on the cut first large model, so as to obtain a simulator that will be passed to the data provider. Exemplarily, the parameters of the first two network layers and the last two network layers of the first large model can be kept unchanged, and the remaining network layers can be pruned with equal steps. For example, for the 3rd to 10th network layers, only the 3rd, 5th, 7th, and 9th network layers are kept and the 4th, 6th, 8th, and 10th network layers are pruned; the uncut first large model is used as the teacher model, and the cut first large model is used as the student model, and the knowledge distillation is performed on the student model so that the student model imitates the input and output behaviors of the teacher model. The student model that has completed the knowledge distillation will be used as a simulator.
[0022] In this implementation, a large number of parameters in the pruned first large model need to be updated in a large number of iteration rounds to obtain a simulator with performance that meets the requirements, which consumes a large amount of computing resources, storage resources, and time resources.
[0023] The embodiments of this specification provide a method and device for generating a simulator for a pre-trained large model. The simulator is used to perform cross-domain fine-tuning on a pre-trained first large model, which includes N functional units. H rounds of first updates can be performed first, and any h-th round of first updates includes: according to the importance scores of each functional unit in the first large model that has been pruned in the h-1th round of first updates, a number of functional units and their corresponding PEFT modules are pruned from the second large model that has been pruned in the h-1th round of first updates to obtain the second large model that has been pruned in the h-1th round of first updates, and the functional units included in the second large model that has been pruned in the h-1th round of first updates are the same as the functional units included in the first large model that has been pruned in the h-1th round of first updates; the first input sample is processed using the second large model that has been pruned in the hth round to obtain a first processing result; according to the first processing result and the second processing result, the parameters in each PEFT module included in the second large model that has been pruned in the h-1th round of first updates are updated, and the second processing result is obtained by processing the first input sample using the first large model that has been pruned in the h-1th round of first updates; and then, according to the second large model that has been pruned in the Hth round of first updates, a simulator corresponding to the first large model is generated.
[0024] In this way, the second largest model serving as the student model is pruned and knowledge distilled by progressive execution. During the knowledge distillation of the second largest model, only the parameters in the PEFT modules that have not been pruned at the current moment are updated. This allows a simulator with performance that meets the requirements to be obtained with relatively less computing resources, storage resources, and time resources.
[0025] The relationship between the uncropped first large model and the uncropped second large model is first described below.
[0026] The pre-trained first large model is used as the source model for cross-domain fine-tuning. When the first large model is not pruned, the first large model usually includes multiple Transformer modules stacked in sequence. The Transformer module can be decomposed into a multi-head attention (MHA) module and a multi-layer perceptron (MLP) module.
[0027] N functional units that allow lossy compression can be determined from the multiple Transfomer modules included in the first large model. If coarse-grained pruning is required, the N functional units can be N Transfomer modules belonging to the multiple Transfomer modules; if fine-grained pruning is required, the N functional units can include N / 2 MHA modules and N / 2 MLP modules included in N / 2 Transfomer modules belonging to the multiple Transfomer modules.
[0028] The MHA module and MLP module included in the same Transfomer module may have different importance in the first large model. Adopting coarse-grained pruning and taking the Transfomer module as the pruning granularity may cause the MHA module or MLP module with relatively large importance to be pruned, which will eventually lead to a significant decline in the performance of the simulator obtained subsequently. Therefore, fine-grained pruning is preferred.
[0029] Currently, there are provided methods for fine-tuning parameters of pre-trained large models, such as adapter-based methods, re-parameterization-based methods, and prompt-based methods, which aim to minimize the number of trainable parameters. The various parameter-efficient fine-tuning methods in the above examples all require setting up multiple PEFT modules for the pre-trained large model, and completing the process of fine-tuning the pre-trained large model by updating the parameters of the multiple PEFT modules. On this basis, the N functional units determined can be configured with N PEFT modules corresponding to the N functional units; and then referring to Figure 1 As shown, the first large model and N PEFT modules corresponding to the N functional units can be used to form a second large model corresponding to the first large model.
[0030] The following takes N PEFT modules specifically including N low-rank adaptation (LORA) modules as an example to exemplarily describe the relationship between the N PEFT modules in the second largest model and the N functional units included in the first largest model.
[0031] First, any nth LoRa module among the N LoRa modules corresponds to the nth functional unit among the N functional units included in the first large model. Here, it is assumed that the parameter matrix of the nth functional unit is W Pn , then the nth LORA module includes the parameter matrix a Mn and b Mn , for a Mn and b Mn The result of multiplication is W Mn , and the parameter matrix W of the nth functional unit Pn The number of rows and columns is the same. Among them, W Mn It represents the parameter matrix W of the nth functional unit during the update of the second largest model. Pn The increments of the parameters included; therefore, in the process of updating the second largest model, only a needs to be updated Mn and b Mn The parameters in are sufficient, and there is no need to update the parameter matrix W of the nth functional unit. Pn A large number of parameters are included; when the iteration termination condition is met, it can be based on a Mn and b Mn Calculate W Mn , the parameter matrix W of the nth functional unit Pn Update to W Pn +W Mn That's it.
[0032] In addition, it is assumed here that in the first large model independent of the second large model, the input of the n+1th functional unit among the N functional units is is the output of the nth functional unit Then in the second largest model corresponding to the first largest model, the corresponding functional units belong to parallel branches, and the input of the n+1th functional unit and the input of the n+1th PEFT module should be the output of the nth functional unit The output of the nth PEFT module The combined result between and In fact, it can be
[0033] It is understandable that when N PEFT modules and the first large model are used to form a second large model corresponding to the first large model, the parameter matrices of the N PEFT modules need to be initialized. For example, for any parameter matrix a included in the nth LORA module Mn and b Mn , we need to set the parameter matrix aMn and b Mn Initialize and ensure that the initialized a Mn and b Mn In the incremental matrix obtained by multiplication, the value of each element is 0, or is less than a preset value close to 0.
[0034] After obtaining the second largest model corresponding to the first largest model, a method for generating a simulator for a pre-trained model provided in an embodiment of this specification can be executed based on the first large model and the second large model, so as to achieve trimming and knowledge distillation of the second large model through progressive execution, and complete the generation of a simulator for cross-domain fine-tuning of the pre-trained first large model.
[0035] A method for generating a simulator for a pre-trained model provided in an embodiment of this specification can be divided into two generation stages, that is, the process of generating a simulator for cross-domain fine-tuning can be divided into two generation stages. In the first generation stage, H rounds of first updates can be performed on the second largest model; in the second generation stage, a simulator corresponding to the first large model can be generated based on the second largest model that has undergone H rounds of first updates, for example, H-1 rounds of second updates are performed on the second largest model that has undergone H rounds of first updates, and then the simulator corresponding to the first large model is determined based on the second largest model that has undergone H rounds of second updates.
[0036] The following is an exemplary description of the process of performing any h-th round of first update on the second largest model in the first generation phase.
[0037] Figure 2 This is a schematic diagram of a process of performing the hth round of first update on the second largest model provided as an example in the embodiments of this specification.
[0038] Reference Figure 2 As shown, any h-th round of first update may include part or all of the following steps S201 to S209.
[0039] In step S201, the importance scores of the functional units in the first large model that has undergone the h-1th round of pruning in the first update of the hth round are determined.
[0040] In one possible implementation, for any i-th functional unit included in the first large model that has undergone the h-1th round of trimming, the i-th functional unit can first be trimmed from the first large model that has undergone the h-1th round of trimming to obtain the i-th test model; then, based on the i-th test model, the importance score of the i-th functional unit in the first update of the h-1th round can be determined.
[0041] Exemplarily, the number of functional units that are allowed to be lossy compressed in the first large model is N, the number of functional units pruned in each round of pruning of the first large model is T, and the number of functional units included in the first large model after the h-1th round of pruning is NT*(h-1). During the process of performing the hth round of first update on the second large model, the importance scores of the NT*(h-1) functional units included in the first large model after the h-1th round of pruning in the hth round of first update can be determined.
[0042] The number of execution rounds H of the first update, and the number of pruned functional units / the number of PEFT modules T for each round of pruning performed on the first large model or the second large model, can be flexibly set based on the compression rate that needs to be met. For example, for the first large model containing 64 functional units, the compression rate of lossy compression is required to be 50%, then the number of execution rounds of the first update can be set to 8, and the number of pruned functional units for each round of pruning performed on the first large model can be set to 4; for another example, for the first large model containing 12 functional units, the compression rate of lossy compression is required to be 50%, then the number of execution rounds of the first update can be set to 6, and the number of pruned functional units for each round of pruning performed on the first large model can be set to 1.
[0043] For example, refer to Figure 3As shown, it is assumed here that the unpruned first large model includes 12 functional units, namely functional units 1 to 12. After 1 to 5 rounds of pruning are performed on the first large model, functional units 9, 8, 3, 11 and 5 are sequentially pruned from the first large model. Then, in the first round of first update executed on the second largest model, it is necessary to calculate 12 importance scores of functional units 1 to 12 in the first round of first update; in the second round of first update executed on the second largest model, the first largest model after the first round of trimming also includes functional units 1 to 8 and functional units 10 to 12, and it is necessary to calculate 11 importance scores of functional units 1 to 8 and functional units 10 to 12 in the first round of first update; in the third round of first update executed on the second largest model, the first largest model after the second round of trimming also includes functional units 1 to 7 and functional units 10 to 12, and it is necessary to calculate 10 importance scores of functional units 1 to 7 and functional units 10 to 12 in the third round of first update; in the fourth round of first update executed on the second largest model, the first largest model after the third round of trimming The model also includes functional units 1, 2, 4-7 and functional units 10-12, and it is necessary to calculate 9 importance scores of functional units 1, 2, 4-7 and functional units 10-12 in the first update of the 4th round; in the first update of the 5th round performed on the second largest model, the first largest model after the 4th round of trimming also includes functional units 1, 2, 4, 6, 7 and functional units 10-12, and it is necessary to calculate 8 importance scores of functional units 1, 2, 4, 6, 7 and functional units 10-12 in the first update of the 5th round; in the first update of the 6th round performed on the second largest model, the first largest model after the 5th round of trimming also includes functional units 1, 2, 4, 6, 7, 10, 12, and it is necessary to calculate 7 importance scores of functional units 1, 2, 4, 6, 7, 10, 12 in the first update of the 6th round.
[0044] In a more specific example, the i-th test model can be used to process the sample word sequence to obtain the predicted probability sequence corresponding to the word sequence; then, the perplexity of the i-th test model is calculated based on the predicted probability sequence; finally, the importance score of the i-th functional unit in the first update of the h-th round is determined based on the perplexity of the i-th test model. It is not difficult to understand that the perplexity and the importance score are positively correlated. The smaller the perplexity of the i-th test model, the better the performance of the i-th test model, which also means that the importance score of the i-th functional unit in the first update of the h-th round is lower.
[0045] The importance score of the i-th functional unit in the first update of the h-th round can also be determined according to the i-th test model in other ways. For example, the first loss of the i-th test model on a certain validation data set can be calculated, and the second loss of the first large model pruned in the h-1th round on the validation data set can be calculated. The loss change rate calculated based on the first loss and the second loss is positively correlated with the importance score of the i-th test model in the first update of the h-th round.
[0046] Alternatively, determining the importance score of each functional unit in the first large model that has been pruned for the h-1th round in the first update of the hth round may be achieved through other implementations. For example, the N importance scores of N functional units in the first large model that has not been pruned may be pre-calculated; for any functional unit among the N functional units, the importance score of the functional unit in the first update of the hth round always remains unchanged, that is, if the functional unit exists in the first large model that has been pruned for the h-1th round, the importance score of the functional unit in the first update of the hth round is still the importance score of the functional unit in the first large model that has not been pruned. This means that the aforementioned step S201 is actually optional and not necessary.
[0047] When the execution round h is less than H, the following step S203 can also be executed, according to the importance scores of each functional unit in the first large model after the h-1 round of trimming in the first update of the h round, several functional units are trimmed from the first large model after the h-1 round of trimming to obtain the first large model after the h round of trimming.
[0048] And, in step S205, according to the importance scores of each functional unit in the first large model that has been pruned in the h-1th round in the first update of the h-1th round, several functional units and their corresponding parameter efficient fine-tuning PEFT modules are pruned from the second large model that has been pruned in the h-1th round for the first time to obtain the second large model that has been pruned in the h-1th round, and the functional units included in the second large model that has been pruned in the h-1th round for the first update are the same as the functional units included in the first large model that has been pruned in the h-1th round.
[0049] Continuing with the previous example, after obtaining the importance scores of the NT*(h-1) functional units included in the first large model that has been pruned after the h-1th round in the first update of the hth round, T functional units can be pruned from the first large model that has been pruned after the h-1th round in order of importance scores from small to large to obtain the first large model that has been pruned after the hth round; and, the T functional units and their corresponding T PEFT modules can be pruned from the second large model that has been pruned after the h-1th round.
[0050] For example, please see Figure 3 As shown. In the first round of the first update performed on the second largest model, after determining the 12 importance scores of the functional units 1 to 12 in the first round of the first update, among the 12 importance scores, the importance score corresponding to the functional unit 9 is the smallest, so the functional unit 9 can be trimmed from the first large model that has been trimmed in the 0th round; and, the functional unit 9 and its corresponding PEFT module 9 can be trimmed from the second large model that has been trimmed in the 0th round of the first update. Similarly, in the second round of the first update performed on the second largest model, after determining the 11 importance scores of the functional units 1 to 7, 10 to 12 in the second round of the first update, among the 11 importance scores, the importance score corresponding to the functional unit 8 is the smallest, so the functional unit 8 can be trimmed from the first large model that has been trimmed in the 1st round; and, the functional unit 8 and its corresponding PEFT module 8 can be trimmed from the second large model that has been trimmed in the 1st round of the first update. By analogy, in the third round of first update performed on the second largest model, functional unit 3 may be cropped from the first large model after the second round of cropping, and functional unit 3 and its corresponding PEFT module 3 may be cropped from the second large model after the second round of first update; in the fourth round of first update performed on the second large model, functional unit 11 may be cropped from the first large model after the third round of cropping, and functional unit 11 and its corresponding PEFT module 11 may be cropped from the second large model after the third round of first update; in the fifth round of first update performed on the second large model, functional unit 5 may be cropped from the first large model after the fourth round of cropping, and functional unit 5 and its corresponding PEFT module 5 may be cropped from the second large model after the fourth round of first update; in the sixth round of first update performed on the second large model, functional unit 6 and its corresponding PEFT module 6 may be cropped from the second large model after the fifth round of first update.
[0051] After obtaining the second largest model after the h-th round of cropping, it can be used as the student model, and the first largest model after the h-1-th round of cropping can be used as the teacher model. On the same input sample, the second largest model after the h-th round of cropping is subjected to knowledge distillation, that is, through the following steps S207 and S209, the knowledge distillation of the second largest model after the h-th round of cropping is realized.
[0052] Step S207, using the second largest model that has undergone the h-th round of pruning to process the first input sample to obtain a first processing result.
[0053] Step S209, updating the parameters of each PEFT module included in the second largest model after the hth round of pruning according to the first processing result and the second processing result, wherein the second processing result is obtained by processing the first input sample using the first largest model after the h-1th round of pruning.
[0054] When updating the parameters in each PEFT module included in the second largest model after the h-th round of pruning, the goal is to minimize the difference between the first processing result and the second processing result. It can be understood that in the h-th round of first update performed on the second largest model, steps S207 and S209 can be performed multiple times, and when the difference between the first processing result and the second processing result meets a certain preset condition or the number of iterations reaches a threshold, the h-th round of first update performed on the second largest model is completed.
[0055] After completing the H rounds of first updates of the second large model through the aforementioned steps S201 to S209, the second generation stage can be entered. As mentioned above, in the second generation stage, a simulator corresponding to the first large model can be generated based on the second large model that has undergone the H rounds of first updates. For example, the second large model that has undergone the H rounds of first updates is then updated in the H-1 round, and then the simulator corresponding to the first large model is determined based on the second large model that has undergone the H-1 rounds of second updates.
[0056] The following is an exemplary description of the process of performing any k-th round of second update on a large model that has undergone H rounds of first updates in the second generation phase.
[0057] Figure 4 This is a schematic diagram of a process for performing the k-th round of second updating on the second largest model provided as an example in the embodiments of this specification.
[0058] Step S401, using the second largest model that has been second updated in the k-1th round to process the second input sample, to obtain a third processing result. It can be understood that when the value of k is 1, the second largest model that has been second updated in the k-1th round in step S401 refers to the second largest model that has been first updated in the Hth round.
[0059] Step S403, based on the third processing result and the fourth processing result, update the parameters of each PEFT module included in the second largest model after the second update in the k-1th round, wherein the fourth processing result is obtained by processing the second input sample using the first largest model pruned in the Hk-1th round.
[0060] When updating the parameters in the PEFT module included in the second largest model after the k-1th round of second update, the goal is to minimize the difference between the third processing result and the fourth processing result. It can be understood that in the kth round of second update performed on the second largest model, steps S401 and S403 can be performed multiple times, and when the difference between the third processing result and the fourth processing result meets a certain preset condition or the number of iterations reaches a threshold, the kth round of second update performed on the second largest model is completed.
[0061] After completing the H-1 round second update of the second largest model after the H round first update, the simulator corresponding to the first largest model can be determined based on the second largest model after the H-1 round second update. Taking the PEFT module as an example, the structure of the second largest model after the H-1 round second update is the same as the structure of the second largest model after the H round tailoring. For example, please continue to refer to Figure 3 As shown, in the second largest model after the second update of the H-1 round, the functional units 3, 5, 6, 8, 9, 11 and their corresponding LORA modules 3, 5, 6, 8, 9, 11 have been trimmed, that is, the second largest model after the second update of the H-1 round also includes functional units 1, 2, 4, 7, 10, 12 and their corresponding LORA modules 1, 2, 4, 7, 10, 12. It is only necessary to update the parameters of the functional units 1, 2, 4, 7, 10, 12 in the second largest model according to the parameter matrix of the LORA modules 1, 2, 4, 7, 10, 12, and delete the LORA modules 1, 2, 4, 7, 10, 12 from the second largest model after the second update of the H-1 round. The obtained second largest model including only the functional units 1, 2, 4, 7, 10, 12 but not the LORA modules 1, 2, 4, 7, 10, 12 can be used as the generated simulator, so that the simulator can be used to perform cross-domain fine-tuning on the first large model later.
[0062] Based on the same concept as the aforementioned method embodiment, the embodiment of this specification also provides a device 500 for generating a simulator for a pre-trained large model, wherein the simulator is used to perform cross-domain fine-tuning on a pre-trained first large model, wherein the first large model includes N functional units, and the device includes: an update processing unit 501, used to perform H rounds of first updates, wherein any h-th round of first updates includes: based on the importance scores of each functional unit in the first large model pruned after the h-1th round in the first update of the h-1th round, pruned a number of functional units and their corresponding parameter efficient fine-tuning PEFT modules from the second large model that has been pruned after the h-1th round, and obtained the first large model pruned after the h-1th round. The second large model, the functional units included in the second large model after the first update in the h-1th round are the same as the functional units included in the first large model after the h-1th round of trimming; the first input sample is processed by the second large model after the h-1th round of trimming to obtain a first processing result; according to the first processing result and the second processing result, the parameters in each PEFT module included in the second large model after the h-1th round of trimming are updated, wherein the second processing result is obtained by processing the first input sample by the first large model after the h-1th round of trimming; the simulator generation unit 503 is used to generate a simulator corresponding to the first large model according to the second large model after the first update in the h-1th round.
[0063] A computer-readable storage medium is also provided in an embodiment of the present specification, on which a computer program / instruction is stored. When the computer program / instruction is executed in a computer, the computer is caused to execute a method for generating a simulator for a pre-trained large model provided in the aforementioned embodiments.
[0064] A computing device is also provided in an embodiment of the present specification, including a memory and a processor, wherein the memory stores a computer program / instructions, and when the processor executes the computer program / instructions, a method for generating a simulator for a pre-trained large model provided in the aforementioned embodiments is implemented.
[0065] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0066] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0067] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present application does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0068] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any specific order.
[0069] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0070] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0071] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0072] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0073] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0074] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0075] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0076] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0077] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0078] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, in the absence of mutual contradiction, a person skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0079] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.
Claims
1. A method for generating a simulator for a pre-trained large model, wherein the simulator is used to perform cross-domain fine-tuning on a pre-trained first large model, wherein the first large model includes N functional units, and the method comprises: Perform H rounds of first updates. Any h-th round of first updates includes: According to the importance scores of the functional units in the first large model that has been pruned in the h-1th round in the first update in the hth round, a number of functional units and their corresponding parameter efficient fine-tuning PEFT modules are pruned from the second large model that has been pruned in the h-1th round for the first update, to obtain the second large model that has been pruned in the h-1th round, and the functional units included in the second large model that has been pruned in the h-1th round for the first update are the same as the functional units included in the first large model that has been pruned in the h-1th round; Processing the first input sample using the second largest model that has been pruned in the hth round to obtain a first processing result; updating the parameters of each PEFT module included in the second largest model pruned in the hth round according to the first processing result and the second processing result, wherein the second processing result is obtained by processing the first input sample with the first largest model pruned in the h-1th round; Based on the second largest model that has undergone H rounds of first updating, a simulator corresponding to the first largest model is generated.
2. According to the method of claim 1, the PEFT module is a low-rank adaptive LORA module.
3. The method according to claim 1, when h is less than H, the h-th round of first update further comprises: The plurality of functional units are cut out from the first large model after the h-1th round of cutting to obtain the first large model after the hth round of cutting.
4. The method according to claim 1, wherein generating a simulator corresponding to the first large model according to the second large model after H rounds of first updating comprises: For the second largest model after H rounds of first updates, perform H-1 rounds of second updates. Any k-th round of second updates includes: Processing the second input sample using the second largest model that has been updated for the second time in the k-1th round to obtain a third processing result; updating the parameters of each PEFT module included in the second largest model that has been second-updated in the k-1th round according to the third processing result and the fourth processing result, wherein the fourth processing result is obtained by processing the second input sample with the first largest model that has been pruned in the Hk-1th round; According to the second largest model after the second update of the H-1 round, the simulator corresponding to the first largest model is determined. 5 . The method according to claim 1 , wherein the first large model comprises a plurality of Transfomer modules stacked in sequence, and the N functional units are N Transfomer modules belonging to the plurality of Transfomer modules.
6. According to the method of claim 1, the first large model comprises a plurality of Transfomer modules stacked in sequence, and the N functional units comprise N / 2 multi-head attention MHA modules and N / 2 multi-layer perceptron MLP modules included in N / 2 Transfomer modules belonging to the plurality of Transfomer modules.
7. The method according to claim 1, wherein the h-th round of first update further comprises: Determine the importance score of each functional unit in the first large model that has been pruned in the h-1th round in the first update of the hth round.
8. The method according to claim 7, wherein determining the importance scores of the functional units in the first large model that has undergone the h-1th round of pruning in the first update of the hth round comprises: For any i-th functional unit included in the first large model that has undergone the h-1-th round of trimming, trim the i-th functional unit from the first large model that has undergone the h-1-th round of trimming to obtain the i-th test model; Determine, according to the i-th test model, the importance score of the i-th functional unit in the h-th round of first updating.
9. The method according to claim 8, wherein determining the importance score of the i-th functional unit in the h-th round of first update according to the i-th test model comprises: Processing a sample word sequence using the i-th test model to obtain a prediction probability sequence corresponding to the word sequence; Calculating the perplexity of the i-th test model according to the predicted probability sequence; According to the perplexity, an importance score of the i-th functional unit in the first update of the h-th round is determined.
10. A device for generating a simulator for a pre-trained large model, the simulator being used to perform cross-domain fine-tuning on a pre-trained first large model, the first large model comprising N functional units, the device comprising: An update processing unit is used to perform H rounds of first updates, where any h-th round of first updates includes: According to the importance scores of the functional units in the first large model that has been pruned in the h-1th round in the first update in the hth round, a number of functional units and their corresponding parameter efficient fine-tuning PEFT modules are pruned from the second large model that has been pruned in the h-1th round for the first update, to obtain the second large model that has been pruned in the h-1th round, and the functional units included in the second large model that has been pruned in the h-1th round for the first update are the same as the functional units included in the first large model that has been pruned in the h-1th round; Processing the first input sample using the second largest model that has been pruned in the hth round to obtain a first processing result; updating the parameters of each PEFT module included in the second largest model pruned in the hth round according to the first processing result and the second processing result, wherein the second processing result is obtained by processing the first input sample with the first largest model pruned in the h-1th round; The simulator generating unit is used to generate a simulator corresponding to the first large model according to the second large model which has undergone H rounds of first updating.
11. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computing device, the computing device executes the method according to any one of claims 1 to 9.
Citation Information
Cited By
Method and apparatus for generating emulator for pre-trained large model
WO2026157402A1