Method for generating simulator for pre-trained large model
By determining the compression score through the compression parameter set and supporting dataset, a simulator is generated, which solves the problem of high resource consumption in existing technologies, realizes efficient cross-domain fine-tuning of large models, and ensures model security and privacy.
Patent Information
- Application Number
- CN202511056404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies consume a large amount of computing, storage, and time resources in the process of generating simulators, and cannot efficiently perform cross-domain fine-tuning of large models.
By acquiring a set of compression parameters, multiple functional units are compressed, a compression score is determined using a support dataset, and compression parameters that meet preset score requirements are selected to generate a simulator, thus avoiding knowledge distillation training.
The simulator achieves performance that meets cross-domain fine-tuning requirements with limited resources, saving computational resources and time, and ensuring model security and privacy.
Smart Images

Figure CN120975137A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present specification belongs to the technical field of large models, and particularly relates to a method for generating an emulator of a pre-trained large model. BACKGROUND
[0002] Based on the need for privacy protection, it can be necessary to perform offsite tuning on a pre-trained large model. In the process of offsite tuning of the pre-trained large model, the model provider can generate an emulator corresponding to the large model by compressing the pre-trained large model through lossy compression technology. The emulator will be passed to the data holder; the data holder can train an adapter corresponding to the large model according to the emulator and the training data set it holds, and the adapter will be passed to the model provider; the model provider can plug the adapter into the pre-trained large model to obtain a target large model that can be used to perform downstream tasks, and complete the offsite tuning process.
[0003] In the current process of generating an emulator, a plurality of network layers can be cut from the pre-trained large model, the cut large model is taken as a student model, the large model that has not been cut is taken as a teacher model, and knowledge distillation is performed on the cut large model to obtain an emulator that will be passed to the data holder. In this embodiment, the cut large model needs to be subjected to a knowledge distillation training process to obtain an emulator with required performance, i.e., a large amount of computing resources, storage resources and time resources are required for model training to obtain an emulator with required performance. SUMMARY
[0004] The purpose of the present application is to provide a method for generating an emulator of a pre-trained large model, saving the computing resources and time required for compressing the large model to obtain the emulator.
[0005] The first aspect of the present specification provides a method for generating an emulator of a pre-trained large model, the pre-trained first large model comprising a plurality of functional units, comprising:
[0006] obtaining a set of compression parameters, the set of compression parameters comprising a plurality of compression parameters, and compressing the plurality of functional units based on the compression parameters to obtain a first emulator;
[0007] obtaining a support data set, the support data set comprising input data and label data;
[0008] Based on the support dataset, a compression score corresponding to each of the compression parameters is determined, the compression score being positively correlated with a gradient difference between the plurality of functional units and the first simulator in the back propagation process, and being negatively correlated with a loss difference between a second large model including the plurality of functional units and a third large model including the first simulator.
[0009] Based on the compression scores corresponding to the plurality of compression parameters respectively, S compression parameters meeting a preset score requirement are determined, and the plurality of functional units are compressed using the S compression parameters to generate a second simulator corresponding to the first large model, the S being determined based on a parameter compression rate.
[0010] The second aspect of the specification provides a computing device including a memory and a processor, the memory storing executable code, and the processor executing the executable code to implement the method of the first aspect.
[0011] In the scheme provided in the above embodiments of the specification, the compression parameters meeting the preset score requirement are determined by taking the compression score as an index, and then the model is compressed by using the compression parameters. The performance of the compressed simulator directly meets the requirements for privacy and practicality in cross-domain fine-tuning, and there is no need to adjust the network parameters of the compressed simulator, that is, there is no need to perform knowledge distillation and other model training, so that a larger model can be compressed under limited resources, greatly saving computing resources and time. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the specification, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0013] Figure 1 A schematic diagram of cross-domain fine-tuning of a large model is exemplarily provided in the embodiments of the specification.
[0014] Figure 2 A flowchart of a method of generating a simulator for a pre-trained large model is exemplarily provided in the embodiments of the specification.
[0015] Figure 3 A schematic diagram of compressing a transform module is exemplarily provided in the embodiments of the specification.
[0016] Figure 4 A structural schematic diagram of a device for generating a simulator for a pre-trained large model is provided in the embodiments of the specification. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the embodiments of the specification. Obviously, the described embodiments are only some of the embodiments of the specification, not all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the scope of protection of the specification.
[0018] Cross-domain fine-tuning refers to a technology of fine-tuning a large model under the condition that a pre-trained large model is provided by a model provider, a training data set is provided by a data holder, and neither the large model of the model provider nor the training data set of the data holder is out of the domain. Through the technology, bidirectional privacy protection can be achieved. Pre-training refers to training in advance on a large data set with a lot of computing power. Fine-tuning refers to training a pre-trained large model with a small amount of data to meet the needs of a downstream task.
[0019] Based on the need for privacy protection, it can be necessary to fine-tune a pre-trained large model (hereinafter referred to as a first large model) across domains. In the process of fine-tuning the first large model across domains, as shown in FIG. 1, the model provider can generate an emulator corresponding to the original model by lossy compression of the first large model through a lossy compression technology, and then the emulator will be passed to the data holder; the data holder can train an adapter corresponding to the first large model based on the emulator and its own training data set, and the adapter will be passed to the model provider after the training is completed; the model provider can access the adapter to the first large model, thereby obtaining a second large model that can be used to perform a downstream task, and the fine-tuning process is completed. Figure 1
[0020] In the cross-domain fine-tuning process, on the one hand, the model provider cannot know the training data set held by the data holder, the data holder cannot know the first large model provided by the model provider, and the data holder cannot know the second large model finally generated by the model provider and capable of being used to perform the downstream task, which can meet the needs of bidirectional privacy protection. On the other hand, the simulator is obtained by lossy compression of the first large model, and the parameter scale is relatively small. The data holder only needs to consume relatively small computing resources, storage resources and time resources to fine-tune the adapter with required performance. On the other hand, since the simulator obtained by the data holder is lossy compressed, the model fine-tuned based on the simulator (the model composed of the simulator and the fine-tuned adapter) cannot achieve the same fine-tuning performance as the second large model. The data holder still needs to use the second large model provided by the model provider to obtain better model performance. This performance difference is regarded as an embodiment of the method to ensure the security of the model, which is also an embodiment of the security of the first large model and the second large model.
[0021] The key challenge of cross-domain fine-tuning is how to effectively perform lossy compression of the first large model, that is, how to obtain a simulator corresponding to the first large model, so as to ensure the security of the model while retaining the performance of the model after cross-domain fine-tuning as much as possible.
[0022] The first large model generally consists of multiple network layers, which can be divided into frozen intermediate layers and network layers to be fine-tuned by the data holder (i.e., adapters). In one possible implementation, the process of compressing the model and making the simulator can be to prune multiple network layers from the frozen intermediate layers, take the pruned first large model as the student model, take the first large model that has not been pruned as the teacher model, perform knowledge distillation on the student model by the teacher model, and thus obtain the simulator to be delivered to the data holder. For example, the first two network layers and the last two network layers of the first large model can be kept unchanged, and the frozen intermediate layers (i.e., network layers from the third layer to the last third layer) can be subjected to UniformLayerDrop (uniform layer drop), i.e., pruning according to equal steps, for example, for the 3rd to 10th network layers, only the 3rd, 5th, 7th, and 9th network layers are retained, and the 4th, 6th, 8th, and 10th network layers are pruned. The un-compressed first large model is taken as the teacher model, and the compressed first large model is taken as the student model, and knowledge distillation is performed on the student model. During the knowledge distillation process, multiple iterations of training are required to update a large number of weight parameters in the student model, so that the student model imitates the input-output behavior of the teacher model, and the frozen intermediate layers in the student model that complete the knowledge distillation will be used as the simulator. After the simulator is made, the model provider will deliver the simulator to the data holder, and the data holder will train the first two layers and the last two layers of the first large model as fine-tuned adapters while keeping the weights of the simulator unchanged. After the data holder fine-tunes, the data holder will return the weights of the adapters to the model provider and access the original first large model to obtain a second large model, which is used to perform the downstream task of the data holder.
[0023] In this implementation, a large number of parameters in the pruned first large model need to be updated during multiple iterations of training in the knowledge distillation process to obtain a simulator with required performance, which consumes a large amount of computing resources, storage resources, and time resources.
[0024] In the embodiments of the present specification, a method for generating a simulator of a pre-trained large model is provided. The first pre-trained large model includes a plurality of functional units, a set of compression parameters can be obtained, the set of compression parameters includes a plurality of compression parameters, and the plurality of functional units are compressed based on the compression parameters to obtain a first simulator. A support data set is obtained, and the support data set includes input data and label data. Then, based on the support data set, a compression score corresponding to each compression parameter is determined. The compression score is positively correlated with the gradient difference between the plurality of functional units and the first simulator in the back propagation process, and is negatively correlated with the loss difference between a second large model and a third large model. The second large model includes a plurality of functional units, and the third large model includes the first simulator. Then, based on the compression scores corresponding to the plurality of compression parameters respectively, S compression parameters satisfying a preset score requirement are determined, and the plurality of functional units are compressed using the S compression parameters to generate a second simulator corresponding to the first large model. S is determined based on a parameter compression rate.
[0025] In this way, the compression parameters satisfying the preset score requirement are determined as the index of the compression score, and then the model is compressed through the compression parameters. The performance of the second simulator after compression directly meets the requirements of privacy and practicality in cross-domain fine-tuning, and there is no need to adjust the network parameters of the simulator obtained after compression, that is, there is no need to perform knowledge distillation and other model training, so that a larger model can be compressed under limited resources, and the computing resources and time are greatly saved.
[0026] First, the relationship between the first large model, the second large model and the third large model is described.
[0027] The first pre-trained large model serves as an original model for cross-domain fine-tuning. In the case where the first large model is not compressed, the first large model usually includes a plurality of Transfomer modules stacked in sequence. The Transfomer module can be further decomposed into a multi-head attention (MHA) module and a multi-layer perceptron (MLP) module.
[0028] The multiple functional units (i.e., the frozen intermediate layers in the above) that allow lossy compression can be determined from the multiple Transformer modules included in the first large model. The multiple functional units can belong to part of the multiple Transformer modules to compress the entire Transformer module; or the multiple functional units can be decomposed into the MHA modules and / or MLP modules included in part of the multiple Transformer modules to compress the MLP modules and MHA modules respectively. The same compression method or different compression methods can be used for the MHA modules and MLP modules, such as low-rank decomposition, selective pruning, etc. To enhance the expression ability, the intermediate layer dimension of the MLP module is empirically set to be very high, significantly exceeding the dimensions of its input and output. For the MLP module, the compression method of selective pruning can be preferably selected to reduce the complexity of the model without weakening the performance of the model, thereby achieving efficient compression.
[0029] For the first large model The optimization goal of lossy compression is to find a lossy compression function to compress it into a model with less parameters for fine-tuning by the data holder. For simplicity, the first large model can be represented as wherein, represents the top and low layers in the first large model that can be trained, i.e., the adapter, and the specific number of layers contained in the top and low layers is not limited in the embodiment, for example, may be the top two layers, i.e., the first two layers, may be the bottom two layers, i.e., the last two layers, and ε represents the frozen intermediate layers. The compressed model can be represented as wherein, the second simulator is the adapter is used as the initialization setting before fine-tuning by the data holder. Then the data holder fine-tunes the adapter using to obtain the third large model after fine-tuning The adapter after fine-tuning is returned to the model provider and merged with the frozen intermediate layers to combine into the second large model In addition, in order to better illustrate the optimization goal of the present scheme, it can also be assumed that the fourth large model is The fourth large model is obtained by directly fine-tuning the first large model, wherein is the adapter fine-tuned based on ε.
[0030] The second major model is located on the model provider's side and includes multiple functional units and adapters for providing services to the data holder for downstream tasks. The third major model is the model obtained after fine-tuning by the data holder and includes a second simulator and adapters.
[0031] The optimization objective for obtaining the second simulator corresponding to the first large model is explained below.
[0032] Given a dataset for downstream tasks Including training dataset D train and test dataset D test The overall optimization objective of cross-domain fine-tuning can be used to find To simultaneously satisfy two objectives: (1) Minimize the value in D test middle and Difference in losses between To maintain the performance of the model (2) maximize in D test middle and Difference in losses between This is to ensure the model's safety. Additionally, a prerequisite must be met: during fine-tuning, ensure... In D train Minimize the loss function And ensure In D train Minimize the loss function To ensure the performance of the fine-tuned model. The overall optimization objective can be expressed as:
[0033]
[0034]
[0035] in, It is the given model (the model whose location represents its position) in the dataset. The average task loss is λ, which controls the balance between model safety and model performance during cross-domain fine-tuning. This represents the loss function.
[0036] It is easy to see from formula (1) that the optimization objective needs to be achieved through the dataset. Training is performed to obtain the parameters required for optimization. In order to achieve the above optimization goal without training the model, this embodiment simplifies the above optimization goal from the perspective of model weight parameters and proposes the concept of compression score.
[0037] The compression score can be defined as: positively correlated with the gradient differences between multiple functional units and the simulator during backpropagation, and negatively correlated with the loss differences between the second and third largest models. The optimization objective of minimizing formula (1) can be transformed into minimizing this compression score. Considering the occurrence of gradient and loss differences, it is essentially due to the weight parameters of the simulator. and the weight parameters w of multiple functional units i The weight difference δ between them i This compression score can actually be expressed as the weight difference δ. i A function, where w i This represents the i-th weight parameter among all weight parameters in the first-largest model. Represents the compressed model Among all the equally effective weight parameters, the i-th weight parameter, when the layer containing i belongs to the frozen intermediate layer. otherwise,
[0038] The following explains one derivation process of the compressed fraction.
[0039] First, in order to satisfy objective (1) in formula (1), it is required to be based on the original first large model. Fine-tuning adapter in With the second largest model Based on simulator The adapter obtained by fine-tuning The closer the better, which is equivalent to making the emulator obtained by compression... The closer the gradients transmitted in the compressed fraction are to those transmitted in the multiple functional units ε before compression, the better. That is, the smaller the gradient difference between the multiple functional units and the simulator in the compressed fraction, the better. Therefore, the compressed fraction is positively correlated with the gradient difference between the multiple functional units and the simulator during backpropagation. Then, in order to satisfy objective (2) in formula (1), the loss difference between the second and third largest models in the compressed fraction should be as large as possible. Therefore, the compressed fraction is negatively correlated with the loss difference between the second and third largest models during forward propagation.
[0040] Furthermore, the difference in loss between the second and third largest models is actually the difference in weights δ between multiple functional units and the simulator. i As a result, the loss difference between the second and third largest models is positively correlated with the loss difference between multiple functional units and simulators. In some implementations, the compression score can also be viewed as being positively correlated with the gradient difference between multiple functional units and simulators during backpropagation and negatively correlated with the loss difference between multiple functional units and simulators.
[0041] The method for generating a simulator of a pre-trained model provided in the embodiments of the present specification can be divided into two stages: first, compressing a plurality of functional units using different compression parameters to obtain a first simulator, and calculating the corresponding compression scores, and then selecting S compression parameters that meet the preset score condition from the compression scores, and compressing the plurality of functional units to obtain the required second simulator. The compression parameters are determined based on the compression method used. When different compression parameters are used for lossy compression, the resulting weight difference δ i is different.
[0042] To calculate the compression score, the gradient difference between the plurality of functional units and the simulator during the backpropagation process and the loss difference between the second largest model and the third largest model during the forward propagation process need to be calculated respectively. The embodiments of the present specification use the data in the support data set to perform the backpropagation process to approximate the training process using the data set , obtaining the parameters required to calculate the gradient difference and the loss difference. The support data set includes input data and label data, wherein the input data is used to input the first large model to obtain output data output by the first large model, and the label data is used to calculate the gradient difference and the loss difference with the output data. It can be understood that the closer the support data set is to the distribution of the data set , the better the approximation effect. For each weight difference δ i , the weight parameters of the compressed first simulator can be calculated through the compression parameters, and then the difference between the weight parameters w i of the plurality of functional units and is calculated.
[0043] The method for generating a simulator of a pre-trained large model provided in the embodiments of the present specification is described below. As shown in Figure 2 , the generation process can be performed by any device, platform or device cluster with computing and processing capabilities, including the following steps S201-S204.
[0044] In step S201, a set of compression parameters is obtained, and a plurality of functional units are compressed based on the compression parameters to obtain a first simulator.
[0045] The set of compression parameters includes a plurality of compression parameters, and the set of compression parameters can be determined based on the compression method used. Through the compression parameters, the weight parameters of the first simulator compressed according to the compression method of the plurality of functional units can be calculated.
[0046] When different compression methods and different parameter compression rates are used, different sets of compressed parameters are generated. The embodiment is not limited to the specific compression method and parameter compression rate used, for example, low-rank decomposition, selective pruning, weight quantization, and shared weight can be used for compression, and the compression parameter rate is set according to actual needs.
[0047] The weight parameters of the plurality of functional units can be referred to as a target weight parameter matrix. As an implementation manner, dynamic rank decomposition (DRD) or low-rank decomposition can be performed on the target weight parameter matrix to compress the weight parameters. Specifically, the set of compressed parameters is a set of singular values in the target weight parameter matrix corresponding to the plurality of functional units. This step can be singular value decomposition of the target weight parameter matrix to obtain a singular value decomposition matrix and a set of singular values, and the set of singular values includes a plurality of singular values. For each singular value, the singular value and the singular vector corresponding to the singular value are retained in the singular value decomposition matrix. Based on the singular value and the singular vector corresponding to the singular value, low-rank decomposition is performed on the plurality of functional units to obtain a first emulator.
[0048] Taking the plurality of functional units as an MHA module as an example, for the target weight parameter matrix of the linear layer in the MHA SVD (Singular Value Decomposition) can be performed:
[0049] W=U∑V T (2)
[0050] wherein the elements on the diagonal of the matrix ∑ are a set of singular values in W, the columns in the matrix U and the rows in the matrix V T are left singular vectors and right singular vectors corresponding to each singular value, respectively.
[0051] When low-rank decomposition (i.e., truncated SVD) is performed on W based on a singular value k in the set of singular values, for the singular value k, the singular value k and the left singular vector and the right singular vector corresponding to the singular value k are retained in the singular value decomposition matrix to obtain a compressed first emulator. The weight parameter matrix of the first emulator, referred to as the first weight parameter matrix, is:
[0052]
[0053] The weight parameter matrix of the first emulator is composed of two low-rank matrices, i.e., U :,k ∑ k,k and
[0054] As a further implementation manner, selective channel pruning (SCP) can be performed on the target weight parameter matrix to compress the weight parameters. Specifically, the compression parameter set is a set of intermediate layer dimensions of the upper weight parameter matrix and the lower weight parameter matrix corresponding to the plurality of functional units. In this step, a set of intermediate layer dimensions can be obtained, and the set of intermediate layer dimensions includes a plurality of intermediate layer dimensions. For each intermediate layer dimension, weight parameters outside the weight parameters corresponding to the intermediate layer dimension in the upper weight parameter matrix and the lower weight parameter matrix are pruned to obtain a first simulator after compression.
[0055] Taking the multi-functional module as an MLP module, the MLP module can be represented as:
[0056] f MLP (X)=σ(XW up )W down (4)
[0057] wherein X is the input of the MLP module, the upper weight parameter matrix the lower weight parameter matrix σ represents an activation function, d h represents a hidden layer dimension, d int represents a set of intermediate layer dimensions, i.e., a plurality of intermediate layer dimensions included in the set of intermediate layer dimensions, d h <<d int Therefore, pruning d int can effectively compress the model.
[0058] For each intermediate layer dimension k in d int , weight parameters outside the weight parameters corresponding to the intermediate layer dimension k in the upper weight parameter matrix and the lower weight parameter matrix are pruned, i.e., only the weight parameters corresponding to the intermediate layer dimension k in the upper weight parameter matrix and the lower weight parameter matrix are retained, to obtain a first simulator after compression. The weight parameter matrix of the first simulator, referred to as the first weight parameter matrix, is:
[0059]
[0060] The weight parameter matrix of the first simulator is composed of two low-rank matrices, i.e., and
[0061] For obtaining the compression parameter set of other compression manners, the above-mentioned embodiments can be referred to, and this embodiment does not limit this.
[0062] In step S202, a support data set is obtained.
[0063] wherein the support data set The support data set includes input data and label data, the input data is used to input the first large model, and the label data is a label of output data of the first large model The support data set can be used to perform a forward propagation process and a backward propagation process to approximate a training process using the data set to obtain parameters required for calculating gradient differences and loss differences. It can be understood that the closer the support data set is to the distribution of the data set , the better the approximation effect is.
[0064] The embodiment does not limit the acquisition of the support data set. For example, in order to protect the data privacy of the data holder, the support data set can be acquired from a public general data set. Further, the general data sets are mixed to acquire the support data set, which can increase the generalization performance of the support data set. For another example, with the permission of the data holder, part of the data set is acquired as the support data set, so that the performance of the compressed model is more matched with the downstream task.
[0065] In step S203, based on the support data set, a compression score corresponding to each compression parameter is determined.
[0066] The compression score is positively correlated with the gradient difference between the plurality of functional units and the first simulator in the backward propagation process, and is negatively correlated with the loss difference between the second large model and the third large model. In this step, the gradient difference and the loss difference in the compression score can be calculated respectively. In order to calculate the required parameters, this specification approximates the calculation of the gradient difference and the loss difference, so that the required values, such as the loss function, the first-order partial derivative of the loss function with respect to the weight, and the second-order partial derivative of the loss function with respect to the weight, are obtained using each input data in the support data set, performing a forward propagation and a backward propagation (without updating the parameters of the model) once.
[0067] 1) Gradient difference between the plurality of functional units and the first simulator
[0068] The gradient difference is calculated, that is, the gradient difference after adding δ i to each weight w i of the plurality of functional units becomes The formula for minimizing the gradient difference can be expressed as:
[0069]
[0070] In fact, due to memory limitations and computational costs, it is almost impossible to estimate the gradient change amount through all weights at the same time. According to the chain rule, the gradient change amount can be minimized to simplify the calculation of the gradient of each layer. For the weight parameter wi corresponding gradient change amount When adding δ i , according to Taylor expansion, we can get:
[0071]
[0072] where the value of high-order term is very small and can be ignored, is each weight parameter w i in the compression fraction. The above process obtains the relationship between the gradient difference and the weight difference δ i , so that we can know how much gradient difference will be caused by the change of model weight without training the model.
[0073] The embodiment is not limited to the specific calculation method of . In an example, is a partial Hessian matrix of a complete model, and the Hessian matrix is the second derivative matrix of the loss function with respect to the model weight parameter. Direct calculation requires a large amount of calculation, in order to reduce the calculation complexity, the calculation complexity is reduced from o(n 3 ) to o(n 2 ), the Hessian matrix can be decomposed into multiple blocks according to the linear number of layers of the model, and then the corresponding Kronecker product is calculated for each block to approximate the Hessian matrix of the layer. The parameters required in this process can be obtained from the back propagation process based on the support dataset, so that the large matrix is decomposed into the product of small matrices, reducing the memory and calculation requirements. In yet another example, the second derivative of the loss function with respect to the weight parameter can be directly calculated .
[0074] 2) the loss difference between the second largest model and the third largest model
[0075] When calculating the loss difference, the loss of the second largest model and the loss of the third largest model can be calculated based on the support dataset, and the loss difference can be calculated. Considering the sum of the product of each derivative and the small change of the corresponding weight variable, the increment of the loss function at a given point can be estimated, and can be expanded in the form of total differential:
[0076]
[0077] where n1 to n2 are the corresponding network layers in the first largest model, w i is the weight parameter in the n1 to n2 layers (i.e., the weight parameter in multiple functional units), and ⊙ represents the inner product.
[0078] By defining total differential, a loss function can be established. and The connection between them, dw i It can be seen as w i and The weight difference δ between i For differentiable functions with small variables The following equations were established:
[0079]
[0080] The above process yields the loss difference and weight difference δ. i By understanding this relationship, we can determine the difference in loss caused by changes in model weights without training the model. This embodiment applies to... The specific calculation method is not restricted. In one example, it can be to directly calculate the partial derivative of the loss function with respect to the weights.
[0081] Thus, two formulas are obtained: formula (7) predicts the gradient difference using the weight difference, and formula (9) predicts the loss difference using the weight difference. The goal of cross-domain fine-tuning is to find a weight difference that minimizes the gradient difference while maximizing the loss difference, which can be expressed by the following formula for the compression score:
[0082]
[0083] Where Δ is an arbitrarily small variable, and n represents w i λ represents the weight parameters in layers n1 to n2 (i.e., the weight parameters in multiple functional units). A larger λ indicates stronger privacy protection, while a smaller λ indicates better model performance. This compression can be viewed as gradient-preserving compression, and the compression score can also be called the gradient-preserving compression score (GCS).
[0084] For each compressed first weight parameter in multiple functional units Its corresponding compression score can be expressed as:
[0085]
[0086] As one implementation method, when determining the compression score corresponding to each compression parameter based on the supporting dataset, this step can involve calculating each weight for each compression parameter. The corresponding compression score is calculated, and then the compression scores corresponding to all weights are summed to obtain the compression score corresponding to the compression parameter.
[0087] As one implementation method, when determining the compression score corresponding to each compression parameter based on the support dataset, this step specifically involves first performing a forward propagation. The input data is fed into the first large model to obtain its output data. Based on the difference between the output data and the label data, the loss of the first large model is determined. This embodiment does not impose restrictions on the loss function used to determine the loss. Then, a backpropagation is performed, based on the loss... Determine the gradient of each weight parameter in multiple functional units during backpropagation. and losses Second-order partial derivatives with respect to the weight parameters Based on the first weight parameter and the weight parameters w of multiple functional units i The weight difference δ between i , and second-order partial derivative product Determine the gradient differences between multiple functional units and the first simulator. Based on weight difference δ i and gradient dot product between Determine the loss difference between the second and third largest models. Finally, a weighted summation based on gradient differences and loss differences. The compression score corresponding to the compression parameters is obtained, and the weight corresponding to the loss difference is negative.
[0088] In this way, for each compression parameter k, we can first obtain the first weight parameter of the first simulator, and then calculate the compression score corresponding to the compression parameter k by combining the weight difference between the first weight parameter and the weight parameters of multiple functional units, combined with the gradient difference and loss difference.
[0089] In step S204, based on the compression scores corresponding to multiple compression parameters, S compression parameters that meet the preset score requirements are determined, and the S compression parameters are used to compress multiple functional units to generate a second simulator corresponding to the first large model.
[0090] Where S is determined based on the parameter compression ratio and the compression method used. When compressing multiple functional units, the number of network parameters, i.e., the number of weight parameters, will be reduced. The parameter compression ratio is used to represent the ratio of the compressed model size to the original model size, i.e., the ratio of the number of weight parameters in the first simulator generated after compression to the number of weight parameters in multiple functional units. The preset score requirement can be a requirement to minimize the compression score as much as possible, or it can be a threshold for the compression score. Compression scores below this threshold are considered to meet the preset score requirement. For example, in order to meet the preset score requirement, the selected S compression parameters can be the S compression parameters with the lowest corresponding compression scores, or any S compression scores from the M compression parameters with the lowest corresponding compression scores, where M is greater than S.
[0091] The following is combined with Figure 3 This example illustrates how to compress the MLP and MHA modules in the transform module using different compression methods. In other implementations, the MLP and MHA modules can be compressed using the same method, such as using low-rank decomposition.
[0092] As one implementation method, when using low-rank decomposition for compression, compressing the target weight parameter matrix W can be viewed as performing truncated SVD. That is, S singular values need to be selected to truncate the target weight parameter matrix using SVD, thus decomposing a large matrix with a large number of weight parameters into two low-rank matrices with a small number of weight parameters. For multiple functional units, such as the MHA module, the target weight parameter matrix is... Based on the given parameter compression ratio r mha The calculation method for S can be as follows: first calculate the compression ratio r. mha The number of rows d in the target weight parameter matrix o With column number d i The product value is then used to calculate the number of rows d in the target weight parameter matrix. o With column number d i The sum of the products is then used to determine the value of S, which is the ratio of the product to the sum.
[0093]
[0094] Then, select the S singular values corresponding to the S lowest compression scores. In the singular value decomposition matrix, retain the S singular values and their corresponding singular vectors. Based on the S singular values and their corresponding singular vectors, perform low-rank decomposition on multiple functional units to obtain the second simulator. The weight parameter matrix of the second simulator is then determined. Given two low-rank matrices B and A, the following formula is used:
[0095]
[0096] in, s is a set of rank indices containing S elements. A rank index is the index position in a matrix corresponding to a singular value k. ∑ s,s Contains the selected S singular values, U :,s and V s,: It contains singular vectors corresponding to S singular values. Additionally, to ensure δ... i To make the compressed score more accurate, we can choose the top 5% of the rank, which is the top 5% of the singular values in ∑.
[0097] As one implementation method, when using selective pruning compression, the compression of the target weight parameter matrix W can be viewed as selecting S compression parameters (intermediate layer dimensions) for pruning to reduce the number of intermediate layer dimensions, thereby reducing the number of weight parameters. For multiple functional units, such as the target weight parameter matrix corresponding to the MLP module, it can be expressed as formula (4), based on the given parameter compression ratio r. mlp With r mha They can be the same or different. S can be calculated using the compression ratio r. mlp and intermediate layer dimension d int The product of the number of channels gives the value of S, i.e.:
[0098] S = r mlp d int (13)
[0099] Then, select the S intermediate dimensions corresponding to the S lowest compressed scores. For each of the S intermediate dimensions (denoted as set s), trim the S intermediate dimensions k in the weight parameter matrix. and lower weight parameter matrix The weight parameters other than the corresponding weight parameters are used to obtain the compressed second simulator. The weight parameter matrix of the compressed second simulator is shown in the following formula:
[0100]
[0101] The weight parameter matrix of the second simulator consists of two low-rank matrices, namely and
[0102] The compressed simulators corresponding to the MLP and MHA modules together form the second simulator corresponding to the first large model.
[0103] The method in the above embodiments determines the compression parameters that meet the preset score requirement based on the compression score, and then compresses the model based on the compression parameters. The performance of the second simulator after compression directly meets the requirements for privacy and practicality in cross-domain fine-tuning, and there is no need to adjust the network parameters of the simulator obtained after compression, that is, there is no need to perform knowledge distillation and other model training. Therefore, a larger model can be compressed under limited resources, which greatly saves computing resources and time. Compared with the training-based method, less than 10% of the computing power and time is required to compress the simulator with the same effect.
[0104] Based on the same concept as the foregoing method embodiments, the present specification embodiments also provide a device for generating a simulator of a pre-trained large model. The device can be applied to any device, platform or device cluster with computing and processing capabilities.
[0105] As shown in Figure 4 The device includes:
[0106] The parameter acquisition module 401 is configured to acquire a compression parameter set, the compression parameter set including a plurality of compression parameters, and compress a plurality of functional units based on the compression parameters to obtain a first simulator. The data set acquisition module 402 is configured to acquire a support data set, the support data set including input data and label data. The score determination module 403 is configured to determine a compression score corresponding to each compression parameter based on the support data set. The compression score is positively correlated with a gradient difference between the plurality of functional units and the first simulator in a back propagation process, and is negatively correlated with a loss difference between a second large model and a third large model. The second large model includes the plurality of functional units, and the third large model includes the first simulator. The input data in the support data set is used to input a first large model to obtain output data output by the first large model, and the label data is used to calculate the gradient difference and the loss difference with the output data. The compression module 404 is configured to determine S compression parameters that meet a preset score requirement based on the compression scores corresponding to the plurality of compression parameters respectively, and compress the plurality of functional units using the S compression parameters to generate a second simulator corresponding to the first large model. S is determined based on a parameter compression rate.
[0107] The present specification embodiments also provide a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed in a computer, the computer executes the method for generating a simulator of a pre-trained large model provided in the foregoing embodiments.
[0108] The present specification embodiments also provide a computing device including a memory and a processor. The memory has a computer program / instruction stored therein. When the processor executes the computer program / instruction, the method for generating a simulator of a pre-trained large model provided in the foregoing embodiments is implemented.
[0109] In the 1990s, it was quite obvious to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structures of diodes, transistors, switches, etc.) or in software (improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flow into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented by hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it, rather than by asking a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented by "logic compiler" software, which is similar to software compilers used in program development, and the original code before compilation is also written in a specific programming language, which is called a hardware description language (HDL), and there are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is quite easy to obtain hardware circuits implementing the logical method flow by only logically programming the method flow in the above-mentioned hardware description languages and programming it into an integrated circuit.
[0110] The controller can be implemented in any suitable way, for example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. The skilled person will also appreciate that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the controller in the form of logic gates, switches, an application specific integrated circuit, a programmable logic controller and an embedded microcontroller, etc. to perform the same functions by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein to perform the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means to perform the various functions can be considered to be both a software module implementing the method and a structure within a hardware component.
[0111] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present application does not rule out that with the development of future computer technologies, computers implementing the functions of the above embodiments can be personal computers, laptop computers, vehicle human-computer interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or combinations of any of these devices.
[0112] Although the method operations of the embodiments of the present disclosure are described in a particular, sequential order, one or more of the method operations can be omitted, or the method operations can be performed in an order other than the described order. Additionally, one or more of the method operations can be performed concurrently, or with partial concurrence. Furthermore, one or more of the method operations can be performed by different entities, or over different time periods. The term "including" as used herein is intended to mean "comprising," such that the process, method, article, or apparatus that includes elements in addition to those specified. As used in this description, the term "coupled" means a direct or indirect connection, which can be physical or logical. The term "coupled" does not relate to a direct connection or wiring.
[0113] For the sake of description, the above-described apparatus is described as various modules to describe the functions of the apparatus. Of course, when implementing one or more embodiments of the present disclosure, the functions of the modules can be implemented in one or more software and / or hardware, or the modules implementing the same functions can be combined into a plurality of sub-modules or sub-units. The apparatus embodiments described above are merely illustrative, for example, the division of the units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0114] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The term "coupled" means a direct or indirect connection, which can be physical or logical. Figure 1 The term "coupled" means a direct or indirect connection, which can be physical or logical.
[0115] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0117] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0118] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory. The memory is an example of computer-readable media.
[0119] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0120] Those skilled in the art will appreciate that the one or more embodiments described herein can be provided as a method, a system or a computer program product. Accordingly, the one or more embodiments described herein can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the one or more embodiments described herein can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable code.
[0121] The one or more embodiments described herein can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The one or more embodiments described herein can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.
[0122] The various embodiments described in this specification are described in the context of progressive embodiments, with each embodiment building on the previous one. The same or similar parts between embodiments are cross-referenced as appropriate. Each embodiment focuses on the differences between that embodiment and the previous one. In particular, the system embodiments are described relatively simply, as they are substantially similar to the method embodiments. In the description of the specification, the use of the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the particular feature, structure, material or characteristic being described is included in at least one embodiment or example of the specification. Illustrative descriptions of the above terms do not necessarily refer to the same embodiment or example in this specification. Moreover, the particular features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples. Furthermore, different embodiments or examples described in this specification can be combined and combined with features of different embodiments or examples, without contradiction, by those skilled in the art.
[0123] The above description merely provides examples of the one or more embodiments described in this specification and does not limit the one or more embodiments described in this specification. The one or more embodiments described in this specification can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the one or more embodiments described in this specification should be included in the scope of the claims.
Claims
1. A method for generating an emulator for a pre-trained large model, a first pre-trained large model comprising a plurality of functional units, the method comprising: obtaining a compression parameter set comprising a plurality of compression parameters, and compressing the plurality of functional units based on the compression parameters to obtain a first emulator; obtaining a support data set comprising input data and label data; based on the support data set, determining a compression score corresponding to each of the compression parameters, the compression score being positively correlated with a gradient difference between the plurality of functional units and the first emulator in a back propagation process, and being negatively correlated with a loss difference between a second large model and a third large model, the second large model comprising the plurality of functional units, the third large model comprising the first emulator, the input data in the support data set being used to input the first large model to obtain output data output by the first large model, and the label data being used to calculate the gradient difference and the loss difference with the output data; based on the compression scores corresponding to the plurality of compression parameters respectively, determining S compression parameters satisfying a preset score requirement, and compressing the plurality of functional units using the S compression parameters to generate a second emulator corresponding to the first large model, the S being determined based on a parameter compression rate.
2. The method of claim 1, wherein, The compression parameter set is a singular value set in a target weight parameter matrix corresponding to the plurality of functional units, and the obtaining a compression parameter set and compressing the plurality of functional units based on the compression parameters to obtain a first emulator comprises: performing singular value decomposition on the target weight parameter matrix to obtain a singular value decomposition matrix and the singular value set, the singular value set comprising a plurality of singular values; for each singular value, retaining the singular value and a singular vector corresponding to the singular value in the singular value decomposition matrix; based on the singular value and the singular vector corresponding to the singular value, performing low-rank decomposition on the plurality of functional units to obtain a first emulator.
3. The method of claim 2, wherein, The using the S compression parameters to compress the plurality of functional units to generate a second emulator corresponding to the first large model comprises: retaining the S singular values and singular vectors corresponding to the S singular values in the singular value decomposition matrix; based on the S singular values and singular vectors corresponding to the S singular values, performing low-rank decomposition on the plurality of functional units to obtain the second emulator, a weight parameter matrix of the second emulator being two low-rank matrices.
4. The method of claim 2, wherein, The S being determined based on a parameter compression rate comprises: calculating a product value of the parameter compression rate, a row number and a column number of the target weight parameter matrix; calculating a sum value of the row number and the column number of the target weight parameter matrix; determining a ratio of the product value to the sum value as a value of the S.
5. The method of claim 1, wherein, The compression parameter set is an intermediate layer dimension set in an upper weight parameter matrix and a lower weight parameter matrix corresponding to the plurality of functional units, and the obtaining a compression parameter set and compressing the plurality of functional units based on the compression parameters to obtain a first emulator comprises: obtaining the intermediate layer dimension set, the intermediate layer dimension set comprising a plurality of intermediate layer dimensions; for each of the intermediate layer dimensions, pruning weight parameters of the intermediate layer dimension outside weight parameters corresponding to the upper weight parameter matrix and the lower weight parameter matrix to obtain a compressed first simulator.
6. The method of claim 5, wherein, The using the S compressed parameters to compress the plurality of functional units to generate a second simulator corresponding to the first large model comprises: for the S intermediate layer dimensions, pruning weight parameters of the S intermediate layer dimensions outside weight parameters corresponding to the upper weight parameter matrix and the lower weight parameter matrix to obtain the compressed second simulator.
7. The method of claim 5, wherein, The S is determined based on a parameter compression rate, comprising: calculating a product value of the parameter compression rate and a channel number of the intermediate layer dimension to obtain a value of the S.
8. The method of claim 1, wherein, The first large model comprises a plurality of Transfomer modules stacked in sequence, and the plurality of functional units are multi-head attention (MHA) modules and / or multi-layer perceptron (MLP) modules in the plurality of Transfomer modules.
9. The method of claim 1, wherein, The first simulator has first weight parameters, and the determining a compression score corresponding to each of the compressed parameters based on the support data set comprises: inputting the input data into the first large model to obtain output data output by the first large model, and determining a loss of the first large model based on a difference between the output data and the label data; based on the loss, determining a gradient of each weight parameter in the plurality of functional units in a back propagation process and a second-order partial derivative of the loss with respect to the weight parameter; based on a weight difference between the first weight parameters and the weight parameters of the plurality of functional units and a product of the second-order partial derivative, determining a gradient difference between the plurality of functional units and the first simulator; based on a dot product between the weight difference and the gradient, determining a loss difference between the second large model and the third large model; based on a weighted sum of the gradient difference and the loss difference, obtaining the compression score corresponding to the compressed parameter, and a weight value corresponding to the loss difference is a negative value.
10. A computing device comprising a memory and a processor, the memory storing executable code, and the processor implementing the method of any one of claims 1-9 when executing the executable code.