Large model miniaturization method based on parameter sharing and knowledge distillation

Through the method of parameter sharing and knowledge distillation, sparse coding and improved KL divergence calculations can achieve miniaturization of large models, solving the problem of large models deploying on resource-constrained devices, and achieving efficient compression and performance maintenance.

CN120258087APending Publication Date: 2025-07-04JIANGSU JIYUAN MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345060.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing large-scale pre-trained models are difficult to deploy on resource-constrained devices, making them difficult to achieve miniaturization and maintain high performance.

Method used

Using parameter sharing and knowledge distillation methods, a miniaturized student model is constructed through teacher model feature extraction, sparse coding and knowledge distillation loss optimization, and combined with sparse autoencoder and Hellinger distance improved KL divergence calculation to achieve model compression and performance maintenance.

Benefits of technology

Effectively compress large models into small models while maintaining high performance. It is suitable for various types of deep learning models, with flexibility and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258087A_ABST
    Figure CN120258087A_ABST
Patent Text Reader

Abstract

The invention discloses a large model miniaturization method based on parameter sharing and knowledge distillation. The method comprises the following steps: firstly, preparing data: preparing a training data set and a test data set; selecting a teacher model: selecting a pre-trained large model as an initial model of the teacher model and the student model; building a student model: building a student model with less parameter quantity; and finally, knowledge distillation training: performing knowledge distillation training on the student model by using the teacher model, and minimizing a loss function. According to the method, the large model can be effectively compressed into the small model, meanwhile, the high performance is kept, and the wide application prospect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for miniaturizing large models based on parameter sharing and knowledge distillation. Background Art

[0002] In recent years, with the rapid development of deep learning technology, large-scale pre-trained models (such as GPT-3, BERT, etc.) have achieved remarkable results in the fields of natural language processing, computer vision, etc. However, these models usually have a huge number of parameters and computational complexity, making it difficult to deploy them on resource-constrained devices (such as mobile devices, embedded devices). Therefore, how to miniaturize large models while maintaining their performance has become an important research topic. Summary of the Invention

[0003] Object of the Invention: The present invention proposes a method for miniaturizing large models based on parameter sharing and knowledge distillation, aiming to solve the above problems existing in the prior art, which can effectively compress large models into small models while maintaining high performance and has broad application prospects.

[0004] Technical Solution Adopted by the Present Invention: A method for miniaturizing large models based on parameter sharing and knowledge distillation, comprising the following steps: Step 1: Prepare a training data set and a test data set, use a pre-trained large model as the teacher model, and select a student model with fewer parameters; Step 2: Calculate the reconstruction error; Feature extraction of the teacher model: Input the data into the teacher model, and extract feature representations at specific layers of the teacher model ; these features contain the teacher model's high-level semantic understanding of the data; Feature extraction of the student model: Similarly, input the data into the student model, and extract feature representations at the corresponding layers of the teacher model ; the reconstruction error can be expressed as ; Step 3: Calculate the difference in model feature interaction; Calculate the covariance matrix of the teacher model features and the covariance matrix of the student model features , the covariance matrix can reflect the linear correlation between features; by calculating the Frobenius norm of the covariance matrix, the interaction relationship between features in the student model and the teacher model can be expressed; the difference in model interaction relationship can be expressed as , F represents calculating the Frobenius norm of the matrix; Step 4: Calculate the parameter sharing loss; To further reduce the model scale, sparse coding is used to learn the sparse representation of features; the goal of the sparse autoencoder is to force the activation values of the neurons in the hidden layer to be as sparse as possible; is the activation value of the neurons in the hidden layer of the student model, and minimize , where i represents the serial number of the neurons in the hidden layer, to achieve further miniaturization of the model; Define as the feature interaction retention factor, as the sparsity penalty coefficient; then the model parameter sharing loss is: ; Step 5: Calculate the knowledge distillation loss; By minimizing the difference between the outputs of the teacher model and the student model, the knowledge of the teacher model is transferred to the student model; the knowledge distillation loss is denoted as ; Define to represent the KL divergence between the outputs of the teacher model and the student model, where T and S represent the outputs of the teacher model and the student model respectively; define to represent the Hellinger distance between the outputs of the teacher model and the student model; Since the KL divergence itself is asymmetric, using the KL divergence to measure the distance between samples, due to its asymmetry, it may make the clustering results complex and difficult to interpret, because the distance between samples will be different depending on the calculation direction; for two similar probability distributions, their differences may not be significantly reflected in the calculation of the Hellinger distance, resulting in difficulty in accurately distinguishing such small differences; to address the above problems, the following constructed function is used to combine the characteristics of both to obtain better results: , represents the KL divergence between the outputs of the student model and the teacher model; The reason for better results is this part makes the metric have a certain symmetry, and then adding , using the symmetry and value range characteristics of the Hellinger distance, further enhances the measurement of the difference between probability distributions, while avoiding the disadvantage that the Hellinger distance itself is difficult to distinguish small differences; this combination method is not a simple weighting, but comprehensively considers the differences between the two distributions from different perspectives; Step 6: Combine parameter sharing and knowledge distillation to jointly optimize the student model; ; During the training process, the constraints brought by knowledge distillation loss and parameter sharing are considered at the same time, so that the student model has smaller parameters and computational complexity while maintaining performance; is the weight of the output difference.

[0005] Furthermore, the specific layer of the teacher model in step 2 is an intermediate hidden layer.

[0006] The present invention first prepares data: prepares training data sets and test data sets; then performs teacher model selection: selects a pre-trained large model as the initial model of the teacher model and the student model; then performs student model construction: constructs a student model with fewer parameters; and finally performs knowledge distillation training: uses the teacher model to perform knowledge distillation training on the student model to minimize the loss function.

[0007] Beneficial effects of the present invention: 1. Efficiency: By optimizing the internal feature representation and structural differences of the model, parameter sharing is achieved. By introducing a sparse mechanism, the parameter scale of the student model is further reduced. The large model is effectively compressed into a small model while maintaining high performance.

[0008] 2. General: This method can be applied to various types of deep learning models.

[0009] 3. Flexible: You can flexibly select the strategies of parameter sharing and knowledge distillation according to actual needs to achieve the best compression effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flowchart of the large model miniaturization method based on parameter sharing and knowledge distillation of the present invention. DETAILED DESCRIPTION

[0011] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0012] like Figure 1 As shown, a large model miniaturization method based on parameter sharing and knowledge distillation includes the following steps: Step 1: Prepare training and test datasets, use the pre-trained large model as the teacher model, and select a student model with fewer parameters; Step 2: Calculate the reconstruction error; Teacher model feature extraction: Input data is passed to the teacher model, and feature representation is extracted at a specific layer (usually the middle hidden layer) of the teacher model. ;These features contain the teacher model’s high-level semantic understanding of the data; Student model feature extraction: similarly pass the input data into the student model and extract feature representation at the layer corresponding to the teacher model ; The reconstruction error can be expressed as ; Step 3: Calculate the difference in model feature interactions; Calculate the covariance matrix of the teacher model features and the covariance matrix of the student model features . The covariance matrix can reflect the linear correlation between features; by calculating the Frobenius norm of the covariance matrix, the interaction relationship between features in the student model and the teacher model can be expressed; the difference in the model interaction relationship can be expressed as , where F represents the calculation of the Frobenius norm of the matrix; Step 4: Calculate the parameter sharing loss; To further reduce the model size, sparse coding is used to learn the sparse representation of features; the goal of the sparse autoencoder is to force the activation values of the hidden layer neurons to be as sparse as possible; is the activation value of the hidden layer neurons of the student model, and minimizing , where i represents the serial number of the hidden layer neurons, realizes the further miniaturization of the model; Define as the feature interaction preservation factor, and as the sparsity penalty coefficient; then the model parameter sharing loss is: Step 5: Calculate the knowledge distillation loss; By minimizing the difference between the outputs of the teacher model and the student model, the knowledge of the teacher model is transferred to the student model; the knowledge distillation loss is denoted as ; Define to represent the KL divergence between the outputs of the teacher model and the student model, where T and S represent the outputs of the teacher model and the student model respectively; define to represent the Hellinger distance between the outputs of the teacher model and the student model; Because the KL divergence itself is asymmetric, using the KL divergence to measure the distance between samples, due to its asymmetry, the clustering results may become complex and difficult to interpret, because the distance between samples will vary depending on the calculation direction; for two similar probability distributions, their differences may not be significantly reflected in the calculation of the Hellinger distance, resulting in difficulty in accurately distinguishing such small differences; to address the above problems, the following constructed function is used to combine the characteristics of both to obtain better results: , represents the KL divergence between the outputs of the student model and the teacher model; The reason for the better effect is This part makes the metric have a certain symmetry, and then adds , making use of the symmetry and value range characteristics of the Hellinger distance to further enhance the measurement of the difference in probability distributions, while avoiding the shortcoming that the Hellinger distance itself is difficult to distinguish small differences; This combination method is not a simple weighting, but comprehensively considers the differences between the two distributions from different perspectives; Step 6: Combine parameter sharing and knowledge distillation to jointly optimize the student model; ; During the training process, considering both the knowledge distillation loss and the constraints brought by parameter sharing, the student model has a smaller number of parameters and computational complexity while maintaining performance; is the weight of the output difference.

[0013] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A method for miniaturizing large models based on parameter sharing and knowledge distillation, characterized in that: It includes the following steps: Step 1: Prepare the training dataset and the test dataset, use a pre-trained large model as the teacher model, and select a student model with fewer parameters; Step 2: Calculate the reconstruction error; Teacher model feature extraction: Input the data into the teacher model and extract the feature representation at specific layers of the teacher model ; These features contain the teacher model's advanced semantic understanding of the data; Student model feature extraction: Similarly, the input data is passed into the student model, and feature representations are extracted at the layers corresponding to the teacher model ; The reconstruction error is expressed as ; Step 3: Calculate the difference in model feature interactions; Calculate the covariance matrix of the teacher model features and the covariance matrix of the student model features , where the covariance matrix reflects the linear correlation between features; by calculating the Frobenius norm of the covariance matrix, the interaction relationship between features in the student model and the teacher model is expressed; the difference in the model interaction relationship is expressed as , where F represents the calculation of the Frobenius norm of the matrix; Step 4: Calculate the parameter sharing loss; To further reduce the model size, use sparse coding to learn the sparse representation of features; The goal of the sparse autoencoder is to force the activation values of the neurons in the hidden layer to be as sparse as possible; is the activation value of the neurons in the hidden layer of the student model, minimizing , where i represents the serial number of the neurons in the hidden layer, to achieve further miniaturization of the model; Definition is the feature interaction retention factor, is the sparsity penalty coefficient; then the model parameter sharing loss is: ; Step 5: Calculate the knowledge distillation loss; Transfer the knowledge of the teacher model to the student model by minimizing the difference between the outputs of the teacher model and the student model; the knowledge distillation loss is denoted as ; Definition Denote the KL divergence between the outputs of the teacher model and the student model, where T and S represent the outputs of the teacher model and the student model respectively; define Denote the Hellinger distance between the outputs of the teacher model and the student model; Adopt a function constructed as follows to combine the characteristics of both to obtain better results: , represents the KL divergence between the outputs of the student model and the teacher model; Step 6: Combine parameter sharing and knowledge distillation, and jointly optimize the student model; ; During the training process, the knowledge distillation loss and the constraints brought by parameter sharing are considered simultaneously, enabling the student model to have a smaller number of parameters and computational complexity while maintaining performance; is the weight of the output difference.

2. A method for miniaturizing a large model based on parameter sharing and knowledge distillation according to claim 1, characterized in that: The specific layer of the teacher model in Step 2 is the intermediate hidden layer.