Iterative pruning optimization method based on dynamic double-teacher multi-stage distillation

By adopting a dynamic dual-teacher multi-stage distillation method during iterative pruning, the problem of difficult model performance recovery after pruning is solved, and more efficient knowledge distillation and performance recovery are achieved.

CN119940454APending Publication Date: 2025-05-06GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510127151.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During iterative pruning, the capacity difference between the fixed teacher model and the student model is too large, making it difficult for the pruning model to effectively restore performance.

Method used

A multi-level distillation method based on dynamic dual teachers is adopted. In the fine-tuning stage, a teacher model with a capacity similar to the student model is selected as the regular teacher by introducing multi-level distillation and dynamic teacher selection mechanisms, and a multi-level distillation fine-tuning is performed using the original model as the associate teacher.

Benefits of technology

The model performance recovery ability after pruning is significantly improved, the capacity gap between the teacher model and the student model is reduced, and the efficiency of knowledge distillation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940454A_ABST
    Figure CN119940454A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses an iterative pruning optimization method based on dynamic double-teacher multistage distillation, comprising the following steps: S1, training a deep learning model as an original model; s2, pruning the model by using a preset filter importance evaluation standard to obtain a pruned student model; s3, selecting a teacher model with the capacity similar to that of the student model from a teacher model warehouse as a primary teacher, and selecting the original model as a secondary teacher; s4, performing multi-stage distillation fine adjustment on the student model by using the primary teacher and the secondary teacher; and S5, repeating the steps of pruning, teacher selection and multi-stage distillation fine adjustment until a preset maximum pruning rate is reached. An iterative pruning optimization method based on dynamic double teachers is provided, and the performance recovery capability of the pruned model is remarkably improved by introducing multi-stage distillation and a dynamic teacher selection mechanism in a fine adjustment stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation. Background Art

[0002] With the rapid development of deep learning technology, the size and computational complexity of the model have gradually become bottlenecks that limit its widespread application. In order to solve this problem, model compression technology has emerged. In particular, iterative pruning technology can significantly reduce the size of the model while maintaining performance by repeatedly removing unimportant neurons and connections in the model. However, although iterative pruning can effectively reduce the computational complexity of the model, the model performance is usually greatly affected during the pruning process, especially in the fine-tuning stage. How to effectively restore the performance lost after pruning has become an urgent problem to be solved.

[0003] Traditional fine-tuning methods usually use supervised learning based on hard labels to restore model performance, that is, by training the pruned student model with real labels. However, this method has several significant defects. First, the supervision information provided by hard label training is relatively single, and it is difficult to effectively guide the learning of the student model after pruning, especially during the pruning process, the capacity of the student model gradually decreases, making model training more difficult. Secondly, the gap between the capacity of the teacher model and the student model gradually increases. Traditional knowledge distillation methods usually use a fixed original teacher model for guidance. As pruning progresses, the capacity of the student model continues to shrink. This capacity difference between the fixed teacher and the constantly changing student model will lead to poor knowledge distillation effects. Therefore, how to overcome these problems in the fine-tuning process after pruning and improve the performance recovery ability of the student model is the main challenge facing current model compression technology.

[0004] Knowledge distillation, as a technology that transfers the knowledge of the teacher model to the student model, can make up for this shortcoming to a certain extent. The distillation process enables the student model to learn richer feature representations from the teacher model when the capacity is small. However, traditional knowledge distillation methods still face some problems, especially in the scenario of iterative pruning. The fixed original teacher model often cannot provide sufficient guidance for the student model whose capacity is greatly reduced after pruning, resulting in unsatisfactory performance recovery of the student model. In addition, how to choose a suitable teacher model is also a major problem in knowledge distillation. As a teacher, although the original model can provide global knowledge, its distillation effect is often not ideal when the capacity difference is too large.

[0005] In order to solve these problems, the dynamic dual-teacher mechanism has gradually become a research hotspot in recent years. The core idea of ​​the dynamic dual-teacher mechanism is to introduce multiple teacher models and improve the effect of knowledge transfer by dynamically selecting a teacher model that is closer to the capacity of the student model. In particular, the original model is used as a deputy teacher and the intermediate model is used as a main teacher to jointly guide the learning of the student model. This method can not only provide more dimensional knowledge information, but also better adapt to changes in the pruning process.

[0006] In general, although the existing technical solutions can reduce the model size during the iterative pruning process, due to the large capacity difference between the fixed teacher model and the student model during the fine-tuning process, it is difficult for the pruned model to effectively restore its performance. Summary of the invention

[0007] The present invention provides an iterative pruning optimization method based on dynamic dual-teachers multi-stage distillation, and proposes an iterative pruning optimization method based on dynamic dual-teachers. In the fine-tuning stage, multi-stage distillation and dynamic teacher selection mechanism are introduced to significantly improve the performance recovery ability of the model after pruning.

[0008] The present invention provides an iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation, comprising the following steps:

[0009] S1: Train a deep learning model as the original model;

[0010] S2: Prune the model using the preset filter importance evaluation criteria to obtain the pruned student model;

[0011] S3: Select a teacher model with a capacity close to that of the student model from the teacher model warehouse as the main teacher, and select the original model as the deputy teacher;

[0012] S4: Perform multi-level distillation fine-tuning on the student model using the main teacher and the assistant teacher;

[0013] S5: Repeat the above steps of pruning, teacher selection, and multi-level distillation fine-tuning until the preset maximum pruning rate is reached.

[0014] Preferably, in step S2, the pruning of the model using a preset filter importance evaluation standard includes: the filter importance evaluation standard uses L1-norm to screen the filters in the convolution layer to determine the importance of each filter and perform pruning.

[0015] Preferably, in step S3, after each round of pruning and fine-tuning, the currently obtained pruned and fine-tuned model is added to the teacher model warehouse so that the subsequent pruning rounds select the model as the positive teacher model.

[0016] Preferably, the conditions for selecting a positive teacher are as follows: the ratio between the capacity of the student model and the capacity of the positive teacher model is not less than a preset threshold, and the teacher model with the best performance within the threshold range is selected as the positive teacher.

[0017] Preferably, in step S4, multi-stage distillation fine-tuning includes intermediate feature distillation and soft label distillation, and the cross entropy loss of the true label is simultaneously introduced to constitute the total loss of the multi-stage distillation fine-tuning, and the student model is trained; wherein, the intermediate features of the shallow network are mainly learned from the main teacher, and the deep unpruned layers learn the intermediate feature knowledge from both the main teacher and the assistant teacher.

[0018] Preferably, the intermediate feature distillation is based on the attention mechanism, and an attention loss function is used to guide the student model; wherein the attention loss function is specifically:

[0019]

[0020] in, is the jth vectorized attention feature of the student model; and They are the j-th vectorized attention feature of the original model and the i-th vectorized attention feature of the intermediate teacher model; and denote the activation tensors of the student and teacher networks respectively, and p denotes the norm type.

[0021] Preferably, the soft label distillation is based on the output distillation technology of the temperature parameter to align the output distribution of the teacher model with the output distribution of the student model, specifically including:

[0022] The softened Softmax function is used to obtain the output probability distribution of the teacher model: where the Softmax function is specifically:

[0023]

[0024] Among them, z S is the logical output of the student network; Z S is the softened output of the student network; T is the parameter temperature, which controls the softening degree of the output result;

[0025] Calculate the output probability distribution of the student model at the same temperature and construct the soft label distillation loss based on the KL divergence; specifically:

[0026]

[0027] L Multi_KD =k1L Dis (S,T1)+k2LDis (S,T2)

[0028] Among them, Z S and are the softened outputs of the student network and the teacher network respectively; T1 and T2 are the main teacher and the assistant teacher respectively; k1 and k2 represent the weight proportion of each loss, and k1+k2=1.

[0029] Preferably, the cross entropy loss of the real label is introduced as follows: each student model simultaneously learns the real label, and uses the cross entropy function to obtain the cross entropy loss for the model to learn the given data set; wherein the cross entropy function is specifically:

[0030] L CE =CrossEntorpy(Z s ,y true )

[0031] Among them, Z s is the result of passing the student network logic output through softmax; y true is the true label.

[0032] Preferably, the total loss function of the multi-stage distillation fine-tuning is composed of attention loss, soft label distillation loss and cross entropy loss weighted, specifically:

[0033] L Total =αL CE +βL Multi_AT +γL Multi_KD

[0034] Among them, α, β and γ represent the weight proportion of each loss, and α+β+γ=1.

[0035] Preferably, in step S5, when the cumulative pruning rate reaches the preset maximum pruning rate, or the degree of model performance recovery after a round of multi-stage distillation fine-tuning does not reach the expected threshold, the iteration is stopped and the final pruning model is output, and the final pruning model has both less model capacity and better performance.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] The present invention discloses an iterative pruning optimization method based on dynamic dual-teachers multi-stage distillation, and proposes an iterative pruning optimization method based on dynamic dual-teachers. In the fine-tuning stage, by introducing multi-stage distillation and dynamic teacher selection mechanism, the performance recovery ability of the model after pruning is significantly improved, the capacity gap between the teacher model and the student model is reduced, and the efficiency of knowledge distillation is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1It is a flow chart of an iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation provided by an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram comparing the present application provided by an embodiment of the present invention with a traditional distillation learning fine-tuning method;

[0040] Figure 3 It is a schematic diagram of the dynamic dual-teacher multi-level knowledge distillation process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0042] like Figure 1 As shown, the present application provides an iterative pruning optimization method based on dynamic dual-teacher multi-level distillation, comprising the following steps:

[0043] S1: Train a deep learning model as the original model;

[0044] S2: Prune the model using the preset filter importance evaluation criteria to obtain the pruned student model;

[0045] S3: Select a teacher model with a capacity close to that of the student model from the teacher model warehouse as the main teacher, and select the original model as the deputy teacher;

[0046] S4: Perform multi-level distillation fine-tuning on the student model using the main teacher and the assistant teacher;

[0047] S5: Repeat the above steps of pruning, teacher selection, and multi-level distillation fine-tuning until the preset maximum pruning rate is reached.

[0048] Preferably, in step S2, the pruning of the model using a preset filter importance evaluation standard includes: the filter importance evaluation standard uses L1-norm to screen the filters in the convolution layer to determine the importance of each filter and perform pruning.

[0049] For the kth filter in the convolutional layer, its importance is calculated by the L1 norm:

[0050]

[0051] Among them, C in is the number of input channels, Kh ×K w is the convolution kernel size, W k,i,j Represents the weight value of the k-th filter at the i-th input channel and the j-th spatial position.

[0052] In the above scheme, the original model is structurally pruned to remove redundant or unimportant filters in the model, thereby reducing the size and computational complexity of the model; the key lies in the "preset filter importance evaluation standard", which is used to measure the importance of each filter in the model and determine which filters can be safely removed with minimal impact on model performance. Common filter importance evaluation standards include but are not limited to L1 norm, L2 norm, gradient-based methods, etc. According to the preset pruning rate, filters with lower importance scores are removed from the original model to obtain a streamlined model, namely the "pruned student model". The student model inherits the network structure of the original model, but some filters have been removed, and its parameter quantity and computational complexity have been reduced.

[0053] Preferably, in step S3, after each round of pruning and fine-tuning, the currently obtained pruned and fine-tuned model is added to the teacher model warehouse so that the subsequent pruning rounds select the model as the positive teacher model.

[0054] Preferably, the conditions for selecting a positive teacher are as follows: the ratio between the capacity of the student model and the capacity of the positive teacher model is not less than a preset threshold, and the teacher model with the best performance within the threshold range is selected as the positive teacher.

[0055] In the above scheme, the concept of dynamic dual teachers is introduced. Specifically, the teacher model warehouse stores multiple intermediate models generated during the iterative pruning process. These intermediate models are models that have undergone previous rounds of pruning and fine-tuning, and they represent different stages of the model compression process. Selecting the main teacher includes: selecting an intermediate model with a capacity similar to the current "student model after pruning" from the teacher model warehouse as the "main teacher". "Similar capacity" means that the number of parameters or the amount of calculation of the main teacher model is within a preset reasonable range with the student model. The purpose of selecting an intermediate model with similar capacity as the main teacher is to reduce the capacity gap between the teacher model and the student model, making the knowledge distillation process more effective. Because when the capacity gap is too large, the complex knowledge contained in the teacher model may be difficult to be effectively absorbed by the student model with a smaller capacity. Selecting the deputy teacher includes: directly selecting the "original model" obtained through training as the "deputy teacher". The original model has the best performance and the most complete knowledge, and can provide global guidance information for the student model. By simultaneously selecting an intermediate model with similar capacity as the main teacher and an original model with excellent performance as the deputy teacher, a dynamic dual teacher structure is constructed to guide the learning of the student model from different angles and levels.

[0056] In one embodiment provided in the present application, Figure 2 As shown, given a CNN with L convolutional layers, we call T a =(C1,C2,,...,C L ) is the original network, where C1 is the number of channels in the first layer, and the model triplet (A, c, p) is defined, where A is the model, c is the size of model A, and p is the performance of model A. That is, the model triplet of the original model is: (T a ,c a ,p a ); Initialize the teacher warehouse Ts = {}, the teacher list TL = {}. Taking advantage of the characteristics of iterative pruning, as pruning proceeds, the model after each round of pruning and fine-tuning is added to the teacher model warehouse as an intermediate model, providing an alternative for the subsequent selection of the positive teacher model, and together with the original model, it constitutes a dual-teacher structure until the student model learns. As pruning proceeds, the number of parameters and the amount of calculation of the pruned model gradually decrease. Each time, the highest performance model within the size gap with the student model is selected from the teacher warehouse and added to the teacher list as the teacher model for this fine-tuning. It not only has a more friendly model capacity than the original model, but also has a highly similar model structure, making its knowledge easier to be absorbed by the student model. In addition, its inference cost is also lower than the original model. The model selection formula is as follows:

[0057] T ’ s ={(A,c,p)|(A,c,p)∈T s ,C s ≤c≤100·C s / gap}

[0058]

[0059] The teacher warehouse is Ts, which will continuously add new teacher models as pruning progresses. New teacher models are added to the teacher warehouse in the form of (A, c, p). The capacity of the student model (pruning model) is Cs, and the capacity range of the teacher model accepted by the student model is [C s ,100·C s / gap], indicating that the ratio of the student model capacity to the teacher model capacity is not less than gap%.

[0060] As mentioned in Multi-Teacher Knowledge Distillation, extracting different knowledge from different teacher models can provide more diverse and rich information for the student model, which helps to improve its performance. The selection formula of dynamic dual teachers is as follows:

[0061] T′ L =T L ∪(Ta ,c a ,p a )

[0062] When iterative pruning just starts and the first round of teacher selection begins, the teacher warehouse is empty. At this time, only the original model is used as the teacher model. At this time, the traditional single-teacher distillation method is used for distillation learning. The original model is used as the deputy teacher, and the main teacher is selected from the model warehouse. The two constitute a dynamic dual teacher.

[0063] Preferably, in step S4, multi-stage distillation fine-tuning includes intermediate feature distillation and soft label distillation, and the cross entropy loss of the true label is simultaneously introduced to constitute the total loss of the multi-stage distillation fine-tuning, and the student model is trained; wherein, the intermediate features of the shallow network are mainly learned from the main teacher, and the deep unpruned layers learn the intermediate feature knowledge from both the main teacher and the assistant teacher.

[0064] In the above scheme, if Figure 3 As shown in the figure, in the process of knowledge distillation, deep intermediate feature knowledge easily leads to over-normalization of the student model, while shallow intermediate feature knowledge cannot play a guiding role. At the beginning of pruning, the number of pruned blocks is small and the capacity gap is not large. Using the intermediate knowledge of the original model as a guide, the performance can also be quickly restored. However, as the number of pruned blocks increases, the capacity gap between each block becomes larger, and the student model cannot effectively learn the teacher model, resulting in poor performance recovery. Therefore, the learning of intermediate feature knowledge selects blocks with similar capacity, and the student model is easier to understand intermediate features with similar capacity, thereby avoiding information loss and compensating for the performance loss due to pruning. Taking advantage of the characteristics of iterative pruning, the original teacher and the intermediate model generated during the pruning process are used as dual teachers to avoid the cost of manually selecting the teacher model, and the student model and the dual teacher model are highly similar in structure, which is more convenient for knowledge transfer. When pruning layer by layer, the pruned blocks and the blocks being pruned are guided by the main teacher model, while the blocks that have not been pruned are jointly guided by the main teacher and the deputy teacher, and the blocks that have not been pruned are jointly guided by the dual teachers. This avoids large capacity deviations between blocks due to pruning, making it easier to transfer intermediate feature knowledge and achieve performance recovery.

[0065] Preferably, the intermediate feature distillation is based on the attention mechanism, and an attention loss function is used to guide the student model; wherein the attention loss function is specifically:

[0066]

[0067] in, is the jth vectorized attention feature of the student model; and They are the j-th vectorized attention feature of the original model and the i-th vectorized attention feature of the intermediate teacher model; and denote the activation tensors of the student and teacher networks respectively, and p denotes the norm type.

[0068] Preferably, the soft label distillation is based on the output distillation technology of the temperature parameter to align the output distribution of the teacher model with the output distribution of the student model, specifically including:

[0069] The softened Softmax function is used to obtain the output probability distribution of the teacher model: where the Softmax function is specifically:

[0070]

[0071] Among them, z S is the logical output of the student network; Z S is the softened output of the student network; T is the parameter temperature, which controls the softening degree of the output result;

[0072] Calculate the output probability distribution of the student model at the same temperature and construct the soft label distillation loss based on the KL divergence; specifically:

[0073]

[0074] L Multi_KD =k1L Dis (S,T1)+k2L Dis (S,T2)

[0075] Among them, Z S and are the softened outputs of the student network and the teacher network respectively; T1 and T2 are the main teacher and the assistant teacher respectively; k1 and k2 represent the weight proportion of each loss, and k1+k2=1.

[0076] Preferably, the cross entropy loss of the real label is introduced as follows: each student model simultaneously learns the real label, and uses the cross entropy function to obtain the cross entropy loss for the model to learn the given data set; wherein the cross entropy function is specifically:

[0077] L CE =CrossEntorpy(Z s ,y true )

[0078] Among them, Z s is the result of passing the student network logic output through softmax; y true is the true label.

[0079] Preferably, the total loss function of the multi-stage distillation fine-tuning is composed of attention loss, soft label distillation loss and cross entropy loss weighted, specifically:

[0080] L Total =αL CE +βL Multi_AT +γL Multi_KD

[0081] Among them, α, β and γ represent the weight proportion of each loss, and α+β+γ=1.

[0082] Preferably, in step S5, when the cumulative pruning rate reaches the preset maximum pruning rate, or the degree of model performance recovery after a round of multi-stage distillation fine-tuning does not reach the expected threshold, the iteration is stopped and the final pruning model is output, and the final pruning model has both less model capacity and better performance.

[0083] In the above scheme, after completing a round of pruning and fine-tuning, if the preset maximum pruning rate has not been reached, the fine-tuned student model will be used as a new model to be pruned; each round of iteration will further reduce the size of the model, and use dynamic dual-teacher multi-level distillation to restore performance; by iterative pruning and fine-tuning, the model can be compressed step by step and finely, avoiding the sharp drop in performance that may be caused by a one-time large-scale pruning; each round of fine-tuning uses a dynamically selected teacher model for knowledge distillation to ensure that while the model is compressed, the performance of the model is maintained or even improved as much as possible.

[0084] The global cumulative pruning rate after t rounds of iteration:

[0085]

[0086] Among them, PruneRate k is the pruning rate in the kth round.

[0087] Model performance recovery after t rounds of iteration:

[0088]

[0089] Among them, AccRecovery t is the performance recovery rate of the student model after the tth iteration.

[0090] The iteration stops when any of the following conditions are met:

[0091] GlobalPruneRate t ≥MaxPruneRate

[0092] AccRecovery t ≤MinAccRecovery

[0093] Among them, MaxPruneRate is the target pruning rate, and MinAccRecovery is the minimum acceptable performance recovery rate.

[0094] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. An iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation, characterized in that: The following steps are involved: S1: Train a deep learning model as the original model; S2: Prune the model using the preset filter importance evaluation criteria to obtain the pruned student model; S3: Select a teacher model with a capacity close to that of the student model from the teacher model warehouse as the main teacher, and select the original model as the deputy teacher; S4: Perform multi-level distillation fine-tuning on the student model using the main teacher and the assistant teacher; S5: Repeat the above steps of pruning, teacher selection, and multi-level distillation fine-tuning until the preset maximum pruning rate is reached.

2. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 1 is characterized in that: In step S2, the pruning of the model using a preset filter importance evaluation standard includes: the filter importance evaluation standard uses L1-norm to screen the filters in the convolution layer to determine the importance of each filter and perform pruning.

3. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 1 is characterized in that: In step S3, after each round of pruning and fine-tuning, the currently obtained pruned and fine-tuned model is added to the teacher model warehouse so that the subsequent pruning rounds select the model as the positive teacher model.

4. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 3 is characterized in that: The specific conditions for selecting the positive teacher are: the ratio between the student model capacity and the positive teacher model capacity is not less than a preset threshold, and the teacher model with the best performance is selected as the positive teacher within this threshold range.

5. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 4 is characterized in that: In step S4, multi-level distillation fine-tuning includes intermediate feature distillation and soft label distillation, and simultaneously introduces the cross entropy loss of the true label to constitute the total loss of multi-level distillation fine-tuning to train the student model; wherein, the intermediate features of the shallow network are mainly learned from the main teacher, and the deep unpruned layers learn the intermediate feature knowledge from both the main teacher and the assistant teacher.

6. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 5, characterized in that: The intermediate feature distillation is based on the attention mechanism and uses the attention loss function to guide the student model; wherein the attention loss function is specifically: in, is the jth vectorized attention feature of the student model; and They are the j-th vectorized attention feature of the original model and the i-th vectorized attention feature of the intermediate teacher model; and denote the activation tensors of the student and teacher networks respectively, and p denotes the norm type.

7. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 5, characterized in that: The soft label distillation is based on the output distillation technology of the temperature parameter to align the output distribution of the teacher model with the student model, specifically including: The softened Softmax function is used to obtain the output probability distribution of the teacher model: where the Softmax function is specifically: Among them, z S is the logical output of the student network; Z S is the softened output of the student network; T is the parameter temperature, which controls the softening degree of the output result; Calculate the output probability distribution of the student model at the same temperature and construct the soft label distillation loss based on the KL divergence; specifically: L Multi_KD =k1L Dis (S,T1)+k2L Dis (S,T2) Among them, Z S and are the softened outputs of the student network and the teacher network respectively; T1 and T2 are the main teacher and the assistant teacher respectively; k1 and k2 represent the weight proportion of each loss, and k1+k2=1.

8. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 5, characterized in that: The cross entropy loss of the real label is introduced as follows: each student model simultaneously learns the real label, and uses the cross entropy function to obtain the cross entropy loss for the model to learn the given data set; wherein the cross entropy function is specifically: L CE =CrossEntorpy(Z s ,y true ) Among them, Z s is the result of passing the student network logic output through softmax; y true is the true label.

9. An iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to any one of claims 6 to 8, characterized in that: The total loss function of the multi-level distillation fine-tuning is composed of attention loss, soft label distillation loss and cross entropy loss weighted, specifically: L Total =αL CE +βL Multi_AT +γL Multi_KD Among them, α, β and β represent the weight proportion of each loss, and α+β+β=1.

10. The iterative pruning optimization method based on dynamic dual-teacher multi-stage distillation according to claim 9, characterized in that: In step S5, when the cumulative pruning rate reaches the preset maximum pruning rate, or the model performance recovery degree after a round of multi-stage distillation fine-tuning does not reach the expected threshold, the iteration is stopped and the final pruning model is output. The final pruning model has both less model capacity and better performance.

Citation Information

Cited By

  • Text error correction method and device, electronic equipment and storage medium

    CN120297267A

  • Text error correction method, device, electronic device and storage medium

    CN120297267B