Knowledge distillation method based on reverse feature fusion and classifier multiplexing

By introducing a decoupled attention projection module and an inverse feature fusion module in knowledge distillation and reusing the teacher classifier, the problem of large gap in feature expression between teachers and students' models is solved, and the prediction performance of the student model and the proximity of the output distribution is improved.

CN120197669APending Publication Date: 2025-06-24CHONGQING UNIV OF TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082092.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In knowledge distillation, there is a big gap in the expression of intermediate features between the teacher model and the student model, which makes it difficult for the student network to fully learn the implicit knowledge of the intermediate features of the teacher network.

Method used

The knowledge distillation method based on reverse feature fusion and classifier multiplexing is adopted. The student model is trained by introducing a decoupled attention projection module and an inverse feature fusion module in the student network, and the teacher classifier and feature distance loss function are multiplexed.

Benefits of technology

It effectively shortens the differences in the characteristics between teachers and students, enables the student network to learn the characteristics of the intermediate layer of the teacher network more fully, improves the prediction performance of the student model, and makes the output distribution of the student model closer to the teacher model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197669A_ABST
    Figure CN120197669A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge distillation method based on reverse feature fusion and classifier multiplexing, and relates to the technical field of computers. The method at least comprises the following steps: S1, obtaining an original data set, and dividing the original data set into a training set and a test set; s2, preprocessing the original data set; s3, pre-training the teacher network model by using the training set, and storing the trained teacher network model; s4, pre-training the weight by freezing the teacher network, introducing a decoupling attention projection module and a reverse feature fusion module into the student network, and training a student model by multiplexing a teacher classifier and a feature distance loss function; and S5, in a reasoning stage, retaining parameters of a student network architecture and a teacher classifier layer. Through the design of the method, the problem that the student network is difficult to fully learn the implicit knowledge of the teacher network interlayer feature due to the large expression difference between the teacher and the student can be solved. And through multiplexing a teacher classifier, output distribution closer to teachers is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and specifically to a knowledge distillation method based on reverse feature fusion and classifier reuse. Background Art

[0002] In recent years, the rapid development of deep neural networks (DNNs) has greatly promoted the progress of the field of computer vision and has been widely applied in task scenarios such as image classification, object detection, and image segmentation. However, the powerful network performance is usually accompanied by a large number of parameters and high computational costs, making it difficult to deploy on devices with limited resources. Therefore, how to reduce the computational cost and ensure the model performance has become a challenging problem. Knowledge distillation provides an effective solution by transferring the knowledge of a pre-trained large model to a lightweight model. Existing knowledge distillation methods are roughly divided into offline knowledge distillation and online knowledge distillation. Offline knowledge distillation adopts a two-stage training method: first, a large teacher model is pre-trained, and then the extracted knowledge is transferred to a smaller student model to help the student learn the complex knowledge in the teacher model. Online knowledge distillation, on the other hand, adopts a single-stage training method, directly optimizing the target model by continuously updating knowledge during the training process, enabling the student model to make full use of the rich information from multiple outputs.

[0003] In the knowledge distillation method, there is a large gap in capabilities between the teacher model and the student model trained from scratch, resulting in a large gap in the intermediate feature expressions between the teacher and the student, making it difficult for the student network to fully learn the implicit knowledge of the intermediate layer features of the teacher network.

[0004] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a knowledge distillation method based on reverse feature fusion and classifier reuse to overcome the problem of a large gap in the intermediate feature expressions between the teacher and the student in knowledge distillation, which makes it difficult for the student network to fully learn the implicit knowledge of the intermediate layer features of the teacher network.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A knowledge distillation method based on reverse feature fusion and classifier reuse, at least including the following steps:

[0007] S1: Obtain the original data set and divide it into a training set and a test set;

[0008] S2: Preprocess the original data set;

[0009] S3: Use the training set to pre-train the teacher network model and save the trained teacher network model;

[0010] S4: Freeze the pre-trained weights of the teacher network, introduce a decoupled attention projection module and an inverse feature fusion module in the student network, and reuse the teacher classifier and the feature distance loss function to train the student model;

[0011] S5: In the inference stage, retain the parameters of the student network architecture and the teacher classifier layer.

[0012] Furthermore, the S4 at least includes the following steps:

[0013] S4.1: Freeze the pre-trained weights of the teacher network model;

[0014] S4.2: Build a decoupled attention projection module at each stage of the student network to match the feature dimensions of each stage of the student with those of the teacher classifier;

[0015] S4.3: Use the inverse feature fusion module for the student network model to inversely fuse the features of each stage of the deep and shallow layers layer by layer;

[0016] S4.4: Reuse the teacher classifier to train on the fused features, and use the feature distance loss function as the distillation loss to train the student model.

[0017] Furthermore, the decoupled attention projection module includes a decoupled attention module and a projection module.

[0018] Furthermore, the decoupled attention module is encapsulated into a basic unit. For the original feature F of the b-th layer b Feed it into two paths for processing. One path is used to perform average pooling on each channel within the height spatial range, and the other path is used to perform average pooling on each channel within the width spatial range;

[0019] Then, apply a one-dimensional convolutional layer to each of the two paths to enhance the original feature information in the decoupled horizontal and vertical directions;

[0020] Subsequently, use group normalization to reduce the interference between the two channels, enhance the position information in the feature space dimension, and reduce the number of parameters;

[0021] The original feature F b The horizontal attention y h Score and the vertical attention y w The decoupled expressions of the scores are:

[0022] y h = σ(G n (H h (F b )))

[0023] y w = σ(G n (Hw (F b )))

[0024] Among them, H h (·) and H w (·) respectively represent the one-dimensional convolutional layers of two paths, G n (·) is used to represent group normalization, σ(·) is used to represent the non-linear activation function Sigmoid, and the attention scores obtained from the two paths are respectively multiplied by the feature F b .

[0025] The feature after passing through the decoupled attention module The output expression is:

[0026]

[0027] Furthermore, the projection module is encapsulated into a basic unit, and the projection module includes two 1×1 and one adaptive convolutional layer;

[0028] The two 1×1 convolutional layers are used to adjust the number of channels to be the same as that of the deep feature, and the adaptive convolutional layer adaptively adjusts the convolutional kernel and stride according to the size of the F b own feature space, so that the student spatial resolution size matches the teacher classifier. The feature after projection is used for subsequent inverse feature fusion and teacher classifier reuse;

[0029] The feature after passing through the projection module The output expression is:

[0030]

[0031] Design the projection function Proj(·) to convert the features of each stage of the student after passing through the decoupled attention module to the same spatial size as the teacher's deep feature, where:

[0032]

[0033] where W ∈ R m×d is the weighted matrix with a spatial size of m×d.

[0034] Furthermore, the S4.3 at least includes the following steps:

[0035] Fuse the projected feature of the b-th layer and the fusion feature of the next layer ;

[0036] Then, the fused features are fed into two paths for processing. One path contains two convolutional layers with a kernel size of 1×1 for capturing local information, and the other path captures global information through an additional global pooling layer;

[0037] The feature information captured by the two paths is activated through the Sigmoid function to obtain weight scores, and they are respectively multiplied with the features to adjust the importance of the deep features and shallow features using the parameter w, and finally the fused features at this level are obtained

[0038] The deep feature map with a total network level of B is used as the fused feature for the first input, that is

[0039] The feature expression of the subsequent reverse fusion output is:

[0040]

[0041] where D(·) is the fusion function from the feature fusion module, and the final fused feature is expressed as

[0042] Furthermore, the S4.4 at least includes the following steps:

[0043] Reuse the classifier layer parameters of the pre-trained teacher for the fused features, enabling the student model to perform inference by directly reusing the pre-trained teacher classifier, and making the feature alignment loss the only source for generating gradients, where F B t is the deep feature map of the teacher model with a total network level of B;

[0044] Construct a feature distance loss function:

[0045]

[0046] Furthermore, the specific parameters of retaining the student network architecture and the teacher classifier layer in the S5 during the inference stage are:

[0047] During the student inference stage, the teacher network model only retains the classifier layer parameters and retains the student network architecture part, with almost no additional cost increase.

[0048] Compared with the prior art, the beneficial effects of the present invention are:

[0049] Through the design of the method, the present invention can overcome the problem that the large difference in the intermediate feature expressions between the teacher and the student makes it difficult for the student network to fully learn the implicit knowledge of the intermediate layer features of the teacher network. And by reusing the teacher classifier, a more teacher-like output distribution is obtained. Brief Description of the Drawings

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of an embodiment of the present invention;

[0052] Figure 2 It is a framework diagram of the overall knowledge distillation method of an embodiment of the present invention;

[0053] Figure 3 It is a decoupled attention projection module diagram of an embodiment of the present invention;

[0054] Figure 4 It is a reverse feature fusion module diagram of an embodiment of the present invention;

[0055] Figure 5 It is a curve graph of the influence of the teacher classifier and label joint guidance on the performance of the invention embodiment;

[0056] Figure 6 It is a heat map of the visualization analysis of the student fusion features of an embodiment of the present invention. Specific embodiments

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0058] Please refer to Figure 1 , a knowledge distillation method based on reverse feature fusion and classifier reuse includes at least the following steps:

[0059] Step 1, obtain the original data set and divide it into a training set and a test set;

[0060] In this embodiment, the obtained original data set is the CIFAR-100 data set, which contains 60,000 RGB images of 100 categories with a size of 32×32. The divided training set and test set are 50k and 10k images respectively.

[0061] Step 2, preprocess the original data set;

[0062] In this embodiment, all images in the original data set are processed and normalized through the channel mean and standard deviation, and standard data augmentation operations are adopted.

[0063] Step 3: Use the training set to pre-train the teacher model and save the trained teacher model;

[0064] In this embodiment, the teacher model includes resnet110×2, resnet32×4, and WRN-40-2. The experimental parameters for training the teacher model are set as follows: BatchSize is 64, and Epoch is 240. The initial learning rate is set to 0.05, and the gradient learning rate method is used for all. The learning rate is divided by 10 at Epoch 150, 180, and 210. The optimizer is selected as Stochastic Gradient Descent, the weight decay is 5e-4, and the momentum is 0.9.

[0065] Step 4: Freeze the pre-trained weights of the teacher network model, and build a decoupled projection module at each stage of the student network; use the reverse feature fusion module to fuse the features of each stage layer by layer; reuse the teacher classifier for the fused features, and train the student model using the feature distance loss function;

[0066] In this embodiment, the student model includes resnet, WRN, ShuffleNet, and VGG. Among the experimental parameters for training the student model, the initial learning rate for ShuffleNet and MobileNet is 0.02, and the initial learning rate for other models is 0.05. Other parameter settings are the same as those for training the teacher model.

[0067] Step 4 includes:

[0068] Step 4-1: Freeze the pre-trained weights of the teacher network model;

[0069] Step 4-2: Build a decoupled attention projection module at each stage of the student network to match the feature dimensions of the student's features at each stage with those of the teacher classifier;

[0070] Step 4-3: Use the reverse feature fusion module for the student network model to reversely fuse the features of the deep and shallow stages layer by layer;

[0071] Step 4-4: Reuse the teacher classifier for the fused features, and train the student model using the feature distance loss function;

[0072] Specifically, during the process of training the student model, freeze the pre-trained weights of the teacher model obtained in Step 3 to ensure that the weights of the teacher model do not change during the training process.

[0073] Secondly, build a decoupled attention projection module at each stage of the student network. Build a decoupled attention projection module at each stage of the student network model, which includes a decoupled attention module and a projection module. Among them, the decoupled attention module is encapsulated into a basic unit. For the original feature F of the b-th layer bSent to two paths for processing: one path is used to perform average pooling on each channel within the height spatial range. The other path is used to perform average pooling on each channel within the width spatial range. Then, a one-dimensional convolutional layer is applied to each path respectively to enhance the original feature information in the decoupled horizontal and vertical directions. Subsequently, group normalization is used to reduce the interference between the two channels, enhancing the position information in the feature space dimension and reducing the number of parameters. The original feature F b The horizontal attention y h score and the vertical attention y w score are decoupled and expressed as:

[0074] y h = σ(G n (H h (F b )))

[0075] y w = σ(G n (H w (F b )))

[0076] where H(·) is used to represent the one-dimensional convolutional layers of the two paths respectively, G n (·) is used to represent group normalization, and σ(·) is used to represent the non-linear activation function Sigmoid, multiplying the attention scores obtained from the two paths with the feature F b respectively. The feature output after passing through the decoupled attention module is expressed as:

[0077]

[0078] The projection module is encapsulated into a basic unit, which contains two 1×1 and one adaptive convolutional layer. Among them, the two 1×1 convolutional layers are used to adjust the number of channels to be the same as that of the deep features. The adaptive convolutional layer adaptively adjusts the convolutional kernel and stride according to the size of the F b own feature space, so that the student spatial resolution size matches that of the teacher classifier. The feature after projection is used for subsequent inverse feature fusion and teacher classifier reuse. The feature output after passing through the projection module is expressed as:

[0079]

[0080] The projection function Proj(·) is designed to transform the features of each stage of the student after passing through the decoupled attention module to the same spatial size as the teacher's deep features. Among them where W ∈ R m×d is a weighted matrix with a spatial size of m×d.

[0081] Therefore, compared with the feature projector constructed by directly using a simple 1×1 convolution kernel or linear method to achieve the feature dimension matching between the student features and the teacher classifier, the addition of the decoupled attention module can enhance the original feature expression ability by focusing on the key spatial channel information of the features, improve the efficiency of the projection module, and achieve more accurate alignment with the teacher classifier. The decoupled attention module decouples the original features into feature vectors in the horizontal and vertical directions in the spatial dimension, respectively capturing the rich spatial information of the feature map in two directions, and more efficiently enhancing the important areas of the features at each stage. The unique characteristics of the feature maps in the two directions are merged to improve the accuracy of attention prediction in each direction, and achieve the enhancement of the original feature expression ability. The projection module can adaptively change the convolution step and the size of the convolution kernel according to the spatial resolution of the original features at each stage through the adaptive convolution layer, ensuring that each feature can be accurately projected to the same spatial size as the teacher's deep features. The original features with enhanced expression ability further improve the alignment efficiency of the projection module and achieve efficient matching with the teacher classifier. Promote the efficient fusion of multi-stage features of different scales and semantics.

[0082] For example Figure 2 As shown in the reverse feature fusion module, the projected features of the bth layer are first And the fusion features of the next level The fused features are sent to two paths for processing: one path contains two convolution layers with a convolution kernel of 1×1 to capture local information; the other path captures global information through an additional global pooling layer. The feature information captured by the two paths is activated by the Sigmoid function to obtain weight scores, and then respectively compared with the feature Multiply them together and use the parameter w to adjust the importance of deep features and shallow features, and finally get the fusion features of this level

[0083] The deep feature map with a total network level of B is used as the fusion feature of the first input, that is, The expression of the reverse feature fusion module is:

[0084]

[0085] Therefore, the designed inverse feature fusion module realizes the fusion of shallow texture and deep abstract features, integrates the feature information of adjacent stages, builds a bridge for the flow of information from deep to shallow layers, enriches student feature information, enhances feature expression, thereby shortening the feature difference between teachers and students, and enabling teachers to transfer knowledge to students more efficiently.

[0086] And again Figure 2As shown, the teacher classifier is reused for the fused features, and the student model is trained using the feature distance loss function, specifically as follows:

[0087] The classifier layer parameters of the pre-trained teacher are reused for the fused features, enabling the student model to perform inference by directly reusing the pre-trained teacher classifier, and making the feature alignment loss the only source for generating gradients. A feature distance loss function is constructed:

[0088]

[0089] Therefore, the designed method of reusing the teacher classifier and the feature distance loss function enables the student model to imitate the teacher's representation space by reusing the powerful discriminative ability of the teacher model's classifier, better learn the representative knowledge of the teacher, and improve the prediction performance of the student. The more informative fused features obtained through the reverse feature fusion module can improve the reuse efficiency and enhance the consistency between the teacher and student models compared to relying only on single-layer features, further shortening the performance gap.

[0090] Step 5, in the inference stage, retain the parameters of the student network architecture and the teacher classifier layer.

[0091] In the student inference stage. The teacher network model only retains the classifier layer parameters and retains part of the student network architecture, with almost no additional cost increase.

[0092] To verify the effectiveness of the proposed knowledge distillation method, the experimental results are shown in Tables 1 and 2.

[0093] Table 1 Experimental results of comparing knowledge distillation methods on CIFAR-100 under the same teacher-student architecture

[0094]

[0095] Table 2 Experimental results of comparing knowledge distillation methods on CIFAR-100 under different teacher-student architectures

[0096]

[0097]

[0098] The experimental results in Table 1 show that the knowledge distillation method of reverse feature fusion and classifier reuse of the present invention shows extensive effectiveness. Compared with the baseline student trained using cross-entropy in Vanilla KD, the accuracy of the vast majority of teacher-student combinations is increased by at least 4% or more. Compared with the SOTA method, the best performance is achieved when the teacher and student use isomorphic and heterogeneous architectures. For example, in the combination of ResNet-32x4 as the teacher model and ResNet-8x4 as the student model, it is 0.73% higher than the SOTA method SimKD

[0099] Next, ablation experiments were conducted on the attention module, where -OA indicates that the features were projected without passing through the decoupled attention module. The experiments demonstrated that the decoupled attention module effectively captured rich spatial information in the horizontal and vertical directions, enhanced the expression ability of the original input features, achieved efficient matching with the teacher classifier, and improved the alignment efficiency of the projection module. Higher accuracy was obtained, proving the feasibility of the present invention.

[0100] Table 3 Comparison of accuracies of RCRKD before and after using the decoupled attention module on CIFAR-100

[0101]

[0102] Then, ablation experiments were conducted on the reverse feature fusion module and classifier reuse. Among them, +RFF indicates that the features only passed through reverse fusion, and +CR indicates that the student network model reused the teacher classifier. The experimental results showed that the student model benefited from the information blending of shallow and deep features, significantly shortened the intermediate feature gap between the teacher and the student, and the more informative fused features could improve the classifier reuse efficiency. Both could effectively improve the recognition effect of the student model, achieving higher accuracy and proving the feasibility of the present invention.

[0103] Table 4 Comparison of accuracies of RCRKD before and after reusing the classifier and using the RFF module on CIFAR-100

[0104]

[0105] As Figure 5 shown, it demonstrates the impact of the joint guidance of the teacher classifier and label information on the performance of the student model under different weight combinations. Under all settings, the student's performance was far inferior to that of the reverse feature fusion and classifier reuse methods when receiving label guidance while also receiving the teacher classifier, that is, similar performance to the teacher could be achieved only by reusing the teacher classifier. This indicates that different types of supervision signals provided by the teacher classifier and the label conflicted, making it difficult for knowledge to be transferred to the student model in a joint training manner. The experimental results show that the student model can significantly bridge the performance gap between the teacher and the student by directly reusing the teacher classifier without the need to additionally receive real labels and soft labels provided by the teacher.

[0106] As Figure 6 shown, to further evaluate the effectiveness of the reverse feature fusion method, we used heatmap visualization to highlight the regions considered important for predicting the corresponding labels. As Figure 3As shown, it illustrates that the proposed reverse feature fusion module successfully guides the regions related to the details of the objects of attention in each stage of the student network to be closer to the teacher, and the fused features successfully combine the attentional spatial information of the deep abstract features and the shallow texture features, highlighting the effect of the reverse feature fusion module in enhancing the feature expression ability of the student and reducing the difference between the teacher and the student. While the student baseline model sometimes regards the correct region as the background and at the same time focuses on spatially adjacent objects.

[0107] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A knowledge distillation method based on reverse feature fusion and classifier reuse, characterized by: At least the following steps are included: S1: Get the original data set and divide it into training set and test set; S2: Preprocess the original data set; S3: Use the training set to pre-train the teacher network model and save the trained teacher network model; S4: By freezing the pre-trained weights of the teacher network, the decoupled attention projection module and the inverse feature fusion module are introduced into the student network, and the teacher classifier and feature distance loss function are reused to train the student model; S5: During the inference phase, the parameters of the student network architecture and the teacher classifier layer are retained.

2. The knowledge distillation method based on inverse feature fusion and classifier reuse according to claim 1, characterized in that: The S4 at least comprises the following steps: S4.1: Freeze the pre-trained weights of the teacher network model; S4.2: Build a decoupled attention projection module at each stage of the student network to match the characteristics of each stage of the student with the feature dimensions of the teacher classifier; S4.3: Use the reverse feature fusion module on the student network model to reversely fuse the deep and shallow features layer by layer; S4.4: Reuse the teacher classifier to the fusion features for training, and use the feature distance loss function as the distillation loss to train the student model.

3. The knowledge distillation method based on inverse feature fusion and classifier reuse according to claim 2, characterized in that: The decoupled attention projection module includes a decoupled attention module and a projection module.

4. The knowledge distillation method based on reverse feature fusion and classifier reuse according to claim 3 is characterized by: The decoupled attention module is encapsulated into a basic unit, for the original feature F of the bth layer b It is sent to two paths for processing, one of which is used to perform average pooling on each channel in the height space range, and the other is used to perform average pooling on each channel in the width space range; Then, a one-dimensional convolutional layer is applied to the two paths respectively to enhance the decoupled original feature information in the horizontal and vertical directions; Subsequently, group normalization is used to reduce the interference between the two channels, which enhances the position information of the feature space dimension and reduces the number of parameters; Original feature F b The horizontal attention y h Score and vertical attention y w The decoupled expression of the score is: y h =σ(G n (H h (F b ))) y w =σ(G n (H w (F b ))) Among them, H h (·) and H w (·) represent the one-dimensional convolutional layers of the two paths, G n (·) is used to represent group normalization, σ(·) is used to represent the nonlinear activation function Sigmoid, and the attention scores obtained from the two paths are respectively combined with the feature F b multiply; Feature F after the decoupled attention module b A The output expression is: F b A =F b ×y h ×y w 。 5. The knowledge distillation method based on reverse feature fusion and classifier reuse according to claim 4 is characterized in that: The projection module is encapsulated into a basic unit, and the projection module includes two 1×1 and one adaptive convolutional layer; The two 1×1 convolutional layers are used to adjust the number of channels to the same number of channels as the deep features, and the adaptive convolutional layer is used according to F b The convolution kernel and step size are adaptively adjusted to the size of the feature space, so that the student spatial resolution size matches the teacher classifier. Used for subsequent reverse feature fusion and teacher classifier reuse; Features after projection module The output expression is: The projection function Proj(·) is designed to transform the student features at each stage after the decoupled attention module to the same spatial size as the teacher's deep features, where: where W∈R m×d is a weighting matrix with a spatial size of m×d.

6. The knowledge distillation method based on inverse feature fusion and classifier reuse according to claim 5, characterized in that: S4.3 at least includes the following steps: The projected features of layer b And the fusion features of the next level to integrate; Then, the fused features are sent to two paths for processing, one of which contains two convolutional layers with a convolution kernel of 1×1 to capture local information, and the other path captures global information through an additional global pooling layer; The feature information captured by the two paths is activated by the Sigmoid function to obtain the weight score, and is respectively compared with the feature Multiply them together and use the parameter w to adjust the importance of deep features and shallow features, and finally get the fusion features of this level The deep feature map with a total network level of B is used as the fusion feature of the first input, that is, The characteristic expression of the subsequent reverse fusion output is: Where D(·) is the fusion function from the feature fusion module, and the final fusion feature is expressed as 7. The knowledge distillation method based on inverse feature fusion and classifier reuse according to claim 6, characterized in that: S4.4 at least includes the following steps: The classifier layer parameters of the pre-trained teacher are reused into the fusion features, so that the student model can directly reuse the pre-trained teacher classifier for reasoning, and the feature alignment loss becomes the only source of generated gradients, where F B t It is the deep feature map of the teacher model with a total network layer of B; Construct feature distance loss function:

8. The knowledge distillation method based on inverse feature fusion and classifier reuse according to claim 1, characterized in that: The parameters of the student network architecture and the teacher classifier layer retained in S5 during the inference phase are specifically: In the student inference phase, the teacher network model only retains the classifier layer parameters and retains the student network architecture part.