A Pruning Method for Token Mixers of a Visual Backbone Model

By constructing teacher and student models in the visual backbone model and replacing the token mixer with knowledge distillation and reparameterization techniques, the problem of high computational cost in the visual backbone model is solved, and the advantages of model pruning and hardware deployment without loss of performance are achieved.

CN116188902BActive Publication Date: 2025-06-03SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310100058.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-06-03
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

In the existing visual backbone model, the token mixer calculation is costly, resulting in increased power consumption and delay, affecting the deployment and user experience of the end-side device.

Method used

By constructing the teacher model and student model, the teacher model includes a token mixer and layer normalization layer, while the student model uses an affine transformation module to replace the token mixer, and trains the student model through the knowledge distillation method, distilling the teacher model's knowledge into the student model, and finally fusion of the parameters of the affine transformation module to the layer normalization layer of the student model through reparameterization.

Benefits of technology

The pruning token mixer module is realized without losing performance. The resulting model has a minimalist architecture, which is conducive to hardware specialization and deployment, and significantly reduces the delay in the model application process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188902B_ABST
    Figure CN116188902B_ABST
Patent Text Reader

Abstract

The present invention discloses a pruning method for a token mixer of a vision backbone model. The method includes: obtaining a target image; inputting the target image into a trained vision model to obtain an image classification result. The training process of the vision model includes: constructing a teacher model, which includes a token mixer and a layer normalization layer is provided between the token mixer and the input mapping; constructing a student model, which includes a layer normalization layer corresponding to the teacher model and uses an affine transformation module to replace the token mixer; training the student model using a set loss function, and during the training process, distilling the knowledge learned by the teacher model into the student model; integrating the relevant parameters learned by the affine transformation module into the layer normalization layer of the student model through a reparameterization process, and using the reparameterized student model as the vision model. The model obtained by the present invention has a simple structure, low computational complexity, small latency, and is easy to deploy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a pruning method for a token mixer of a vision backbone model. Background Art

[0002] Existing vision backbone architectures (such as DeiT, Swin Transformer, MLP-Mixer, etc.) mostly follow the general macro-architecture MetaFormer of a general vision backbone model, that is: a pyramid architecture composed of several blocks, and each block internally consists of a token mixer that allows different spatial positions (tokens) to communicate and a multi-layer perceptron (ChannelMLP) in the channel dimension that allows different channels at the same position to communicate. The token mixer brings a relatively large computational cost to the model, resulting in more power consumption and latency, which is not conducive to deploying the model on edge devices and affects the user experience.

[0003] Traditional model compression techniques for general vision backbone models (such as vision transformers, ViT) can be classified into several categories, such as model architecture search (such as GLIT, AutoFormer, HR-NAS), knowledge distillation (such as DeiT, MiniViT, TinyViT), and designing more efficient token mixers (such as Swin Transformer, PoolFormer). The model architecture search technology of vision transformers refers to obtaining a model with the best trade-off between performance and computational complexity by automatically searching for the most appropriate number of channels, operation operators, and connection methods in each layer. The knowledge distillation technology refers to guiding the training of a more lightweight student model by means of the features or supervision information of a more complex teacher model (CNN, ViT, etc.). The design of an efficient token mixer generally refers to replacing the traditional self-attention module with a more efficient token mixer module, such as window-based self-attention and simple pooling operations (Pool). These new token mixing modules can simplify the computational complexity of the traditional self-attention mechanism and enable the entire model to achieve performance equivalent to or even better than before.

[0004] After analysis, the prior art mainly has the following defects:

[0005] 1) Although the ViT model obtained by the method based on model architecture search can achieve a good trade-off between accuracy and computational complexity, the obtained model architecture is not sufficiently regular and neat. Each block of the model has different operations, number of channels, connection methods, and number of heads. Although the carefully searched model architecture has certain advantages on the surface, it is not conducive to the implementation of the model in actual business and is also not conducive to hardware deployment. Moreover, the process of architecture search also requires a large search cost and training cost, which limits the actual application in development and business scenarios.

[0006] 2) Currently, the method based on the design of efficient token mixers aims to design token mixers with lower computational complexity to achieve an efficient general vision backbone model. Although such carefully designed token mixers can reduce the computational overhead to a certain extent, they still require token mixers and will incur corresponding latency costs.

[0007] 3) In the existing distillation methods for ViT, the target student model is generally obtained by directly reducing the depth, number of heads, and number of channels based on the teacher model. Although this method can reduce the computational complexity of the teacher model, the student model has a high latency and is not easy to deploy. Moreover, the target student model has the same structure during training and deployment, resulting in low flexibility in model training. This is because the architecture of the target student model cannot decouple the training and inference processes, and the student model also has a token mixer with a high computational complexity, resulting in a high latency during actual business deployment. Summary of the Invention

[0008] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for pruning the token mixer of a vision backbone model. The method includes the following steps:

[0009] Obtain a target image;

[0010] Input the target image into a trained vision model to obtain an image classification result;

[0011] Among them, the vision model is obtained according to the following steps:

[0012] Construct a teacher model, which includes a token mixer, and a layer normalization layer is provided between the token mixer and the input mapping;

[0013] Construct a student model, which includes a layer normalization layer corresponding to the teacher model, and replaces the token mixer with an affine transformation module;

[0014] Train the student model using a set loss function, and during the training process, distill the knowledge learned by the teacher model into the student model;

[0015] Integrate the relevant parameters learned by the affine transformation module into the layer normalization layer of the student model through a reparameterization process, and use the reparameterized student model as the visual model.

[0016] Compared with the prior art, the advantages of the present invention are that the obtained compressed model can directly remove the most complex token mixer module without loss of performance, and this model has a minimalist architecture, which is beneficial to the specialization of hardware and the deployment in practical applications, and significantly reduces the latency in the model application process.

[0017] Other features and advantages of the present invention will become clear from the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0019] Figure 1 is a flowchart of a method for pruning the token mixer of a visual backbone model according to an embodiment of the present invention;

[0020] Figure 2 is an architecture diagram of a visual backbone model of the prior art;

[0021] Figure 3 is an architecture diagram of a student model according to an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of the training and testing process of the overall architecture of a visual model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Now, various exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0024] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use.

[0025] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0026] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.

[0027] It should be noted that similar reference numerals and letters refer to similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0028] See Figure 1 As shown, the provided method for pruning the token mixer of the visual backbone model includes the following steps:

[0029] Step S110: Construct a teacher model and a student model, where the teacher model includes a token mixer, and a layer normalization layer is provided between the token mixer and the input mapping. The student model replaces the token mixer with an affine transformation module.

[0030] Figure 2 is the model architecture in the prior art. The left figure is the general architecture (MetaFormer) of the general visual backbone model, and each block contains two residual structures. The residual structure with the token mixer is used to fuse information from different spatial positions, and the residual structure with the channel MLP (multi-layer perceptron) is used to fuse information from different channels. The middle figure is the visual Transformer structure, and its token mixer is the self-attention module. The right figure is the visual MLP structure, and its token mixer is the spatial MLP module (Spatial MLP).

[0031] In this step S110, a teacher model is constructed using the visual backbone model, that is, the teacher model includes a token mixer, and a layer normalization layer is provided between the token mixer and the input mapping. Different from the prior art, in the embodiment of the present invention, the student model no longer includes a token mixer, but is replaced by an affine transformation module. See Figure 3 the architecture of the student model shown.

[0032] Step S120: Train the student model using a set loss function. During the training process, distill the knowledge learned by the teacher model into the student model.

[0033] In this step, the student model is trained using the knowledge distillation method. The proposed knowledge distillation method mimics the input-output mapping of the complex token mixer module using affine transformation operations. For example, define the affine transformation module as f(·), the token mixer module as g(·), and the knowledge distillation objective is:

[0034]

[0035] where, is the optimization objective of soft knowledge distillation, λ 1 , λ 2 , λ 3 are hyperparameters. is defined as follows:

[0036]

[0037] where T is the model, m is the layer index, and N, C, H, and W are the Batch Size, number of channels, height, and width of the features, respectively.

[0038]

[0039] where is the correlation metric function.

[0040]

[0041] where LN is layer normalization. a and t represent the relevant parameters of the affine transformation and the relevant parameters of the teacher token mixer, respectively.

[0042] The design of this knowledge distillation loss function hopes to imitate the behavior of a complex token mixer using a simple affine transformation operation, that is, using a pre-trained model with a token mixer as the teacher model and a vision model without a token mixer as the student model, and teaching the relevant knowledge of the token mixer to the student model through feature knowledge distillation and relational knowledge distillation strategies at appropriate positions.

[0043] Furthermore, the distribution of model features and the effective receptive field are analyzed. The analysis results show that the feature distillation method of the present invention can effectively change the feature distribution of the Token Mixer-Free type architecture to make it closer to the model with Token Mixer. Moreover, the trained student model can make the effective receptive field of the Token Mixer-Free type architecture larger. At the same time, some practical suggestions for the corresponding distillation strategies are given. For example, the size of the receptive field of the teacher Token Mixer affects the distillation effect: for teacher models with Token Mixer of comparable performance, the largest receptive field of Token Mixer (GFNet [global]) helps to improve the performance of the Token Mixer-Free student model.

[0044] Step S130, fuse the relevant parameters of the affine transformation module into the layer normalization layer of the student model through a reparameterization process, and use the simplified student model as the final vision model.

[0045] During the above distillation training, replace the token mixer with an affine transformation operation. Further, the parameters of this affine transformation can be fused into the previous layer normalization operation using the reparameterization process.

[0046] Specifically, the reparameterization process "absorbs" the affine transformation module into layer normalization. Therefore, it is necessary to obtain γ′ and β′ of the layer normalization layer. The derivation process is as follows:

[0047] M (2) = Affine(LN(M (1) , μ, σ, γ, β), s, t)

[0048] M (2) = LN(M (1) , μ, σ, γ′, β′)

[0049] Among them, M (1) is the input feature, and M (2) is the output feature. μ, σ, γ, and β are the mean, standard deviation, scaling parameter, and shift parameter of the LN layer, respectively.

[0050] The above two equations are equivalent, and we get:

[0051]

[0052] Among them, M is the input feature, and s and t are the parameters of the affine transformation. n, i, h, and w are the batch number, number of channels, height, and width of the feature, respectively.

[0053] Solving gives:

[0054] γ′ s = γ i s i , β′ i = β i s i + t i

[0055] Among them, γ′ i and β′ i are the variables after reparameterization.

[0056] The above is the process of reparameterization. The following introduces the improved knowledge distillation method:

[0057] After the above distillation training and reparameterization process, a simplified student model is obtained. In the present invention, the architecture after pruning the visual backbone model is also called IdenFormer, and its overall architecture is shown in Figure 4 (a), the training process is shown in Figure 4 (b), and the testing process is shown in Figure 4 (c). Figure 4 (c) is the final visual model architecture obtained after pruning the token mixer in the visual backbone model.

[0058] Step S140: Use the obtained vision model to classify or semantically segment the target.

[0059] With the obtained vision model, real-time classification of the target image can be performed. Pruning the overall part of the token mixer and combining it with appropriate training and optimization strategies, a series of models after pruning can be extremely simple architectures that only contain 1×1 convolutions and downsampling operations, which is beneficial for hardware specialization and deployment in practical applications.

[0060] The present invention can be applied to the fields related to computer vision image task recognition. For example, for the application scenarios of image recognition tasks (such as face recognition), when using models that follow the general macro-architecture MetaFormer of the general vision backbone model (such as DeiT, Swin Transformer, MLP-Mixer, etc.), these models have a large computational cost, resulting in more power consumption and latency, which is not conducive to deployment on edge devices and affects the user experience. At this time, the technical means proposed by the invention can be used to train a more concise vision model after removing the token mixer, reducing the latency of the model without affecting the accuracy and improving the user's satisfaction.

[0061] It should be noted that, in addition to being applied to image classification, object detection, and semantic segmentation tasks, the present invention may also be used in dense prediction tasks such as pose estimation.

[0062] In summary, the present invention provides a training strategy for the Token Mixer-Free architecture, that is, combining the improved feature distillation method with the structure reparameterization method. Use the Affine (affine) operation during training and incorporate its parameters into the pre-positioned LN (layer normalization layer) during inference. By virtue of the characteristics of the Affine operation in Token Mixer-Free, additional parameters are introduced into the model during training. Whether using multiple (the number is irrelevant to the model complexity during final inference) parallel or serial Affine operations during training, these additional parameters can be removed during inference, enhancing the expressive ability of the model. At the same time, make the simple Affine operation imitate the behavior of the complex Token Mixer (with the knowledge distillation technology of Feature Distillation), that is, use the pre-trained model with Token Mixer as the teacher model and the vision model without Token Mixer as the student model, and teach the student model the knowledge related to Token Mixer through feature knowledge distillation and relational knowledge distillation strategies at appropriate positions, hoping that the function of the simple Affine operation of the student model is as close as possible to the function of the Token Mixer of the teacher model.

[0063] It should be noted that, without departing from the spirit and scope of the present invention, those skilled in the art can make appropriate changes or modifications to the above embodiments. The core point of the present invention is how to ensure the model accuracy after pruning the token mixer. Therefore, other methods that can improve the model accuracy (such as other knowledge distillation schemes, model fusion strategies, etc.) belong to variants of the present invention.

[0064] To further verify the effectiveness of the present invention, experiments were conducted. Specifically, experiments were carried out on the large-scale image classification task ImageNet, and the experimental results of different-sized models have all proven the effectiveness of the present invention. The experimental results are shown in Tables 1 and 2.

[0065] Table 1: Experimental Results of Model Accuracy

[0066]

[0067] As can be seen from Table 1, compared with the model before pruning, the performance of the present invention after pruning the token mixer has basically no performance loss compared with that before pruning, and the accuracy can be guaranteed during the actual business use process.

[0068] Table 2: Experimental Results of Model Throughput

[0069] Model / Throughput (sheets / ms) Before Token Mixer Pruning Token Mixer Pruning PoolFormer-S12 4160.18 4899.60 PoolFormer-S24 2140.20 2530.48 PoolFormer-S36 1440.37 1699.94 PoolFormer-M36 1009.45 1185.33

[0070] As can be seen from Table 2, after pruning the token mixer, the model throughput has increased by approximately 18%, significantly reducing the latency of the model in the actual business scenario and bringing a better experience to users.

[0071] In summary, compared with the prior art, the present invention has the following advantages:

[0072] 1) By means of the structure reparameterization technology, the present invention uses affine transformation operations during training and integrates its parameters into the pre-positioned layer normalization layer during inference. The affine transformation operation introduces additional parameters for the model during training. Whether multiple parallel or serial affine transformation operations are used during training, these additional parameters can be removed during inference, enhancing the expression ability of the model.

[0073] 2) The purpose of the feature distillation method of the present invention is to enable simple affine transformation operations to mimic the behavior of complex token mixers (with the help of feature knowledge distillation technology): that is, using a pre-trained model with a token mixer as the teacher model and a vision model without a token mixer as the student model. By performing feature knowledge distillation and relational knowledge distillation strategies at appropriate positions, the relevant knowledge of the token mixer is taught to the student model, hoping that the function of the simple affine transformation operation of the student model is as close as possible to the function of the token mixer of the teacher model. In addition, since the architecture of the target student model does not have a token mixer, affine transformation operations can be used during training to introduce additional parameters into the model, enhancing its representational ability, enabling the student model to decouple the training and inference processes, and increasing the flexibility of the model optimization process.

[0074] 3) The model without a token mixer after pruning in the present invention uses the pre-trained weights of the model with a token mixer as initialization, and does not use the labels of the training set and the cross-entropy loss function during training, but only uses the distillation objective function as supervision, which is more conducive to improving the performance of the student model.

[0075] 4) The present invention directly removes the token mixer that occupies a relatively high latency as a whole, and keeps the other parts unchanged. The resulting student model can be a minimalist architecture that only includes 1×1 convolutions and downsampling operations, which is conducive to hardware specialization and deployment in practical applications.

[0076] 5) The present invention achieves extremely coarse-grained pruning. The structure of the lightweight vision backbone model obtained during inference only has 1×1 convolutions and several downsampling operations, and the structure is very regular, with obvious latency advantages.

[0077] 6) IdenFormer obtained based on knowledge distillation was evaluated on various vision tasks, including image classification, etc. The evaluation results show that the model obtained in the present invention achieves comparable or even better performance compared with SOTA CNNs, vision transformers, and vision MLPs. Therefore, IdenFormer can also be used as a general vision architecture.

[0078] 7) The present invention does not involve a complex architecture search process, is simple to implement, and the obtained model does not change in other aspects except that the token mixer is removed compared with the original teacher model, which is convenient for deployment and implementation.

[0079] The present invention can be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0080] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0081] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0082] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0083] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0084] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture comprising instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0085] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0086] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0087] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A pruning method for the token mixer of a visual backbone model, comprising the following steps: Obtain a target image; Input the target image into a trained visual model to obtain an image classification result; Wherein, the visual model is obtained according to the following steps: Construct a teacher model, which includes a token mixer and has a layer normalization layer between the token mixer and the input mapping; Construct a student model, which includes a layer normalization layer corresponding to the teacher model and replaces the token mixer with an affine transformation module; Train the student model using a set loss function, and during the training process, distill the knowledge learned by the teacher model into the student model; Integrate the relevant parameters learned by the affine transformation module into the layer normalization layer of the student model through a reparameterization process, and use the reparameterized student model as the visual model; Wherein, the loss function is expressed as: Among them, is the optimization objective of soft knowledge distillation, , , are hyperparameters, , , The definitions of are as follows: Among them, LN represents layer normalization, is a correlation metric function, T is the model, m is the layer index, N is the number of Batches of features, C, H, and W are the number of channels, height, and width of the features respectively; a and t represent the relevant parameters of the affine transformation and the relevant parameters of the teacher token mixer respectively, and F represents the norm.

2. The method according to claim 1, wherein, the reparameterization process includes: Integrate the affine transformation module into the layer normalization layer of the student model, expressed as: Equivalently transform to: Solve to get: Among them, is the input feature, is the output feature, are respectively the mean, standard deviation, scaling parameter and shift parameter of the layer normalization layer in the Affine process of the affine transformation, is the parameter of the affine transformation, are respectively the Batch number index, channel number index, height index and width index of the feature, and are the variables after reparameterization. Affine represents the affine transformation, and LN represents the layer normalization, and are respectively the scaling parameter and shift parameter of the layer normalization layer in the process of calculating the output feature , and are the variables before reparameterization of the i-th channel, and are the parameters of the affine transformation corresponding to the i-th channel.

3. The method according to claim 1, wherein, the teacher model is a general visual backbone model structure, or a vision Transformer structure, or a visual multi-layer perceptron MLP structure.

4. The method according to claim 1, wherein, the visual model is an architecture that only includes 1×1 convolutions and downsampling operations.

5. The method according to claim 1, wherein, the affine transformation module performs multiple parallel affine transformation operations or serial affine transformation operations.

6. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

7. A computer device, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, wherein, when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Sequence recommendation method for knowledge distillation based on land movement distance

    CN112507209A

  • System and method for knowledge-preserving neural network pruning

    US11200497B1