A Network Pruning Method Based on Attention Heads and Self-Distillation

By integrating token and attention head pruning with self-distillation, the method addresses the limitations of static pruning in multi-modal models, improving computational efficiency and accuracy by adaptively pruning based on input-specific importance and cross-modal interactions.

CN119886259BActive Publication Date: 2025-07-15BEIJING INST OF CONTROL & ELECTRONICS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510347973.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-15
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing multimodal large-modal pruning method cannot perform dynamic pruning, and fails to effectively consider the interaction between visual modes and linguistic modes, resulting in pruning inaccuracy and inefficiency.

Method used

Using a network pruning method based on attention head and self-distillation, combining token pruning and model structure pruning, student models are optimized through the self-distillation training process, the importance of token and attention head is evaluated using a learnable lightweight module, and the model is trained through the self-distillation method learned in the course to share parameters.

Benefits of technology

Dynamic pruning is achieved based on input, reducing the complexity of model calculations and parameter amount, improving the accuracy and efficiency of pruning, and avoiding additional fine-tuning processes in traditional knowledge distillation methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886259B_ABST
    Figure CN119886259B_ABST
Patent Text Reader

Abstract

This specification discloses a network pruning method based on attention heads and self-distillation, belonging to the technical field of multi-modal large model pruning. It includes obtaining a student model and a teacher model based on a token pruner, an attention head pruner, and a base model. The token pruner is used to evaluate the importance and prune the tokens input to the base model. The attention head pruner is used to evaluate the importance and prune the attention heads of the base model. The student model and the teacher model are trained based on training samples and a self-distillation objective function to obtain an optimized student model. The token pruner and the attention head pruner in the student model are activated. The token pruner and the attention head pruner in the teacher model are frozen, solving the problems existing in current multi-modal pruning methods, such as the inability to perform dynamic pruning and the failure to effectively consider the interaction relationship between the visual modality and the language modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-modal large model pruning, and particularly to a structured network pruning method based on minimum entropy. Background Art

[0002] The pruning methods of the Transformer model can be roughly classified into token pruning and model structure pruning. Token pruning refers to a pruning method that reduces the computational complexity of the model by reducing the number of input tokens. Model structure pruning refers to a pruning method that reduces the computational complexity and the number of parameters of the model by removing redundant model structures.

[0003] For multi-modal large model pruning methods, since multi-modal large models are usually models based on the Transformer architecture, the classification of Transformer model pruning methods can also be followed. In previous studies, Gan et al. directly applied the pruning method of a single-modal model to the model pruning of multi-modal large models, verifying the existence of the lottery hypothesis in multi-modal large models. Fang et al. proposed DistillVLM, proving that the knowledge distillation method can be used to imitate the attention distribution of multi-modal large models. Shi et al. proposed UPop, which completes the pruning of multi-modal large models through two stages of consistent search and aggressive pruning.

[0004] Most of the current multi-modal large model pruning methods are static model pruning. That is, the pruning mask of the model structure has nothing to do with the input token sequence and is determined once the pruning is completed. However, for different input token sequences, the importance of each attention head is different, and the importance of each token is also different.

[0005] Secondly, many multi-modal large model pruning methods do not consider the interaction relationship between the visual modality and the language modality, and simply apply the pruning method of a single-modal Transformer model to the multi-modal model. Such a pruning method may cause inaccuracy of the pruned model.

[0006] The method of using knowledge distillation for model pruning requires fine-tuning the teacher network on the target dataset. The fine-tuning of the teacher network and the knowledge distillation process cannot be carried out simultaneously. This results in low efficiency of model pruning.

[0007] The problems of the above pruning methods are that they cannot achieve dynamic pruning according to the input. And they either only perform model structure pruning or only perform token pruning. Moreover, the interaction relationship between the visual modality and the language modality is not effectively considered. Summary of the Invention

[0008] The purpose of the present invention is to provide a network pruning method based on attention heads and self-distillation, which solves the problems existing in current multi-modal pruning methods, namely, the inability to perform dynamic pruning and the failure to effectively consider the interaction relationship between the visual modality and the language modality.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] On the one hand, this specification provides a network pruning method based on attention heads and self-distillation, including:

[0011] Step 102: Based on a token pruner, an attention head pruner, and a base model, obtain a student model and a teacher model; the token pruner is used to evaluate the importance of and prune the tokens input to the base model; the attention head pruner is used to evaluate the importance of and prune the attention heads of the base model;

[0012] Step 104: Based on training samples and a self-distillation objective function, train the student model and the teacher model to obtain an optimized student model; the token pruner and the attention head pruner in the student model are activated; the token pruner and the attention head pruner in the teacher model are frozen.

[0013] On the other hand, this specification provides a network pruning device based on attention heads and self-distillation, including:

[0014] A pruning model construction module, which is used to obtain a student model and a teacher model based on a token pruner, an attention head pruner, and a base model; the token pruner is used to evaluate the importance of and prune the tokens input to the base model; the attention head pruner is used to evaluate the importance of and prune the attention heads of the base model;

[0015] A model self-distillation module, which is used to train the student model and the teacher model based on training samples and a self-distillation objective function to obtain an optimized student model; the token pruner and the attention head pruner in the student model are activated; the token pruner and the attention head pruner in the teacher model are frozen.

[0016] Based on the above technical solutions, this specification can achieve the following technical effects:

[0017] (1) Combine token pruning and model structure pruning. The pruning method in this application includes both screening input tokens, reducing the number of tokens based on a learnable lightweight module, and pruning attention heads, that is, pruning the model structure by removing unimportant attention heads. The computational complexity and the number of parameters of the model are minimized based on the hybrid paradigm.

[0018] (2) Self-distillation training process. During the process of training the parameters of adaptive pruning, we adopt a self-distillation training method based on curriculum learning. This enables the pruned model and the pruning model to share the same parameters. The only difference is that the pruned model activates the model pruner during the forward propagation process. In this way, the separate fine-tuning process for the teacher network in traditional knowledge distillation methods is avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic flowchart of a network pruning method based on attention heads and self-distillation in an embodiment of the present invention.

[0020] Figure 2 Framework diagram of the network pruning method based on attention heads and self-distillation in an embodiment of the present invention.

[0021] Figure 3 Schematic structural diagram of a network pruning device based on attention heads and self-distillation in an embodiment of the present invention.

[0022] Figure 4 Schematic diagram of an electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. According to the following description and the claims, the advantages and features of the present invention will be more clearly understood. It should be noted that the drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.

[0024] It should be noted that, in order to clearly illustrate the content of the present invention, the present invention specifically gives multiple embodiments to further explain different implementation manners of the present invention. Among them, the multiple embodiments are illustrative rather than exhaustive. In addition, for the sake of simplicity of description, the content already mentioned in the previous embodiments is often omitted in the subsequent embodiments. Therefore, the content not mentioned in the subsequent embodiments can be correspondingly referred to the previous embodiments. Embodiment 1

[0025] Please refer to Figure 1 , Figure 1 as shown in the network pruning method based on attention heads and self-distillation provided in this embodiment. In this embodiment, the method includes:

[0026] Step 102, obtaining a student model and a teacher model based on a token pruner, an attention head pruner, and a base model; the token pruner is used to evaluate the importance and prune the tokens input to the base model; the attention head pruner is used to evaluate the importance and prune the attention heads of the base model;

[0027] In this embodiment, the token pruning device is a lightweight module based on MLP; the lightweight module based on MLP includes a local policy network and a global policy network; the local policy network is used to evaluate the importance of a single token; the global policy network is used to evaluate the importance of cross-modal tokens.

[0028] In this embodiment, the token pruning device performs importance evaluation and pruning on the tokens input to the base model, including:

[0029] Step 202: Input the target token into the local policy network, and calculate the local importance score of the target token.

[0030] In this embodiment, step 202 includes:

[0031] Step 2021: Perform dimensionality reduction processing on the target token to obtain the target token after dimensionality reduction.

[0032] Step 2022: Through stochastic gradient descent learning, obtain the learnable parameters of the local policy network.

[0033] Step 2023: Input the target token after dimensionality reduction into the local policy network, and calculate the local importance score of the target token.

[0034] Step 204: Perform fusion processing and mapping processing on the global representations of the visual modality and the text modality to obtain the cross-modal global representation.

[0035] In this embodiment, the global representations based on the visual modality and the text modality are the corresponding class tokens.

[0036] Step 206: Input the cross-modal global representation and the target token into the global policy network, and calculate the cross-modal global importance score of the target token.

[0037] Step 208: Based on the sum of the local importance score and the global importance score of the target token, obtain the final importance score of the target token.

[0038] Step 210: Prune or retain the target token based on the final importance score of the target token.

[0039] In this embodiment, one implementation manner of step 210 is:

[0040] Step 2101: Map the final importance score of the target token to between 0 and 1 based on the neural network threshold function to obtain the pruning mask of the target token.

[0041] Step 2102: If the pruning mask of the target token is 1, retain the target token; otherwise, remove the target token.

[0042] In this embodiment, the attention head pruning device performs importance evaluation and pruning on the attention heads of the base model, including:

[0043] Step 302: Input the global representations of the current modality and another modality into the attention pruning device, and calculate the importance scores of the attention heads of the current modality;

[0044] Step 304: Map the importance scores of the attention heads of the current modality to between 0 and 1 based on the neural network threshold function to obtain the pruning mask of the attention heads of the current modality;

[0045] Step 306: If the pruning mask of the attention head is 1, retain the attention head of the current modality; otherwise, discard the attention head of the current modality.

[0046] Step 104: Train the student model and the teacher model based on the training samples and the objective function of self-distillation to obtain an optimized student model; the token pruning device and the attention head pruning device in the student model are activated; the token pruning device and the attention head pruning device in the teacher model are frozen.

[0047] In this embodiment, before step 104, it further includes:

[0048] Step 103: Obtain the objective function of self-distillation based on the parameters of the student model and the teacher model and the objective loss function.

[0049] In this embodiment, one implementation manner of step 104 is:

[0050] Step 402: Input the training samples into the teacher model to obtain the prediction confidence of the training samples in the teacher model;

[0051] Step 404: Sort the training samples in descending order based on the prediction confidence of the training samples in the teacher model to obtain descending-order samples;

[0052] Step 406: Extract target samples from the descending-order samples and input them into the teacher model and the student model, and update the model parameters in the method of batch gradient descent according to the objective function of self-distillation to obtain the student model and the teacher model with updated parameters.

[0053] In this embodiment, the output result of the student model with updated parameters is aligned with the output result of the teacher model.

[0054] Specifically, referring to Figure 2 , a network pruning method based on attention heads and self-distillation is specifically implemented as follows:

[0055] The first step: Adaptive token pruning

[0056] By evaluating the importance of the tokens input to the model and removing the tokens with lower importance, the computational complexity of the model is reduced. The evaluation of token importance is achieved through a learnable lightweight module. The calculated token importance combines the local importance of the tokens and the cross-modal global importance.

[0057] In this step, the input tokens of each encoder block are gradually pruned to eliminate unimportant tokens, and then the more important tokens are passed to the next block. To evaluate the importance of each token, a lightweight MLP-based module, named the XModal-aware trimmer, is inserted before the encoder block of each multimodal model. Suppose tokens are represented as , then first is input to the Local Policy Network to extract the local importance, and this process can be expressed as follows:

[0058]

[0059] where represents the importance of each token. is obtained by reducing the dimension of . When calculating , each token is calculated independently without considering the contribution of tokens in other modalities in the multimodal input. a is a learnable parameter obtained by stochastic gradient descent learning.

[0060] To introduce the importance of tokens in other modalities to the current token, cross-modal interaction is required. Specifically, the global representations of the visual and text modalities (using the class token as the global representation of each modality) are fused, and then mapped to obtain the cross-modal global representation g. Finally, it is input to the global policy network together with to calculate the cross-modal token importance score :

[0061]

[0062] where represents the mapping layer, represents the hybrid paradigm. Finally, the importance of each token is calculated as follows:

[0063]

[0064] During inference, the pruning mask is directly obtained from Sampling results in: 1 indicates retaining the corresponding token, otherwise it indicates removing the corresponding token. It is a neural network threshold function used to map variables between 0 and 1.

[0065] Step 2: Adaptive Attention Head Pruning

[0066] This method designs a modality - adaptive attention head pruner to estimate the importance of attention heads and integrates it into the multi - head self - attention module and the multi - head cross - attention module. Using the importance estimated by the attention head pruner, attention heads with low importance can be removed, thereby reducing the computational complexity and the number of parameters of the model.

[0067] The multi - modal large model can capture the interactions between tokens within and across modalities through the multi - head self - attention mechanism and the multi - head interactive attention mechanism. However, the complexity of attention calculation is related to the input, that is, simple samples do not require the model to spend too much attention on modalities, while complex samples require more attention from the model. For this reason, this project designs a modality - adaptive attention head pruner and integrates it into the attention module. Specifically, first, the global representation of the input sequence is input into the attention head pruner, which is shown as follows:

[0068]

[0069] Among them, respectively represent the class token representations of the current modality and another modality. Similar to the token pruner, during inference, directly obtain to decide whether to remove or retain the corresponding attention head.

[0070] Step 3: Self - Distillation Performance Improvement

[0071] The token pruner and the attention head pruner are trained on the dataset simultaneously with the base model. In this method, we call it self - distillation. The self - distillation process simultaneously trains a pruned student model and the teacher model before pruning, and aligns the output of the student model with that of the teacher model. The two models share parameters. The difference is that during forward calculation, the pruner is enabled for the student model, while the teacher model does not enable the pruner during forward calculation.

[0072] During model training, this project designs a self - distillation strategy to prompt the pruned model to output prediction results that are aligned with the model before pruning . Among them, and are models that share parameters. The difference between them is that during the forward process the adaptive pruner is activated, while The adaptive pruning mechanism is frozen. During training, and are optimized and updated simultaneously. The objective function of self-distillation is expressed as follows:

[0073]

[0074] where represent the input and output logits (fully connected layers) respectively, and is the recognition loss function. This approach can avoid the need for additional fine-tuning of the teacher model in traditional knowledge distillation methods.

[0075] During the process of optimizing this objective function, this project proposes an optimization method based on curriculum learning. Curriculum learning mimics the way humans learn things through a curriculum, that is, starting from simple samples and gradually increasing the learning difficulty. The optimization method based on curriculum learning is as follows:

[0076] (1) Using the prediction confidence of samples in the teacher model as an indicator, sort the samples in descending order to obtain .

[0077] (2) Extract samples in the sorted order, pass them through the teacher model and the student model, and update the parameters using the SGD method according to the objective function in 4.5.

[0078] (3) After the parameter update is completed, fix the pruning weights of the student network for evaluating the importance of input tokens and attention heads during the test phase.

[0079] During the update process of curriculum learning, samples with high confidence are regarded as simple samples, while samples with low confidence are regarded as difficult samples. Through the sample selection strategy of starting from the easy ones and then moving to the difficult ones, the weights of the student model are better optimized to approximate the output distribution of the teacher model.

[0080] The above is the process of the network pruning method based on attention heads and self-distillation.

[0081] In summary, this method can achieve the following technical effects:

[0082] (1) Combine token pruning and model structure pruning. The pruning method in this application includes both screening input tokens and reducing the number of tokens based on a learnable lightweight module, as well as pruning attention heads to remove unimportant attention heads from the model structure. Based on the hybrid paradigm, the computational complexity and the number of parameters of the model are minimized to the greatest extent.

[0083] (2) Self-distillation training process. During the process of training the parameters of adaptive pruning, we adopt a self-distillation training method based on curriculum learning. This enables the pruned model and the pruning model to share the same parameters. The only difference is that the pruned model activates the model pruner during the forward propagation process. In this way, the separate fine-tuning process of the teacher network in traditional knowledge distillation methods is avoided. Embodiment 2

[0084] Please refer to Figure 3 , Figure 3 shown in the network pruning device based on attention heads and self-distillation provided in this embodiment. In this embodiment, the device includes:

[0085] A pruning model construction module, configured to obtain a student model and a teacher model based on a token pruner, an attention head pruner, and a base model; the token pruner is used to evaluate the importance and prune the tokens input to the base model; the attention head pruner is used to evaluate the importance and prune the attention heads of the base model;

[0086] A model self-distillation module, configured to train the student model and the teacher model based on training samples and a self-distillation objective function to obtain an optimized student model; the token pruner and the attention head pruner in the student model are activated; the token pruner and the attention head pruner in the teacher model are frozen.

[0087] Optionally, the token pruner is a lightweight module based on an MLP; the lightweight module based on an MLP includes a local policy network and a global policy network; the local policy network is used to evaluate the importance of a single token; the global policy network is used to evaluate the importance of cross-modal tokens.

[0088] Optionally, it further includes a token pruning module, configured to evaluate the importance and prune the tokens input to the base model.

[0089] Optionally, the token pruning module includes:

[0090] A token local importance analysis unit, configured to input a target token into the local policy network and calculate the local importance score of the target token;

[0091] A cross-modal global representation acquisition unit, configured to fuse and map the global representations of the visual modality and the text modality to obtain a cross-modal global representation;

[0092] A token global importance analysis unit, configured to input the cross-modal global representation and the target token into the global policy network and calculate the cross-modal global importance score of the target token;

[0093] A token importance score determination unit for obtaining the final importance score of a target token based on the sum of the local importance score and the global importance score of the target token;

[0094] A token pruning unit for pruning or retaining the target token based on the final importance score of the target token.

[0095] Optionally, the token local importance analysis unit includes:

[0096] A token dimensionality reduction processing subunit for performing dimensionality reduction processing on the target token to obtain the target token after dimensionality reduction;

[0097] A learnable parameter determination subunit for obtaining the learnable parameters of the local policy network through stochastic gradient descent learning;

[0098] A local analysis subunit for inputting the target token after dimensionality reduction into the local policy network and calculating the local importance score of the target token.

[0099] Optionally, the global representation based on the visual modality and the text modality is the corresponding category token.

[0100] Optionally, the token pruning unit includes:

[0101] A token pruning mask calculation subunit for mapping the final importance score of the target token to between 0 and 1 based on the neural network threshold function to obtain the pruning mask of the target token;

[0102] A token pruning subunit for retaining the target token if the pruning mask of the target token is 1, otherwise, removing the target token.

[0103] Optionally, it further includes an attention head pruning module for performing importance evaluation and pruning on the attention heads of the base model;

[0104] Optionally, the attention head pruning module further includes:

[0105] Inputting the global representations of the current modality and another modality into the attention pruner to calculate the importance score of the attention heads of the current modality;

[0106] Mapping the importance score of the attention heads of the current modality to between 0 and 1 based on the neural network threshold function to obtain the pruning mask of the attention heads of the current modality;

[0107] If the pruning mask of the attention heads is 1, then retain the attention heads of the current modality, otherwise, the attention heads of the current modality.

[0108] Optionally, it further includes:

[0109] The objective function determination module for self-distillation is used to obtain the objective function of self-distillation based on the parameters of the student model and the teacher model and the objective loss function.

[0110] Optionally, the model self-distillation module includes:

[0111] The prediction confidence determination unit is used to input the training samples into the teacher model to obtain the prediction confidence of the training samples in the teacher model;

[0112] The sample descending sorting unit is used to sort the training samples in descending order based on the prediction confidence of the training samples in the teacher model to obtain the descending samples;

[0113] The model parameter update unit is used to extract the target samples from the descending samples and input them into the teacher model and the student model, and update the model parameters by the method of batch gradient descent according to the objective function of self-distillation to obtain the student model and the teacher model with updated parameters.

[0114] Optionally, the output result of the student model with updated parameters is aligned with the output result of the teacher model.

[0115] Based on this, the present device can obtain the following technical effects:

[0116] (1) Combine token pruning and model structure pruning. The pruning method in the present application includes both screening the input tokens and reducing the number of tokens based on a learnable lightweight module, and pruning the attention heads to remove the unimportant attention heads in the model structure pruning. The computational complexity and the number of parameters of the model are minimized based on the hybrid paradigm.

[0117] (2) The training process of self-distillation. During the process of training the parameters of the adaptive pruning, we adopt the training method of self-distillation based on curriculum learning. The pruned model and the pruning model share the same parameters, and the only difference is that the pruned model will activate the model pruner during the forward propagation process. In this way, the separate fine-tuning process of the teacher network in the traditional knowledge distillation method is avoided. Example 3

[0118] Please refer to Figure 4, this embodiment provides an electronic device, which includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a network pruning method based on attention heads and self-distillation at the logical level. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.

[0119] The network interface, processor, and memory can be interconnected through a bus system. The above bus can be divided into an address bus, a data bus, a control bus, etc.

[0120] The memory is used to store programs. Specifically, the program can include program code, and the above program code includes computer operation instructions. The memory can include a read-only memory and a random access memory, and provides instructions and data to the processor.

[0121] The processor is used to execute the program stored in the above memory, and specifically execute:

[0122] Step 102, obtain a student model and a teacher model based on a token pruner, an attention head pruner, and a base model; the token pruner is used to evaluate and prune the importance of tokens input to the base model; the attention head pruner is used to evaluate and prune the importance of the attention heads of the base model;

[0123] Step 104, train the student model and the teacher model based on training samples and a self-distillation objective function to obtain an optimized student model; the token pruner and the attention head pruner in the student model are activated; the token pruner and the attention head pruner in the teacher model are frozen.

[0124] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed through the integrated logic circuit of the processor's hardware or instructions in software form.

[0125] Based on the same inventive concept, an embodiment of this specification also provides a computer-readable storage medium. The above computer-readable storage medium stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the above electronic device is caused to execute Figures 1-2 The corresponding embodiment provides a network pruning method based on attention heads and self-distillation.

[0126] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-readable storage media containing computer-usable program code.

[0127] In addition, for the specific implementation of the above system, since it is basically similar to the method implementation, the description is relatively simple. For the relevant parts, please refer to the partial description of the method implementation. Moreover, it should be noted that in each module of the system of this application, the components are logically divided according to the functions to be realized. However, this application is not limited thereto, and the components can be re-divided or combined as needed.

[0128] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0129] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, in the process depicted in the drawings, it is not necessarily required to show a specific order or a continuous order to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.

[0130] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A network pruning method based on attention heads and self-distillation, characterized in that Including: Based on a token pruning device, an attention head pruning device, and a base model, a student model and a teacher model are obtained; the token pruning device is used to evaluate the importance and prune the tokens input to the base model; the attention head pruning device is used to evaluate the importance and prune the attention heads of the base model; Based on training samples and a self-distillation objective function, the student model and the teacher model are trained to obtain an optimized student model; the token pruning device and the attention head pruning device in the student model are activated; The token pruning device and the attention head pruning device in the teacher model are frozen; The token pruning device evaluating the importance and pruning the tokens input to the base model includes: Inputting a target token into a local policy network to calculate the local importance score of the target token; Fusing and mapping the global representations of the visual modality and the text modality to obtain a cross-modal global representation; Inputting the cross-modal global representation and the target token into a global policy network to calculate the cross-modal global importance score of the target token; Based on the sum of the local importance score and the global importance score of the target token, obtaining the final importance score of the target token; Pruning or retaining the target token based on the final importance score of the target token; The inputting the target token into the local policy network to calculate the local importance score of the target token includes: Performing dimensionality reduction processing on the target token to obtain a dimensionally reduced target token; Through stochastic gradient descent learning, obtaining the learnable parameters of the local policy network; Inputting the dimensionally reduced target token into the local policy network to calculate the local importance score of the target token; The attention head pruning device evaluating the importance and pruning the attention heads of the base model includes: Inputting the global representations of the current modality and another modality into the attention pruning device to calculate the importance score of the attention head of the current modality; Based on a neural network threshold function, mapping the importance score of the attention head of the current modality to between 0 and 1 to obtain the pruning mask of the attention head of the current modality; If the pruning mask of the attention head is 1, retaining the attention head of the current modality, otherwise, removing the attention head of the current modality; The training of the student model and the teacher model based on training samples and a self-distillation objective function to obtain an optimized student model includes: Inputting the training samples into the teacher model to obtain the prediction confidence of the training samples in the teacher model; Based on the prediction confidence of the training samples in the teacher model, performing a descending order sorting on the training samples to obtain descending order samples; Extracting target samples from the descending order samples and inputting them into the teacher model and the student model, and updating the model parameters in accordance with the self-distillation objective function by means of batch gradient descent to obtain the student model and the teacher model with updated parameters.

2. The method according to claim 1, wherein The token pruning device is a lightweight module based on an MLP; the lightweight module based on an MLP includes a local policy network and a global policy network; the local policy network is used to evaluate the importance of a single token; the global policy network is used to evaluate the importance of cross-modal tokens.

3. The method according to claim 1, characterized in that, The global representations of the visual modality and the text modality are corresponding class tokens.

4. The method according to claim 1, wherein Pruning or retaining the target token based on the final importance score of the target token includes: Mapping the final importance score of the target token to between 0 and 1 based on the neural network threshold function to obtain the pruning mask of the target token; If the pruning mask of the target token is 1, retain the target token, otherwise, remove the target token.

5. The method according to claim 1, characterized in that, Before training the student model and the teacher model based on the training samples and the self-distillation objective function to obtain the optimized student model, it further includes: Obtaining the self-distillation objective function based on the parameters of the student model and the teacher model and the target loss function.

6. The method according to claim 1, wherein The output result of the student model after the parameter update is aligned with the output result of the teacher model.

Citation Information

Patent Citations

  • Pollen image classification method based on cross attention distillation Transformer

    CN113887610A