Model pruning method and system based on weighted attention alignment

By adopting a weighted attention alignment mechanism during the model pruning process and combining pruning and migration for joint training, the problem of degradation in the existing technology of models when migrating to small data sets is solved, and efficient model pruning and migration is achieved, reducing calculation and time costs.

CN119990231AActive Publication Date: 2025-05-13SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510000758.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-13
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

Existing model pruning methods have degraded performance when migrating to small datasets, and the lack of effective combination of pruning and migration, resulting in high computational and time costs.

Method used

The model pruning method based on weighted attention alignment is adopted, and the initial model is obtained through pre-fine-tuning and pre-pruning. Then, the guidance of the pre-trained model is introduced in the pruning fine-tuning training through the weighted attention alignment mechanism, and combined training is carried out in combination with pruning and migration.

Benefits of technology

Effectively utilizing pre-training knowledge improves the performance of the model on small data sets, reduces the time cost of the training process, and improves the interpretability of pruning and feature alignment effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990231A_ABST
    Figure CN119990231A_ABST
Patent Text Reader

Abstract

The invention discloses a model pruning method and system based on weighted attention alignment, relates to artificial intelligence, and provides a scheme for solving the problems of lack of association and the like in the prior art. Comprising the following steps: S1, pre-fine tuning: initializing a target model by using pre-training model parameters, and then performing L1 regularization fine tuning on batch normalization layer parameters for a plurality of times by using data of a target domain; s2, pre-pruning: globally sorting absolute values of batch normalized layer parameters of the trained target model, and cutting off a tail channel to obtain a pre-pruned model; and S3, model pruning training based on weighted attention alignment: performing pruning fine tuning training on the pruned model, and introducing guidance of a pre-training model through a weighted attention alignment mechanism to obtain a pruning model for the target domain data set. The method has the advantages that efficient utilization of pre-training knowledge is achieved, effective combination of pruning and migration is achieved, pruning performance is improved, and training time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence technology, and in particular to a model pruning method and system based on weighted attention alignment. Background Art

[0002] In recent years, deep learning has developed rapidly, affecting various fields including computer vision, natural language processing, and speech recognition. Among them, convolutional neural networks for images have shown strong feature extraction capabilities and have become the basic model in image recognition tasks. In order to improve the performance of the model, convolutional neural networks have developed in the direction of deeper, wider, and more complex, resulting in a sharp increase in model storage space and computational complexity. In response to this, more and more people are studying model compression, that is, how to reduce the size, complexity, and computational complexity of deep models through various technical means, improve the inference speed and efficiency of the model, and maintain the performance of the model as much as possible so that it can be deployed and run on resource-constrained devices. Among them, model pruning can accelerate neural networks by removing unimportant components, retain a high degree of accuracy, and is simple and effective to operate. Filter pruning has received a lot of attention because it is not only applicable to any convolutional model architecture, but also has deployment-friendly properties.

[0003] However, the generalization of the pruned model is greatly weakened. Specifically, after pruning the large model pre-trained on ImageNet, although the performance on ImageNet is not significantly reduced, if it is migrated to a small dataset, the performance is greatly reduced compared to the migration result of the full model. In other words, the pruned model is more sensitive to the distribution differences between different datasets and cannot be effectively migrated to other datasets. That is, only the pre-trained full model can improve the performance of the pruned model on a small dataset.

[0004] To achieve the above goals, some methods have been proposed by Liu et al. (Liu B, Cai Y, Guo Y, et al. TransTailor: Pruning the Pre-trained Model for Improved Transfer Learning. [Z]. 20218627-8634). The algorithm flow of fine-tuning first and then pruning is proposed: the pre-trained model is first fine-tuned for a small data set, and then the fine-tuned model is pruned during training. The pruning basis mainly includes Taylor expansion, cumulative sum of activation values, attention mechanism, etc. However, the existing methods have the following problems: First, pre-training knowledge is the basis for improving performance, but they only use pre-training knowledge as a simple initialization and do not make full use of the large amount of useful knowledge in the pre-training model; second, there is a lack of effective combination of pruning and transfer. Pruning naturally ranks the importance of feature channels, while the transfer used in the existing methods is only fine-tuning and does not involve feature alignment; third, for the two training processes of transfer and pruning, they simply perform them sequentially or even iteratively, which consumes a lot of computational and time costs. Summary of the invention

[0005] The purpose of the present invention is to provide a model pruning method and system based on weighted attention alignment to solve the problems existing in the above-mentioned prior art.

[0006] The model pruning method based on weighted attention alignment described in the present invention comprises the following steps:

[0007] S1. Pre-fine-tuning: First, use the pre-trained model parameters to initialize the target model, and then use the data of the target domain to perform L1 regularization fine-tuning on the batch normalization layer parameters several times;

[0008] S2. Pre-pruning: globally sort the absolute values ​​of the batch normalization layer parameters of the trained target model, prune the channels at the end, and obtain the pre-pruned model;

[0009] S3. Model pruning training based on weighted attention alignment: Prune and fine-tune the pruned model, introduce the guidance of the pre-trained model through the weighted attention alignment mechanism, and obtain the pruned model for the target domain dataset.

[0010] The model pruning system based on weighted attention alignment described in the present invention uses the model pruning method to prune the model to obtain a pruned model for a small data set.

[0011] The model pruning method and system based on weighted attention alignment described in the present invention have the advantages of making up for the defects in the prior art:

[0012] 1. Weighted attention alignment: To address the problem of insufficient utilization of pre-trained knowledge, we first calculate the weighted attention feature map through the parameters of the batch normalization layer, and then align the weighted attention maps of the large model and the small model to achieve efficient utilization of pre-trained knowledge.

[0013] 2. Combining pruning and migration: To address the issue of pruning and migration being unrelated, the batch normalization layer parameters are used to effectively combine the two. The batch normalization layer parameters are a quantitative assessment of the importance of each channel, which is not only used as the basis for channel pruning, but also used to calculate the weighted attention feature map, achieving an effective combination of pruning and migration.

[0014] 3. Optimization of joint training strategy: To address the problems of step-by-step and iterative algorithm training, the designed algorithm process combines migration and pruning training, and performs weighted attention migration while pruning and fine-tuning training, which not only improves pruning performance but also significantly reduces the time cost of the training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the principle of the model pruning method described in the present invention.

[0016] Figure 2 It is a schematic diagram of the calculation process of the weighted attention map described in the present invention. DETAILED DESCRIPTION

[0017] The present invention can make full use of pre-training knowledge and improve the performance of the pruning model for small data sets by calculating the weighted attention map of batch normalization layer parameters and applying alignment constraints. At the same time, it also effectively combines pruning and migration through batch normalization layer parameters to improve the interpretability of the method, and performs weighted attention migration while pruning fine-tuning training, which not only improves pruning performance, but also significantly reduces the time cost brought by the existing step-by-step or even iterative training process.

[0018] A model pruning method based on weighted attention alignment is Figure 1 As shown, the specific steps include:

[0019] S1. Pre-fine-tuning: First, use the pre-trained model parameters to initialize the target model, and then use the data from the target domain to perform N-times L1 regularization fine-tuning on the batch normalization layer parameters, where N is usually set to 10. This step can effectively avoid cutting off channels that are useful for the target domain data and affect subsequent fine-tuning.

[0020] S2. Pre-pruning: The absolute values ​​of the batch normalization layer parameters of the trained target model represent the importance of the corresponding channels. They are globally sorted and the channels at the end are pruned to obtain the pre-pruned model.

[0021] S3. Model pruning training based on weighted attention alignment: Prune and fine-tune the pruned model, introduce the guidance of the pre-trained model through the weighted attention alignment mechanism, and obtain the pruned model for the target domain dataset.

[0022] Since the batch normalization layer parameters are the basis for judging the importance of channels, they are widely used in model pruning tasks. Based on the weighted attention alignment mechanism, the parameters are extended to the pruning and migration tasks. On the one hand, pruning and migration are effectively combined to solve the problem of the two being unrelated. On the other hand, after model channel pruning, the number of channels is reduced, and conventional feature alignment methods in transfer learning are difficult to apply. The attention mechanism can integrate channel dimensions and solve the problem of feature alignment with inconsistent channel numbers.

[0023] In the weighted attention alignment mechanism, the calculation steps of the weighted attention map are as follows: Get the batch normalization layer parameters of the convolutional neural network Where C represents the number of channels, and the intermediate tensor output by the activation layer after the batch normalization layer is obtained Where H and W represent the height and width of the feature map respectively;

[0024] The weighted attention map is calculated using the function WAM(w,x), such as Figure 2 As shown, it specifically includes the following sub-steps:

[0025] S31. Perform softmax normalization on the batch normalization layer parameter w to obtain the weight vector

[0026] S32. Pooling the intermediate tensor Where A represents the size of the feature map after pooling;

[0027] S33. The weight vector And the intermediate tensor after pooling Multiply by channel to get the weighted tensor

[0028] S34. Weighted tensor summed by channel in Representing a tensor The forward slice of ; then the weighted attention map is obtained after Frobenius normalization where F i,j Represents the element in the i-th row and j-th column of the matrix.

[0029] In step S3, the pre-trained model is for a large data set and cannot be changed, and the pruned model is for a small data set and is generated by pruning the pre-trained model. A weighted attention alignment module, an output feature alignment module, and a target domain perception training module are set.

[0030] The weighted attention alignment module uses the weighted attention alignment mechanism to obtain the weighted attention graphs of the pre-trained model and the pruned model, and aligns the weighted attention graphs of the two at the corresponding positions, using the knowledge of the pre-trained model to guide the fine-tuning training of the pruned model to improve the pruning performance. The specific application location of the weighted attention alignment module is all the "convolutional layer-batch normalization layer-activation layer" network structure blocks in the network. For a network with a total of B such network structure blocks, its loss function is in denote the parameters of the batch normalization layer in the kth network structure block of the pre-trained model and the pruned model, respectively. denote the output feature tensor of the kth network structure block of the i-th image input pre-trained model and pruned model, respectively. WAM(·,·) is the weighted attention map function. Represents the square of the Frobenius norm.

[0031] Since the two models target different data sets, the output logical probability distribution sizes are different and cannot be aligned directly. Therefore, the output feature alignment module replaces the classification head of the pruned model with the classification head of the pre-trained model and then performs probability distribution alignment. Its loss function is in Denote the output probability distribution of the i-th image after it is input into the feature extractor of the pre-trained model and the pruning model, and then input into the classification head of the pre-trained model, D KL (·) is the KL divergence function.

[0032] The target domain-aware training module fine-tunes the pruning model on a small dataset and uses data labels for supervised training. Its loss function is in represents the output probability distribution of the i-th image after it is input into the feature extractor of the pruning model and then into the classification head of the pruning model, y i is the label of the i-th image, and CE(·) is the cross entropy function.

[0033] In summary, the loss function for a single image is:

[0034] Where α and β are balance coefficients. The pre-trained model is frozen, and the trainable model is the pruned model. Finally, the system obtains the pruned model for the small data set.

[0035] The model pruning system based on weighted attention alignment described in the present invention utilizes the model pruning method to prune the model to obtain a pruned model for a small data set.

[0036] The model pruning method and system based on weighted attention alignment described in the present invention have the following advantages:

[0037] (1) The weighted attention map calculation method combines pruning and migration by using the batch normalization layer parameter to quantify channel importance, which increases the interpretability of the pruning-migration task. On the other hand, the channel fusion improves the attention map and successfully solves the problem of feature alignment with inconsistent channel numbers, making it feasible to perform pruning and migration tasks simultaneously.

[0038] (2) The weighted attention alignment training mechanism successfully realizes the joint training of pruning and transfer, solving the problem of high cost of step-by-step and iterative training in existing methods. It not only improves the performance of pruning by using pre-trained knowledge, but also significantly reduces the time cost of the training process.

[0039] For those skilled in the art, various other corresponding changes and deformations can be made according to the technical solutions and concepts described above, and all of these changes and deformations should fall within the protection scope of the claims of the present invention.

Claims

1. A model pruning method based on weighted attention alignment, characterized in that: The following steps are involved: S1. Pre-fine-tuning: First, use the pre-trained model parameters to initialize the target model, and then use the data of the target domain to perform L1 regularization fine-tuning on the batch normalization layer parameters several times; S2. Pre-pruning: globally sort the absolute values ​​of the batch normalization layer parameters of the trained target model, prune the channels at the end, and obtain the pre-pruned model; S3. Model pruning training based on weighted attention alignment: Prune and fine-tune the pruned model, introduce the guidance of the pre-trained model through the weighted attention alignment mechanism, and obtain the pruned model for the target domain dataset.

2. According to claim 1, a model pruning method based on weighted attention alignment is characterized in that: In the weighted attention alignment mechanism, the calculation steps of the weighted attention map are as follows: Get the batch normalization layer parameters of the convolutional neural network Where C represents the number of channels, and the intermediate tensor output by the activation layer after the batch normalization layer is obtained Where H and W represent the height and width of the feature map respectively; The weighted attention map is calculated using the function Obtaining includes the following sub-steps: S31. Perform softmax normalization on the batch normalization layer parameter w to obtain the weight vector S32. Pooling the intermediate tensor Where A represents the size of the feature map after pooling; S33. The weight vector And the intermediate tensor after pooling Multiply by channel to get the weighted tensor S34. Weighted tensor summed by channel in Representing a tensor The forward slice of ; then the weighted attention map is obtained after Frobenius normalization where F i,j Represents the element in the i-th row and j-th column of the matrix.

3. According to claim 2, a model pruning method based on weighted attention alignment is characterized in that: In the step S3, a weighted attention alignment module, an output feature alignment module and a target domain perception training module are set; The weighted attention alignment module uses the weighted attention alignment mechanism to obtain the weighted attention graphs of the pre-trained model and the pruned model, and aligns the weighted attention graphs of the two at corresponding positions, and uses the knowledge of the pre-trained model to guide the fine-tuning training of the pruned model to improve the pruning performance; for a network with a total of B network structure blocks, the loss function is in denote the parameters of the batch normalization layer in the kth network structure block of the pre-trained model and the pruned model, respectively. denote the output feature tensor of the kth network structure block of the i-th image input pre-trained model and pruned model, respectively. WAM(·,·) is the weighted attention map function. represents the square of the Frobenius norm; The output feature alignment module replaces the classification head of the pruned model with the classification head of the pre-trained model, and then performs probability distribution alignment; the loss function is Where P i s,s , Denote the output probability distribution of the i-th image after it is input into the feature extractor of the pre-trained model and the pruning model, and then input into the classification head of the pre-trained model, D KL (·) is the KL divergence function; The target domain-aware training module fine-tunes the pruned model on a small data set and uses data labels for supervised training; The loss function is in represents the output probability distribution of the i-th image after it is input into the feature extractor of the pruning model and then into the classification head of the pruning model, y i is the label of the i-th image, CE(·) is the cross entropy function; The loss function for a single image is: Among them, α and β are balance coefficients; Finally, we get a pruned model for a small dataset.

4. A model pruning system based on weighted attention alignment according to claim 1, characterized in that: The model is pruned using the model pruning method as described in any one of claims 1 to 3 to obtain a pruned model for a small data set.

Citation Information

Patent Citations

  • Convolutional neural network channel pruning method based on model fine tuning

    CN111931914A

  • Multi-modal large model training optimization method and device in electric power vertical field

    CN118643470A

  • Method for realizing a multi-channel convolutional recurrent neural network EEG emotion recognition model using transfer learning

    US20230039900A1

  • Data processing method and apparatus

    WO2024245061A1