Small sample fine-grained block-level model pruning method for visual Transform model

By subdividing the Transformer layer of the visual Transformer model into multi-head self-attention and multi-layer perceptron blocks, and gradually pruning combined with fine-tuning recovery ability evaluation indicators, the problem of low pruning efficiency of the visual Transformer model in small sample compression scenarios is solved, and efficient deployment on resource-constrained devices is achieved.

CN120068978APending Publication Date: 2025-05-30UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510225625.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing visual Transformer models are difficult to achieve efficient pruning in small sample compression scenarios, resulting in inefficient deployment on resource-constrained devices.

Method used

A fine-grained block-level model pruning method is proposed. By subdividing the Transformer layer into multi-head self-attention (MSA) blocks and multi-layer perceptron (MLP) blocks, combining fine-tuning recovery ability evaluation indicators, gradually pruning the least important blocks, and ensuring the coherence of model performance through an integrated pruning and fine-tuning framework.

Benefits of technology

The best model pruning effect is achieved under small sample conditions, improving the deployment efficiency and performance of the model in resource-constrained scenarios such as mobile devices and embedded systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068978A_ABST
    Figure CN120068978A_ABST
Patent Text Reader

Abstract

The invention provides a small sample fine granularity block level model pruning method aiming at a visual Transform model. Compared with a traditional structured model pruning method, the fine-grained block pruning method provided by the invention can more accurately identify and cut redundant parts in the model. According to the method, a Transform coding layer of a visual Transform (ViT) model is subdivided into fine-grained blocks, so that the fine processing of the model is realized. Under the condition of small samples, the accuracy of a pruning decision is further ensured by utilizing a pruning candidate block importance evaluation index based on the recovery capability, the model recovery capability and the calculation efficiency after candidate block pruning are measured, the block importance is comprehensively evaluated, accurate pruning is realized, and the precision and the efficiency are balanced. Further, a pruning and fine tuning integrated frame is constructed, candidate blocks are sequenced through the fine-tuned model performance, it is ensured that the pruning process is coherent, the performance loss is minimum, and the method adapts to the resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of algorithm models, and particularly relates to a small-sample fine-grained block-level model pruning method for a vision Transformer model. Background Art

[0002] To address the problem of scarce resources, many model compression techniques for vision Transformer (ViT) have been widely explored, such as quantization, distillation, and pruning.

[0003] Regarding ViT pruning, existing methods can be divided into two main categories. The first category is unstructured pruning represented by Token pruning, which aims to dynamically or statically delete or merge Tokens according to different inputs to accelerate the model. Since image data in vision tasks often contains a large amount of redundant information, and Tokens are the basic units for ViT to process images, the analysis and pruning of their redundancy are crucial for improving model efficiency. Although this method can maintain model performance, it is not ideal in terms of hardware utilization. Another common ViT pruning technique is structured pruning represented by channel pruning, which achieves compression by removing less important parameters in the ViT model. However, in small-sample compression scenarios, the granularity of channel pruning may be too coarse to achieve a high acceleration ratio on real-world devices such as GPUs.

[0004] In contrast, block pruning is a technique commonly used in the compression of convolutional neural networks and large language models (LLMs). It involves dividing the network into blocks according to the layer structure and performing pruning at the block level. Compared with channel pruning, the throughput of this method is much higher, so it is more beneficial to completely remove blocks. In convolutional neural networks, the block pruning method removes a complete residual block from the network structure. Generally speaking, block pruning is a relatively aggressive technique, and its potential risk lies in that it may directly delete important layers because it ignores the fine-grained architecture within each layer. The focus of the present invention is on how to perform fine-grained block-level redundancy extraction in the ViT model to improve pruning accuracy and achieve more efficient model compression without relying on other techniques.

[0005] It is worth noting that recent research progress in large language model pruning shows that jointly pruning the multi-head self-attention and multi-layer perceptron blocks usually outperforms pruning a single block alone. Considering the architectural similarity between large language models and ViT, it can be speculated that there is also fine-grained redundancy in ViT, and more effective pruning can be achieved by mining this redundancy. However, due to the significant differences in data representation, model structure, and task requirements between vision tasks and language tasks, directly migrating large language model pruning techniques to ViT poses certain challenges.

[0006] On the other hand, resource scarcity also includes practical challenges such as data privacy issues and rapid deployment requirements, which usually limit the acquisition of large-scale training datasets. Therefore, how to achieve efficient model pruning under data scarcity has become a key issue in improving the performance of models in practical applications. For many years, there has been little research on small-sample model compression, and existing methods focus on the compression of convolutional neural networks. Although there have been certain advancements in the pruning techniques of convolutional neural networks and large language models, the research on the compression and acceleration of ViT is still relatively limited. Existing pruning techniques, especially small-sample compression methods, are mainly designed for convolutional neural networks, and most methods do not fully consider the unique architectural characteristics of ViT. To address these challenges, this paper explores the structured pruning technique of ViT in small-sample scenarios, aiming to uncover the fine-grained redundancy in the ViT architecture, thereby improving its deployment efficiency on resource-constrained devices. Summary of the Invention

[0007] The present invention proposes a fine-grained block pruning method for small samples (Fine-grained Block Pruning with Tiny Sets for Vision Transformers, FBP-ViT), and its core technical points are as follows:

[0008] I. Definition of Fine-grained Block (FGB)

[0009] The novel concept of "fine-grained block" is introduced into the Vision Transformer (ViT) model for the first time. Each Transformer layer is subdivided into two fine-grained residual blocks, namely the multi-head self-attention (MSA) block and the multi-layer perceptron (MLP) block. Each block contains the corresponding functional modules and residual connections, and can independently process and enhance the input features. The MSA block is good at capturing global features and long-range dependencies, while the MLP block can effectively extract local features and perform non-linear transformations. Through the division of fine-grained blocks, the model can more accurately identify and process different types of features, improving the comprehensive capture ability of local and global information in the input data. This division method makes full use of the structural characteristics of the ViT model, laying a solid foundation for the pruning operation in the subsequent model optimization process, opening up a new exploration path, and is expected to promote the performance breakthrough of the ViT model in more complex task scenarios.

[0010] II. Importance Evaluation of Pruning Candidate Blocks

[0011] A pruning candidate block importance evaluation metric specifically designed for the ViT model is proposed. This metric is based on the model recovery ability of the candidate block after pruning, and measures the importance of the candidate block by the ability of the model to regain accuracy after fine-tuning on a small number of samples. The lower the score, the smaller the impact of the candidate block on the overall performance of the model, and it can be preferentially pruned and removed.

[0012] III. Effective Pruning under Data Scarcity

[0013] To ensure the computational efficiency and practicality of the pruning strategy under resource constraints, a fine-grained block pruning framework that tightly combines pruning and fine-tuning is constructed. The model is immediately fine-tuned after each pruning to evaluate the pruning effect and optimize the remaining model structure. This integrated process avoids the performance loss caused by the separation of pruning and fine-tuning in traditional pruning methods, ensuring the coherence and effectiveness of the pruning process. Through fine-grained pruning and timely fine-tuning, while maintaining the model performance, the computational resource consumption of the model can be significantly reduced, improving the practicality and deployment flexibility of the ViT model in resource-constrained scenarios such as mobile devices and embedded systems.

[0014] Benefiting from the above three designs, FBP-ViT achieves the best current model pruning effect in block-level pruning under small sample conditions, and under small sample conditions (even with only 50 samples), the performance of FBP-ViT is better than that of existing ViT pruning techniques using a relatively large dataset (1000 samples) in small samples.

[0015] To solve the above technical problems, the specific technical solution of a small sample fine-grained block-level model pruning method for a vision Transformer model of the present invention is as follows:

[0016] A small sample fine-grained block-level model pruning method for a vision Transformer model, characterized by including the following steps:

[0017] Step 1: Obtain an image dataset;

[0018] Step 2: Based on the vision Transformer model ViT, decompose each Transformer encoding layer in the model ViT into two fine-grained blocks: a multi-head self-attention MSA block and a multi-layer perceptron MLP block;

[0019] The model ViT reshapes the input image into a flattened sequence of blocks and linearly projects each block into tokens; for a model ViT with L layers, each layer is further subdivided to obtain 2L fine-grained blocks, and the output of each block is denoted as x l , l = 1,..., 2L;

[0020] Step 3: Importance evaluation of pruning candidate blocks;

[0021] Using the definition of fine-grained blocks, candidate models are obtained by removing a custom number of MSA or MLP blocks from the original ViT model;

[0022] Use fine-tuning recoverability to evaluate the ability of the candidate model after pruning to recover accuracy; calculate the speed improvement rate of the candidate model for one inference operation, and finally combine the fine-tuning recoverability and the speed improvement rate as the importance evaluation index of the pruning candidate blocks; the fine-tuning recoverability is calculated by the L2 distance:

[0023]

[0024] where FR(B i ) represents the fine-tuning recoverability, and respectively represent the sets of all tokens output by the original model and the candidate model, represents the small sample set, represents the second norm.

[0025] The speed improvement rate is obtained as follows: After removing block , when inputting an image, the speed improvement calculation formula for one inference operation of the corresponding candidate model is:

[0026]

[0027] where S(B i ) represents the speed improvement rate, represents the running speed of the original model for one inference operation,

[0028] represents the running speed of the corresponding candidate model for one inference operation;

[0029] The importance evaluation index of the pruning candidate blocks is calculated by the following formula:

[0030]

[0031] where I(B i ) represents the importance of block .

[0032] Step 4: After deleting each block, use the small sample dataset to fine-tune the adjacent blocks of the deleted block;

[0033] Perform iterative pruning search to gradually prune K blocks; in each iteration, remove each fine-grained block of the model in turn. After deleting each block, use a small-sample dataset to fine-tune the adjacent blocks of the deleted block; then evaluate the importance of the fine-grained blocks to generate the importance distribution of all blocks, thereby gradually discarding the least important blocks.

[0034] After pruning, a sequence of pruned blocks is obtained, denoted as and the corresponding pruned model It is constructed by recording the type and ID of each pruned block during the iteration process, where type specifies the type of the block, including MSA and MLP blocks, and ID represents the position of the block in the original model.

[0035] The visual Transformer model architecture consists of 1 embedding layer, L Transformer encoder layers, and 1 prediction head. Each Transformer encoder layer has the same structure and contains an MSA module and an MLP module. Both of these modules are after the normalization layer and are followed by a residual connection; after partitioning, the MSA block includes a normalization layer and a multi-head self-attention module, and the MLP block includes a normalization layer and a multi-layer perceptron module.

[0036] For the input token x of the k-th encoder layer k-1 , k = 1, …, L, the overall calculation process of the Transformer encoder layer is defined as:

[0037] x′ k = MSA(norm(x k-1 )) + x k-1

[0038] x k = MLP(norm(x′ k )) + x′ k

[0039] where MSA(·) and MLP(·) represent the MSA block and the MLP block respectively, and norm(·) represents the normalization operation; for a model V with L layers, each layer is further subdivided to obtain 2L fine-grained blocks, and the calculation process within each block can be expressed by the following formula:

[0040] x l = FGB(norm(x l-1 )) + x l-1

[0041] Here, l = 1, …, 2L, and FGB(·) represents the fine-grained block.

[0042] Compared with traditional structured model pruning methods, the fine-grained block pruning method proposed by the present invention can more accurately identify and remove redundant parts in the model. By subdividing the Transformer encoding layer of the Vision Transformer (ViT) model into multi-head self-attention (MSA) blocks and multi-layer perceptron (MLP) blocks, refined processing of the model is achieved. This fine-grained partitioning method enables pruning operations to identify fine-grained redundancies rather than blindly cutting entire layers, promoting more precise model optimization. By providing flexible pruning units, it ensures the accurate capture of local and global features, optimizes the pruning effect, and retains the core feature extraction ability.

[0043] Under small-sample conditions, using a pruning candidate block importance evaluation metric based on recovery ability further ensures the accuracy of pruning decisions. It measures the model's recovery ability and computational efficiency after pruning candidate blocks, comprehensively evaluates block importance, achieves precise pruning, and balances accuracy and efficiency. By measuring the importance of candidate blocks through performance changes after fine-tuning with a small number of samples, this method can effectively avoid mispruning key blocks, thus greatly reducing the number of model parameters and computational volume while maximizing the retention of model performance.

[0044] Furthermore, by constructing an integrated pruning and fine-tuning framework, candidate blocks are sorted based on the performance of the fine-tuned model, ensuring that the pruning process is coherent and the performance loss is minimized, adapting to resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 FIG. is a schematic diagram of the decomposition of the Transformer encoding layer into blocks;

[0046] Figure 2 FIG. is a schematic diagram of the fine-grained block pruning framework structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To better understand the purpose, structure, and function of the present invention, the following further describes in detail a small-sample fine-grained block-level model pruning method for a Vision Transformer model of the present invention with reference to the accompanying drawings.

[0048] For the structured pruning of vision Transformer models, the present invention proposes a fine-grained block-level pruning method using few-shot fine-tuning. This method consists of three parts: 1) Model structure analysis and block-level division. The structure of the vision Transformer (ViT) model is deeply analyzed, and each Transformer encoding layer is subdivided into two fine-grained blocks: the multi-head self-attention (MSA) block and the multi-layer perceptron (MLP) block. This division method makes full use of the structural characteristics of the ViT model and provides a basis for subsequent pruning operations. By decomposing the Transformer layer into smaller blocks, the pruning process can be more finely controlled, avoiding a large impact on the overall performance of the model by directly removing the entire layer. The candidate pruning blocks can be MSA blocks or MLP blocks, or a combination of these two blocks. In this way, flexible pruning of the model can be achieved to meet different compression requirements. 2) Importance evaluation of pruning candidate blocks. Use a recovery ability-based metric to accurately measure the importance of each block. 3) An integrated pruning and fine-tuning framework that combines the pruning and fine-tuning processes to ensure high performance of the model after compression.

[0049] In this method, the fine-grained block pruning strategy optimizes the compression process of the vision Transformer (ViT) model, and tightly combines the pruning operation with the model performance recovery through an integrated pruning and fine-tuning framework, which specifically includes the following steps:

[0050] Step 1: Obtain the IMAGENET-1K image dataset;

[0051] Step 2: Based on the vision Transformer (ViT) model, decompose each Transformer encoding layer in the model into two fine-grained blocks: the multi-head self-attention (MSA) block and the multi-layer perceptron (MLP) block; Specifically, a typical vision Transformer (ViT) model architecture consists of 1 embedding layer, L Transformer encoder layers, and 1 prediction head. Each Transformer encoder layer has the same structure and includes a multi-head self-attention (MSA) module and a multi-layer perceptron (MLP) module. Both of these modules are after the normalization (Norm) layer and are followed by a residual connection. Considering the existence of the residual connection, each Transformer encoder layer can be further subdivided into two fine-grained residual blocks: the MSA block and the MLP block. As Figure 1 shown, the MSA block includes a normalization layer and a multi-head self-attention module, and the MLP block includes a normalization layer and a multi-layer perceptron module.

[0052] The ViT model reshapes the input image into flattened sequence blocks and linearly projects each block into an embedding vector, i.e., a token. For the input token x of the k-th encoder layerk-1 , where \(k = 1,\ldots,L\), the computational process inside the entire Transformer encoder layer can be defined as:

[0053] x' k = MSA(norm(x k-1 )) + x k-1 ,

[0054] x k = MLP(norm(x' k )) + x' k .

[0055] Here, MSA(·) and MLP(·) represent the MSA block and the MLP block respectively, and norm(·) represents the normalization operation. For a model \(V\) with \(L\) layers, each layer is further divided to obtain \(2L\) fine-grained blocks, and the computational process within each block can be expressed by the following formula:

[0056] x l = FGB(norm(x l-1 )) + x l-1 .

[0057] Here, \(l = 1,\ldots,2L\), and FGB(·) represents the fine-grained block. From the perspective of mathematical operation logic, the MSA block and the MLP block can be treated equally during the pruning process. This partitioning method exploits the redundant relationship between the MSA block and the MLP block, thereby improving the overall performance of the model.

[0058] Step 3: Importance evaluation of pruning candidate blocks.

[0059] In fine-grained block pruning, it is crucial to evaluate the importance of each candidate block. Measuring the KL-divergence and L2 distance between the original model and the candidate model after pruning before and after deleting each block are two common practices; this method makes an improvement to the evaluation metric by using fine-tuning recoverability to measure the ability of the pruned model to recover accuracy.

[0060] Using the definition of the fine-grained block, candidate models can be obtained by removing a user-defined number of MSA or MLP blocks from the original model \(V\)

[0061] The concept of fine-tuning recoverability is defined as follows:

[0062]

[0063] where \(FR(B i )\) represents fine-tuning recoverability, and represent the sets of all tokens output by the original model and the candidate model respectively. Denotes a small sample set consisting of training samples from a small number of IMAGENET-1K image datasets for fine-tuning. Denotes the square of the Frobenius norm, i.e., the second norm, L2 distance.

[0064] When removing blocks, it must be recognized that the removal of different blocks will result in different running speeds of the corresponding candidate models. When removing blocks After that, when inputting 64 color pictures of size 224×224, the calculation formula for the speed improvement of the corresponding candidate model for one inference operation is:

[0065]

[0066] S(B i ) represents the speed improvement rate, represents the running speed of the original model for one inference operation, represents the running speed of the corresponding candidate model for one inference operation.

[0067] Under the condition of the same fine-tuning recoverability, blocks with lower speed improvement should be removed preferentially. To balance performance and speed improvement, a new metric - importance is introduced for each block The definition is as follows:

[0068]

[0069] where I(B i ) represents the importance of the block . This metric comprehensively evaluates the performance of the model, considering various factors that affect the overall accuracy and efficiency of the model. The block with the lowest importance indicates that its removal has the least impact on performance, meaning that this block contributes less to the overall effect of the model.

[0070] Step 4: Integrated pruning and fine-tuning framework

[0071] After deleting each block, use a small sample dataset consisting of training samples from a small number of IMAGENET-1K image datasets to fine-tune the adjacent blocks of the deleted block, rather than directly fine-tuning all parts of the entire model comprehensively. This strategy aims to reduce the computational time required for gradient backpropagation while restoring the model performance.

[0072] The cumulative error caused by the removal of multiple blocks may grow non-linearly. As the number of pruned blocks increases, the context background information will be significantly lost, leading to an obvious decline in the model performance. As the adjacent blocks are adjusted to make up for the modifications made during the pruning process, the iterative process ensures the integrity of the model. The overall framework is as Figure 2As shown, the ratio of the colored area in the block to the entire block reflects the importance of the corresponding block. By iteratively evaluating the importance of each block (MSA or MLP), the importance distribution of all blocks is generated, and then the least important blocks are gradually discarded. Then, the adjacent blocks of the pruned blocks are fine-tuned, as Figure 2 Only the blocks with the flame logo are trainable, and the blocks with the snowflake logo are frozen blocks that do not participate in the gradient descent process of training.

[0073] Perform iterative pruning search and gradually prune K blocks. In each iteration, each fine-grained block of the model is removed in turn, and a small sample training set is used to fine-tune the corresponding candidate model, and then evaluate the importance of the fine-grained blocks. By strictly analyzing the contribution of the blocks, the impact on the overall performance of the model is evaluated, thus further improving the effectiveness of the pruning strategy. After pruning, a sequence of pruned blocks is obtained, denoted as and the corresponding pruned model It is gradually constructed by recording the type and ID of each pruned block during the iteration process, where type specifies the type of the block (such as MLP, MSA), and ID represents the position of the block in the original model.

[0074] The finally obtained model is applied to image classification.

[0075] This method has been extensively experimented on multiple standard image classification datasets. The results show that this method has reached the current known state-of-the-art level in the field of small sample pruning and has demonstrated excellent convergence speed in the compression and optimization of the Vision Transformer (ViT) model.

[0076] IMAGENET-1K: A large-scale image classification dataset containing over 1,000 categories of annotated images, with approximately 1,000 images per category, totaling over 1 million annotated images. These images cover a variety of natural scenes, objects, animals, etc., and are widely used for image classification and model training benchmark tests in computer vision tasks.

[0077] CIFAR-10 & CIFAR-100: Two commonly used computer vision datasets mainly used for image classification tasks. The CIFAR-10 dataset contains 60,000 32×32-sized color images of 10 categories, divided into 50,000 training images and 10,000 test images. The CIFAR-100 dataset is similar to CIFAR-10 but contains 100 more fine-grained categories and also has 60,000 32×32-sized color images.

[0078] Flowers: An image dataset for flower classification, containing 1360 images of 17 types of flowers, with 80 images for each type. These images have significant variations in shape, scale, and perspective, and some types of flowers are difficult to distinguish in terms of color, shape, or texture, increasing the challenge of the classification task. This dataset covers a wide range of fields from natural scenes to artificial environments, providing rich data resources for training and evaluating image classification models.

[0079] Evaluation Metrics:

[0080] Top-1 Accuracy: The proportion of correct predictions made by the model on the test set, which is the main metric for measuring the classification performance of the model.

[0081] Speedup: The multiple of the computational speed improvement of the pruned model, used to evaluate the computational efficiency of the model.

[0082] Comparison Methods:

[0083] Original: The original ViT model without pruning.

[0084] EViT: An existing ViT pruning method that prunes by dynamically selecting important tokens.

[0085] S 2 ViTE: A ViT pruning method based on sparse training that prunes by removing tokens and attention heads.

[0086] SPViT: A ViT pruning method that prunes tokens by clustering.

[0087] GOHSP: A ViT pruning method based on graph and optimization for heterogeneous structured pruning that measures the importance of attention heads by graph ranking and imposes a structured sparsity pattern through an optimization process.

[0088] UP: A ViT pruning method that prunes by evaluating the L1 distance of each component.

[0089] LPViT: A ViT pruning technique using a low-power semi-structured pruning method that focuses on optimizing the multi-head self-attention module and the feed-forward network, achieving model compression and efficiency improvement by introducing convolutional layers and adjusting the MLP expansion ratio.

[0090] Table 1. Comparison of Top-1 validation accuracy results of various ViT prunings on ImageNet-1k using 1000 training samples

[0091]

[0092] Table 1 shows the result comparison of this method with different types of baseline methods in small-sample fine-tuning. On the ImageNet-1K dataset, FBP-ViT (this method) always maintains the best performance. This indicates that the FBP-ViT obtained by this method not only improves in speed but also significantly outperforms other baseline methods in terms of accuracy when dealing with the pruning task of the Vision Transformer (ViT) model in data-scarce scenarios.

[0093] Table 2. Top-1 validation accuracy (%) on ImageNet-1k after pruning DeiT-BASE (Data-efficient Image Transformer) using 50, 100, 500, and 1000 training samples

[0094]

[0095]

[0096] Table 2 shows the comparison of this method with traditional block-level pruning methods in terms of running speed and Top-1 accuracy. It can be seen that this method not only improves in running speed compared to traditional block-level pruning methods but also has an approximate 10% increase in Top-1 accuracy.

[0097] Table 3. Performance of FBP-ViT on out-of-domain images after removing 4 fine-grained blocks from DeiT-Base

[0098] Model CIFAR-10 CIFAR-100 Flowers DeiT-Base 99.1 90.8 98.4 FBP-ViT 97.8 87.6 98.9

[0099] Table 3 lists the Top-1 accuracy results of the original DeiT-Base model and the model pruned by FBP-ViT on out-of-domain datasets. Compared with the original model, the pruned model shows comparable performance in all out-of-domain image classification tasks and even performs better in some cases. This indicates that the model capabilities obtained on ImageNet through the fine-grained compression method are effectively retained in various downstream tasks.

[0100] It should be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A small sample fine-grained block-level model pruning method for a visual Transformer model, characterized in that: The following steps are involved: Step 1: Get the image dataset; Step 2: Based on the visual Transformer model ViT, each Transformer encoding layer in the model ViT is decomposed into two fine-grained blocks: multi-head self-attention MSA block and multi-layer perceptron MLP block; The model ViT reshapes the input image into a flattened sequence of blocks and linearly projects each block into a token; for a model ViT with L layers, each layer is further subdivided to obtain 2L fine-grained blocks, and the output of each block is represented as x l , l = 1, ..., 2L; Step 3: Importance evaluation of pruning candidate blocks; Using the definition of fine-grained blocks, candidate models are obtained by removing a custom number of MSA or MLP blocks from the original model ViT; Use fine-tuning recoverability to evaluate the ability of the candidate model to recover accuracy after pruning; Calculate the speed improvement rate of a candidate model for an inference operation, and finally combine fine-tuning recoverability and speed improvement rate as the importance evaluation index of pruning candidate blocks; Step 4: Perform iterative pruning search and gradually prune K blocks. In each iteration, remove each fine-grained block of the model in turn. After deleting each block, use a small sample dataset to fine-tune the adjacent blocks of the deleted block. Then the importance of fine-grained blocks is evaluated, and the importance distribution of all blocks is generated, so as to gradually discard the least important blocks; After pruning, a sequence of pruned blocks is obtained, recorded as And the corresponding pruning model It is constructed by recording the type and ID of each pruned block during the iteration process, where type specifies the type of block, including MSA and MLP blocks, and ID represents the position of the block in the original model.

2. According to claim 1, a small sample fine-grained block-level model pruning method for a visual Transformer model is characterized in that: The visual Transformer model architecture consists of 1 embedding layer, L Transformer encoder layers and 1 prediction head. Each Transformer encoder layer has the same structure and contains an MSA module and an MLP module. Both modules are after the normalization layer and followed by a residual connection. After blocking, the MSA block includes a normalization layer and a multi-head self-attention module, and the MLP block includes a normalization layer and a multi-layer perceptron module.

3. According to claim 2, a small sample fine-grained block-level model pruning method for a visual Transformer model is characterized in that: For the input token x of the kth encoder layer k-1 , k = 1, ..., L, the overall calculation process of the Transformer encoder layer is defined as: x′ k =MSA(norm(x k-1 ))+x k-1 x k =MLP(norm(x′ k ))+x′ k Among them, MSA(·) and MLP(·) represent MSA blocks and MLP blocks respectively, and norm(·) represents the normalization operation. For a model V with L layers, each layer is further subdivided to obtain 2L fine-grained blocks. The calculation process in each block can be expressed by the following formula: x l =FGB(norm(x l-1 ))+x l-1 Here l = 1, ..., 2L, and FGB(·) represents a fine-grained block.

4. According to claim 3, a small sample fine-grained block-level model pruning method for a visual Transformer model is characterized in that: The fine-tuning recoverability is calculated by the following formula: Among them, FR(B i ) indicates fine-tuning recoverability, and Represents the set of all tokens output by the original model and the candidate model respectively. represents a small sample set, represents the second norm; The speed increase is achieved by removing the block After that, when an image is input, the speed increase calculation formula for performing an inference operation on the corresponding candidate model is: Among them, S(B i ) represents the speed increase rate, Indicates the running speed of the original model for one inference operation, Indicates the running speed of one inference operation of the corresponding candidate model; The importance evaluation index of the pruning candidate block is calculated by the following formula: Among them, I(B i ) indicates a block The importance of.

Citation Information

Cited By

  • Transform model lightweight method and power transmission line image analysis method

    CN120910703A

  • Visual Transform accelerator based on grouped fine-grained structured pruning

    CN121009925A