Compression method of visual Transform model
By evaluating the importance of patch tokens of the visual Transformer model, dynamic sampling and fusion, the model calculation complexity is reduced, the problem of high computing resource consumption is solved, and the model performance and universality are maintained.
Patent Information
- Application Number
- CN202510182409.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The computational complexity of the visual Transformer model is high, resulting in huge consumption of computing resources during training and inference, and the existing optimization methods are insufficient in generality and deployment difficulty.
Reduce computational complexity by compressing the patch token of Transformer Encoder. Specific methods include calculating important scores of patch tokens, dynamic sampling and fusion, and generating compressed image embedding.
It significantly reduces the computing resources required for model operation, maintains the model's high-precision processing capabilities, ensures the accuracy of image processing results, and has high versatility.
Smart Images

Figure CN120124699A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and more specifically, relates to a method for compressing a Vision Transformer model. Background Art
[0002] With the rapid development of computer vision, deep learning has achieved remarkable success in tasks such as image classification, object detection, and semantic segmentation. Among them, Convolutional Neural Networks (CNNs) have always been the mainstream method for traditional vision tasks due to their local receptive fields and parameter sharing mechanisms. However, the emergence of Vision Transformers (ViTs) marks a paradigm shift in the field of computer vision. ViTs utilize the self-attention mechanism to capture global features in images, demonstrating powerful feature modeling capabilities, and their performance even exceeds that of traditional CNN models in multiple tasks.
[0003] The advantages of ViTs lie in their flexible feature expression ability and global information modeling ability. By dividing the input image into fixed-size image patches and then embedding these patches, ViTs serialize the image data into a word sequence similar to that in natural language processing tasks. In this way, ViTs can capture the global correlations between all pixels in the image through the self-attention mechanism, thereby generating richer context feature representations. This method breaks the limitation of CNNs relying on local receptive fields and shows better performance especially in complex scenarios or tasks that require global perception.
[0004] However, behind this performance improvement is accompanied by a huge computational cost. The computational complexity of the ViTs model mainly comes from its self-attention mechanism, which needs to calculate the similarity between all input tokens. Specifically, for an input sequence containing N tokens, the computational complexity of the self-attention mechanism is O(N 2 ²). When the image resolution is high, the number of tokens increases rapidly, resulting in the model consuming a large amount of computational resources during training and inference. In addition, ViTs models usually contain tens of millions to hundreds of millions of parameters. For example, a typical ViT-Base model contains 86M parameters and requires large-scale datasets (such as ImageNet) to support training.
[0005] In response to these challenges, researchers have proposed various optimization strategies to reduce the computational complexity of ViTs.
[0006] First, efficient architecture design is one of the main means to reduce the computational requirements of ViTs. This approach reduces the computational volume while maximizing model performance by redesigning the basic structure of the model or introducing optimization mechanisms. Its core idea is to allocate computational resources in a more economical way, thus achieving a good trade-off between performance and efficiency. Swin Transformer significantly reduces the computational complexity of high-resolution images through a hierarchical structure and a shifted window mechanism. Its method restricts self-attention calculations to local windows and realizes interactions between windows through window shifting, which not only enhances the local feature modeling ability but also retains the global long-range dependence information. EfficientViT combines a multi-scale attention mechanism and lightweight convolution to optimize model efficiency. It reduces the complexity from quadratic to linear by restricting the scope of global attention calculations and embeds lightweight convolution to complement the deficiencies of ViTs in local feature extraction, thereby capturing local context information at a lower cost and further enhancing the expressive ability. These methods maximize the balance between performance and efficiency through architecture optimization, but they also suffer from the problem of insufficient generality. The designs are usually targeted at specific model structures and are difficult to simply transfer to other models, thus limiting their applicability in diverse application scenarios.
[0007] Second, model quantization techniques are also widely applied to the optimization of ViTs. Quantization effectively reduces the computational complexity and memory consumption by reducing the model parameters and calculations from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers). However, the application of quantization techniques also faces many limitations. For example, Quantization-Aware Training (QAT) improves the model's adaptability by simulating the effects of quantization during the training process, but this usually increases the training time and complexity. In addition, Post-Training Quantization (PTQ) avoids the additional training overhead, but it is often difficult to maintain the model performance at extremely low precision (such as 4 bits or below), so it has high requirements for the quality and quantity of calibration data. These limitations also pose challenges to the actual deployment of quantization techniques. Since the quantization precision and format supported by hardware platforms vary, the model needs to be optimized for specific hardware environments, which increases the deployment difficulty.
[0008] Token compression technology effectively reduces the computational complexity by decreasing the number of input tokens processed by ViTs models, which is an important direction for optimizing model efficiency. Its core lies in retaining key tokens based on the importance of input tokens, reducing unnecessary consumption of computing resources, and minimizing performance loss. For example, DynamicViT dynamically evaluates the importance scores of each token through a lightweight prediction module and discards low-score tokens, thereby adapting to different input characteristics while reducing complexity. However, the method of directly deleting tokens may lead to the loss of useful information, especially in fine-grained feature modeling. In addition, the introduction of the dynamic evaluation module also increases the computational overhead. To address these issues, methods such as Tome propose calculating the similarity between tokens, weighted-merging highly similar tokens, retaining more key information, and reducing the negative impact of direct deletion. Although this approach is an effective means of reducing computational costs, it still needs to balance efficiency and performance when dealing with tokens, and a more general compression strategy needs to be designed to adapt to diverse application scenarios. Summary of the Invention
[0009] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a compression method for visual Transformer models, which reduces the computational complexity of visual Transformer models, improves the inference speed, and simultaneously maintains the performance and generality of the models by compressing the patch tokens of the Transformer Encoder.
[0010] To achieve the above-mentioned invention purpose, the compression method for the visual Transformer model of the present invention includes the following steps:
[0011] S1: Denote the number of Transformer Blocks in the Transformer Encoder of the visual Transformer model as M. Select K Transformer Blocks from the M Transformer Blocks as the Blocks to be compressed according to actual needs, and denote the original serial number corresponding to the k-th Block to be compressed as m k , k = 1, 2,..., K;
[0012] S2: Input the input image into the embedding layer of the visual Transformer model. The embedding layer generates an image embedding x and inputs it into the Transformer Encoder. The method for generating the image embedding is as follows:
[0013] Divide the input image evenly into N patches, encode each patch to obtain an embedding vector n = 1, 2,..., N, D represents the dimension of the embedding vector. At the same time, encode the position of the patch to obtain a position vector e n , and stack them to obtain the patch representation vector z n=f n +e n , and then each patch represents the vector z n Image embedding constructed as patch tokens where z cls Represents a classification token;
[0014] S3: Transformer Encoder extracts features from image embeddings and then outputs them to subsequent modules for classification. During the feature extraction process, for the mth layer Transformer Block, when m≠m k , then no model compression is performed, when m = m k , the following method is used to compress the model:
[0015] S3.1: The image embedding received by the current Transformer Block is recorded as The 0th token Token for classification 1st to Nth m Patch tokens j=1,2,…,N m , N m Represents image embedding X m The number of patch tokens in the current Transformer Block; the value matrix of the multi-head self-attention module input where d v represents the dimension of the value, and the calculated attention matrix is The importance score of each patch token is calculated using the following formula:
[0016]
[0017] Among them, A 0,n Represents a classification token With patch tokens The attention score between Represents the value matrix V m Medium Patch Token The corresponding value vector, |||| means to find the norm;
[0018] S3.2: The image embedding obtained after processing by the multi-head self-attention module and the stacked normalization module is Set the compression ratio R according to the actual situation and calculate the number of patch token sampling H m =R×N m , according to the importance score of the patch token from N m Patch Tokens H mPatch tokens form a patch token queue P. The higher the importance score, the higher the probability of being selected;
[0019] S3.3: Set the fusion ratio T according to the actual situation, and calculate the number of fused tokens G m = T × H m , then divide the patch token queue P into two queues C 1 、C 2 , where queue C 1 contains G m patch tokens, and queue C 2 contains H m - G m patch tokens; traverse each patch token in queue C 1 , and filter out the patch token with the highest similarity to it from queue C 2 and delete the patch token from queue C 2 , then use the following formula for fusion to obtain the fused patch token:
[0020]
[0021] where, represents the g-th patch token in queue C 1 , g = 1, 2,..., G m , j g represents the original serial number of the patch token , represents the fused patch token, w g represents the weight, represents the patch token with the highest similarity to the patch token 2 in queue C , represents the original serial number of the patch token ;
[0022] Combine the classification tokens, the G m fused patch tokens and the remaining patch tokens in set C 2 to form an image embedding and output it to the subsequent module.
[0023] Compression method for the vision Transformer model of the present invention. According to actual needs, determine the block to be compressed in the Transformer Block of the Transformer Encoder of the vision Transformer model. Input the input image into the embedding layer of the vision Transformer model to generate an image embedding and then input it into the Transformer Encoder. When the Transformer Block is the block to be compressed, calculate the importance score of the patch tokens, sample and fuse the patch tokens processed by the multi-head self-attention module and the stacked normalization module according to the importance score to obtain the image after compressing the patch tokens, and then output it to the subsequent module.
[0024] The present invention has the following beneficial effects:
[0025] 1) By comprehensively considering the importance of the patch tokens of the vision transformer model, the present invention accurately identifies the patch tokens with low importance degree in the model inference process, effectively reduces redundant tokens without introducing additional trainable parameters, and maximally retains important information. Thus, without affecting the model accuracy, the model is maximally compressed, significantly reducing the computing resources required for model operation.
[0026] 2) Although the present invention reduces the number of patch information, by precisely controlling the selection and fusion of the patch tokens, it still maintains the high-precision processing ability of the model, ensuring the accuracy of the image processing result.
[0027] 3) The compression method designed by the present invention has high generality and can be applied to multiple models with only minor modifications and achieve good results. Description of the Drawings
[0028] Figure 1 is the flowchart of the specific implementation manner of the compression method for the vision Transformer model of the present invention;
[0029] Figure 2 is the flowchart of the compression of the patch tokens in the present invention;
[0030] Figure 3 is the structural diagram of the improved Transformer Encoder using the present invention;
[0031] Figure 4 is the schematic diagram of the processing process of the block to be compressed in this embodiment. Specific Implementation Manner
[0032] The specific embodiments of the present invention will be described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0033] Embodiment
[0034] Figure 1 is a flowchart of the specific implementation of the visual Transformer model compression method of the present invention. As Figure 1 shown, the specific steps of the visual Transformer model compression method of the present invention include:
[0035] S101: Determine the Transformer Block to be compressed:
[0036] Denote the number of Transformer Blocks in the Transformer Encoder of the visual Transformer model as M. Select K Transformer Blocks from the M Transformer Blocks as the Transformer Blocks to be compressed according to actual needs. Denote the original serial number corresponding to the kth Transformer Block to be compressed as m k , k = 1, 2,..., K.
[0037] In practical applications, it can be made that the Transformer Blocks to be compressed are evenly distributed among all Transformer Blocks. Therefore, the original serial number m k of the Transformer Blocks to be compressed can be calculated by the following formula:
[0038]
[0039] where denotes rounding down.
[0040] S102: Generate image embeddings:
[0041] Input the input image into the embedding layer of the visual Transformer model. The embedding layer generates the image embedding x and inputs it into the Transformer Encoder. The method for generating the image embedding is:
[0042] Evenly divide the input image into N patches. These patches are regarded as the basic units processed by the model, that is, "tokens". Encode each patch to obtain the embedding vector n = 1, 2,..., N, D represents the dimension of the embedding vector. At the same time, encode the position of the patch to obtain the position vector e n, the patch representation vectors z are obtained by superimposition n = f n + e n , and then each patch representation vector z n is used as a patch token to construct an image embedding where z cls represents a classification token.
[0043] In this embodiment, the input image uses an image sample in the ImageNet dataset, and the image resolution is adjusted to 224x224 pixels to adapt to the model input requirements. Then, the image is normalized, and the pixel values are scaled from [0, 255] to the [0, 1] interval, and data augmentation techniques such as random cropping and flipping are used to improve the generalization ability of the model. These normalization processes ensure the quality and consistency of the input image, laying a solid foundation for the training and testing of the model.
[0044] Common patch sizes are 16×16 pixels or 32×32 pixels. In this embodiment, the specification of 16×16 is used. For example, if the input image size is 224×224 and the size of each patch is 16×16, then the image will be divided into patches of size 16×16, and each patch will contain local information of the image.
[0045] In this embodiment, the patch encoding is implemented using a linear layer. First, each patch is flattened into a one-dimensional vector. When the patch size is 16×16 pixels and the image is an RGB image, the flattened vector of each patch is a 768-dimensional vector of 16×16×3. Then, the flattened patch is mapped through a linear layer to be converted into a D-dimensional vector. In this embodiment, D = 768. Since the Transformer model cannot automatically process the order information of elements in the sequence, position vectors need to be added to each patch. The position vectors provide the relative or absolute position of the patch in the image. These position encodings can be predefined and are generated using sine and cosine functions. Finally, the embedding vector of the patch and the position encoding are added together to obtain the final patch representation vector. To enhance the aggregation ability of global information, a special classification token ([CLS] token) also needs to be added at the very front of the input sequence, and the final input image embedding is constructed from the classification token and the patch tokens (patch representation vectors).
[0046] S103: Patch token compression:
[0047] The Transformer Encoder extracts features from the image embedding and then outputs to subsequent modules for classification. During the feature extraction process, for the m-th Transformer Block, when m ≠ m k, no patch token compression is performed. When m = m k , patch token compression is performed. Figure 2 is the flowchart of patch token compression in the present invention. As Figure 2 shown, the specific steps of the model compression of the present invention include:
[0048] S201: Calculate the importance score of patch tokens:
[0049] To analyze the influence degree of each patch token on the final output and determine the priority of its retention or discard, it is necessary to evaluate the importance of each patch token. The present invention uses the self-attention mechanism of the Transformer model to calculate the mutual relationship between each patch token and other patch tokens, so as to realize the importance evaluation of patch tokens. Each token input to the Vision Transformer model will generate an "attention weight matrix" through self-attention calculation with other tokens. These weights represent the dependence and correlation of each token with other tokens in the model. The self-attention mechanism can be expressed by the following formula:
[0050]
[0051] Among them, respectively represent the query matrix, the key matrix and the value matrix. The superscript T represents the transpose, and d k represents the dimension of the key, and d v represents the dimension of the value, and N represents the number of input features.
[0052] And among them, the calculation formula of the attention weight matrix A is as follows:
[0053]
[0054] Based on the above analysis, the calculation method of the importance score of patch tokens in the present invention is as follows:
[0055] Denote the image embedding received by the current Transformer Block as where the 0th token is the classification token The 1st to the N m th tokens are patch tokens j = 1, 2,..., N m , N m represents the number of patch tokens in the image embedding X m ; Denote the value matrix input to the attention module in the current Transformer Block as where d v represents the dimension of the value, and the calculated attention matrix is The importance score of each patch token is calculated using the following formula:
[0056]
[0057] Among them, A 0,n represents the classification token and the attention score between the patch tokens represents the value matrix V m in the patch token corresponding value vector (i.e., the j-th row vector), || || represents taking the norm.
[0058] S202: Patch token sampling:
[0059] Next, samples are taken according to the importance of the tokens, and the tokens to be retained are dynamically selected. The specific method is as follows: Denote the image embedding obtained after being processed by the multi-head self-attention module and the stacked normalization module in the current Transformer Block as Set the compression ratio R according to the actual situation, and calculate the patch token sampling quantity H m = R×N m , and sample H m from N m patch tokens to form a patch token queue P. The greater the importance score, the higher the probability of being selected.
[0060] In this embodiment, the quota sampling method is adopted for patch token sampling. The specific method is as follows:
[0061] Sort the N m patch tokens in descending order of importance score, and then divide the first H m patch tokens into the first group, and the last N m -H m patch tokens into the second group. Extract the first h m,1 patch tokens from the first group to form a queue Extract the first h m,2 patch tokens from the second group to form a queue where h m,1 = γ×H m , h m,2 = (1 - γ)H m , γ represents the preset quota ratio, γ > 0.5. Then merge the queues to obtain the patch token queue P.
[0062] In this embodiment, the set The insertion merging method is adopted during merging. The specific method is as follows: Insert the patch tokens in the queue into the queue starting from the beginning in the order of an arithmetic progression, where the common difference d is calculated using the following formula:
[0063]
[0064] where q 1 represents the preset insertion starting position, represents rounding down.
[0065] S203: Patch token fusion:
[0066] Set the fusion ratio T according to the actual situation, and calculate the number of fused tokens G m = T × H m , then divide the patch token queue P into two queues C 1 , C 2 , where the queue C 1 contains G m patch tokens, and the queue C 2 contains H m - G m patch tokens. Traverse each patch token in the queue C 1 , filter out the patch token with the highest similarity to it from the queue C 2 and delete the patch token from the queue C 2 , then use the following formula for fusion to obtain the fused patch token:
[0067]
[0068] where, represents the g-th patch token in the queue C 1 , g = 1, 2,..., G m , j g represents the original serial number of the patch token , represents the fused patch token, w g represents the weight, represents the patch token with the highest similarity to the patch token 2 in the queue C , represents the original serial number of the patch token .
[0069] Construct an image embedding from the classification token, the G m fused patch tokens and the remaining patch tokens in the set C 2 and output it to the subsequent module.
[0070] Figure 3 It is the structural diagram of the Transformer Encoder improved by the present invention. As Figure 3 shown, after being improved by the present invention, two modules, namely Token Sampling and Token Fusing, can be regarded as added in the original Transformer Encoder. Figure 4 It is the schematic diagram of the processing process of the Block to be compressed in this embodiment. As Figure 4 shown, in a certain layer of the Block to be compressed, after patch token sampling and fusion, the size of the output image embedding is smaller than that of the input image embedding. After passing through K Blocks to be compressed, the size of the image embedding can be greatly reduced, thus realizing the compression of the image embedding.
[0071] In order to better illustrate the technical effects of the present invention, specific examples are used to conduct experimental verification on the present invention.
[0072] In this embodiment, 4 existing model compression methods are selected for comparison, including:
[0073] DynamicViT, see the literature "Y. Rao, Z. Liu, W. Zhao, J. Zhou, and J. Lu, Dynamics spatial sparsification for efficient vision transformers and convolutional neural networks[C]. 2023";
[0074] Evo-vit, see the literature "Y. Xu, Z. Zhang, M. Zhang, K. Sheng, K. Li, W. Dong, L. Zhang, C. Xu, and X. Sun, Evo-vit: Slow-fast token evolution for dynamic vision transformer[C], 2022";
[0075] EViT, see the literature "Y. Liang, C. Ge, Z. Tong, Y. Song, J. Wang, and P. Xie, Not all patches are what you need: Expediting vision transformers via token reorganizations, [C]. 2022";
[0076] Ia-red^2, see the literature "B. Pan, R. Panda, Y. Jiang, Z. Wang, R. Feris, and A. Oliva, Ia-red^2: Interpretability-aware redundancy reduction for vision transformers[C]. 2021".
[0077] In this embodiment, experiments were carried out in terms of Top-1 Acc (model classification accuracy), Params (model parameter quantity), Flops (model computational volume), Throughput (model throughput), etc. Table 1 is a comparison table of evaluation indicators when the present invention and other existing methods are applied to the DeiT benchmark model.
[0078]
[0079] Table 1
[0080] As shown in Table 1, in this embodiment, the top-1 accuracy, FLOPs, and throughput of each model were tested. It can be seen from Table 1 that on the standard model DeiT-S, when the present invention is adopted and the computational cost is reduced by 37%, the top-1 precision only decreases by 0.1%. It is worth noting that on the standard model DeiT-B, the present invention can reduce the computational cost by 35% while keeping the top-1 precision unchanged, and increase the throughput by 1.52 times. At the same time, compared with other methods, the present invention achieves a good balance between model accuracy and model running speed.
[0081] In order to verify the effectiveness of token sampling and token fusion in the present invention, a comparative experiment was also carried out. The comparative methods include Average fusing and Max fusing, and the evaluation indicators include Top-1 Acc (model classification accuracy) and Flops (model computational volume). Table 2 is a comparison table of evaluation indicators when the token sampling and token fusion method and the comparative method in the present invention are applied to the DeiT benchmark model.
[0082]
[0083] Table 2
[0084] As shown in Table 2, compared with the original method, the method proposed in the present invention has achieved a significant improvement in the model classification accuracy index under the same computational volume.
[0085] Although the above description of the illustrative embodiments of the present invention has been given for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A visual Transformer model compression method, characterized in that: The following steps are involved: S1: Let the number of Transformer Blocks in the Transformer Encoder of the visual Transformer model be M. According to actual needs, select K from the M Transformer Blocks as the blocks to be compressed. Let the original serial number corresponding to the kth block to be compressed be m k , k=1,2,…,K; S2: The input image is input into the embedding layer of the visual Transformer model. The embedding layer generates the image embedding x and inputs it into the Transformer Encoder. The image embedding is generated as follows: The input image is evenly divided into N patches, and each patch is encoded to obtain an embedding vector n=1,2,…,N, where D represents the dimension of the embedding vector and the position of the patch is encoded to obtain the position vector e n , superimposed to obtain the patch representation vector z n =f n +e n , and then each patch represents the vector z n Image embedding constructed as patch tokens where z cls Represents a classification token; S3: Transformer Encoder extracts features from image embeddings and then outputs them to subsequent modules for classification. During the feature extraction process, for the mth layer Transformer Block, when m≠m k , then no model compression is performed, when m = m k , the following method is used to compress the model: S3.1: The image embedding received by the current Transformer Block is recorded as The 0th token Token for classification 1st to Nth m Patch tokens j=1,2,…,N m , N m Represents image embedding X m The number of patch tokens in the current Transformer Block; the value matrix of the multi-head self-attention module input where d v represents the dimension of the value, and the calculated attention matrix is The importance score of each patch token is calculated using the following formula: Among them, A 0,n Represents a classification token With patch tokens The attention score between Represents the value matrix V m Medium Patch Token The corresponding value vector, || || means to find the norm; S3.2: The image embedding obtained after being processed by the multi-head self-attention module and the stacked normalization module in the current Transformer Block is Set the compression ratio R according to the actual situation and calculate the number of patch token sampling H m =R×N m , according to the importance score of the patch token from N m Patch Tokens H m Patch tokens form a patch token queue P. The larger the importance score, the higher the probability of being selected. S3.3: Set the fusion ratio T according to the actual situation and calculate the number of fusion tokens G m =T×H m , then divide the patch token queue P into two queues C1 and C2, where queue C1 contains G m patch tokens, queue C2 contains H m -G m patch tokens; traverse each patch token in queue C1, select the patch token with the greatest similarity from queue C2 and delete the patch token from queue C2, and then use the following formula to fuse to obtain the fused patch token: in, represents the g-th patch token in queue C1, g = 1, 2, ..., G m , j g Represents a patch token The original serial number, represents the fused patch token, w g represents the weight, Indicates the patch token in queue C2 The patch token with the greatest similarity, Represents a patch token The original serial number; The classification tokens and the fused G m Patch Tokens The remaining patch tokens in set C2 constitute the image embedding and are output to subsequent modules.
2. The visual Transformer model compression method according to claim 1, characterized in that: The original sequence number m of the block to be compressed k The calculation is done using the following formula: in, Indicates rounding down.
3. The visual Transformer model compression method according to claim 1, characterized in that: In step S3.2, the quota sampling method is used to sample patch tokens. The specific method is: N m Patch Tokens Sort by importance score from largest to smallest, and then m patch tokens are divided into the first group, and the next N m -H m The patch tokens are divided into the second group, and the first h are extracted from the first group. m,1 Patch tokens form a queue Extract the first h from the second group m,2 Patch tokens form a queue where h m,1 =γ×H m ,h m,2 =(1-γ)H m , γ represents the preset quota ratio, γ>0.5; then the queue The merged patch token queue P is obtained.
4. The visual Transformer model compression method according to claim 3, characterized in that: The collection The insertion merge method is used when merging. The specific method is: The patch tokens in are inserted into the queue from the beginning in the order of arithmetic progression. The tolerance d is calculated using the following formula: Among them, q1 represents the preset insertion starting position, Indicates rounding down.
Citation Information
Patent Citations
Dynamic pruning method of visual Transform
CN116933859A
Visual Transform lightweight method and system based on token encapsulation and enhancement
CN117592525A
Efficient visual Transform method for aggregating semantic mark angles
CN118710964A
Efficient fine-grained image classification model based on improved ViT
CN118887445A
System and method for self-distilled vision transformer for domain generalization
US20240203098A1