Visual Transform lightweight method based on Token fusion
Through the visual Transformer lightweight method based on Token fusion, the redundant computing problem existing in the actual application of ViT model is solved, and the calculation redundancy is reduced while maintaining the model accuracy and improving the computing efficiency, which is suitable for the deployment of mobile devices.
Patent Information
- Application Number
- CN202411883110.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-13
AI Technical Summary
In practical applications, the ViT model has the problem of excessive redundant calculations leading to performance degradation, especially when deploying to mobile devices, it is difficult to achieve efficient computing.
The visual Transformer lightweight method based on token fusion is adopted. By segmenting the input image into multiple tokens and introducing Class Tokens, the Encoder module and multi-head attention mechanism in the Transformer model are used to extract the token information, calculate the interaction score between the tokens, and judge the tokens with low contribution for fusion processing, thereby reducing the calculation redundancy of the model.
While maintaining model accuracy, the calculation redundancy of the model is reduced through Token fusion, the computing efficiency is improved, and the ViT model can better adapt to the deployment needs of mobile devices.
Smart Images

Figure CN119992158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the reasoning acceleration of large models and related fields of deep learning, and in particular to a lightweight method for large visual Transformer models. Background Art
[0002] With the emergence and gradual development of the ViT model in the field of vision, it has gradually become a research focus. However, as the model has been widely used and the demand for its deployment has increased, researchers have begun to realize the challenges and limitations that the ViT model does have in practical applications. In order to overcome these challenges, researchers have begun to focus on lightweight optimization of the model to improve its efficiency and performance. By learning on different types of tasks, the ViT model is able to obtain rich knowledge representations, thereby demonstrating excellent performance in multiple visual tasks such as image classification, object detection, and behavior recognition, surpassing traditional models. The successful application of the ViT model not only proves its great potential in visual tasks, but also promotes the continuous optimization and improvement of model complexity and performance. This continuous research effort heralds new breakthroughs and progress in the field of vision in the future, providing more efficient and convenient solutions for solving complex visual tasks in the real world.
[0003] In the deployment and application of the ViT model, one of its biggest drawbacks is that too much redundant calculation leads to reduced performance, making it difficult to further improve in the field of vision. Therefore, this paper will focus on reducing unnecessary redundant calculations, that is, calculations between different Tokens, and explore methods to optimize the model effect, so as to reduce redundant calculations while maintaining good performance, so that the model can be easily deployed on mobile devices. In response to the challenge of computational redundancy in the ViT model, researchers are committed to finding effective methods to evaluate the degree of correlation between the tokens of the model to improve the efficiency and scalability of the model. By carefully designing the model architecture and embedding new modules, researchers try to reduce unnecessary calculations of the model without sacrificing performance. This optimization of model complexity aims to enable the ViT model to better adapt to actual application needs and further expand its application scope in visual tasks. By reducing the interaction of additional Tokens, the ViT model is expected to improve its deployment efficiency and application flexibility in the real world while maintaining superior performance. This series of research and exploration provides important guidance and direction for the continuous improvement and development of the ViT model. In summary, this chapter summarizes the research issues of the lightweight ViT model as follows:
[0004] (1) Computational redundancy problem: The global interaction characteristics of the model lead to redundant calculations between different tokens with small correlations;
[0005] (2) Structural optimization problem: The model’s multi-head attention structure may lead to overfitting in the deep extraction of information;
[0006] (3) Model deployment problem: Unnecessarily large calculations make it difficult to deploy the model on mobile devices. Summary of the invention
[0007] Purpose of the invention: The purpose of the present invention is to solve the problem that the current model cannot be deployed on a larger scale due to the large number of parameters.
[0008] The purpose of the present invention is to realize a visual Transformer lightweight method based on Token fusion through the following technical scheme, comprising the following steps:
[0009] S1: For the TF-ViT model, the input image is segmented into multiple image blocks according to the ViT model, multiple tokens are obtained, and the Class Token is introduced;
[0010] S2: Through the Encoder module in the Transformer model, the Token is input and the multi-head attention mechanism is used to extract the information of each Token and calculate the interaction score between each Token and other Tokens;
[0011] S3: Calculate the contribution of each Token according to the interaction score, determine the Token with low contribution, and merge it;
[0012] S4: in the fusion process, performing complexity estimation on the ViT model and the TF-ViT model;
[0013] S5: While maintaining the accuracy of the model, reduce the computational redundancy of the model and improve the computational efficiency through Token fusion.
[0014] Optionally, the input image in S1 is assumed to be X×X. According to the ViT model, the image is cut into W×H image blocks, where W=H. At this time, the number of image blocks is:
[0015]
[0016] After the image block passes through the linear layer, a series of tokens are obtained. The linear layer is a layer in the neural network. It performs a linear transformation on the input through a weight matrix and a bias term. At this time, the Class Token is introduced. The ClassToken is the result of all Token processing. The number of tokens at this time is:
[0017]
[0018] Optionally, the Token in S2 enters the Encoder module of the first-layer Transformer model, and first uses multi-head attention to extract the Token information and fuse it into the Class Token. The multi-head attention mechanism is a core component of the Transformer model. In the multi-head attention mechanism, each Token is associated with other Tokens, and the degree of attention of each Token to other Tokens is determined by calculating the attention score.
[0019] Optionally, the interaction score between each Token in S2 and other Tokens:
[0020] Attention i =[a i,1 ,a i,2 ,...,a i,i ,...,a i,M ]
[0021] where a i,M represents the interaction score between M Tokens and the i-th Token, and also represents the contribution of M-1 Tokens to the i-th Token in this single attention. And so on. i Addition:
[0022]
[0023] Where ai is the sum of the attention scores of the i-th Token on M-1 Tokens in a single attention.
[0024] Alternatively, if a token’s attention scores for other M-1 tokens are generally not high, we cannot directly assume that its contribution is low. Considering that this means that in the current layer of the self-attention mechanism, the token does not form a strong association with other tokens, this may indicate that the information contained in the token is not particularly important in the current context, or that other tokens already contain enough information to make the additional information of the token less critical. Therefore, considering multi-head attention, assuming the number of heads is N, we get:
[0025]
[0026] Optionally, the interaction score calculates the contribution of each Token by adding the attention scores of the i-th Token and other M-1 Tokens from different heads, and reorganizing the scores to obtain:
[0027] Score=[s1,s2,s3,...,s M ]
[0028] where s1≥s2≥s3,...≥s M , that is, for S core Rearrange the values in and sort them in descending order.
[0029] Optionally, the comprehensive contribution of each Token to other Tokens in different self-attention heads is judged according to the score value in Score, and the model complexity can be reduced and the convergence speed can be improved by integrating Tokens with low comprehensive contribution.
[0030] Optionally, the complexity estimation of the ViT model and the TF-ViT model in S4 refers to the complexity estimation of S core The complexity of the model is proportional to the square of the Token. For the original ViT model, the complexity can be estimated as:
[0031]
[0032] Where k is the coefficient of complexity, O is the approximate representation in the computer, and for the TF-ViT model, its complexity is approximately:
[0033]
[0034] Optionally, the complexity of the TF-ViT model fusion varies with the size of the window, and as the window becomes larger, the complexity of the TF-ViT model becomes lower and the computational redundancy becomes lower.
[0035] Compared with the prior art, the present invention has the following advantages:
[0036] (1) It adopts the token fusion method, which preserves the integrity of image information as much as possible and reduces useless calculations compared with the existing technology;
[0037] (2) Combining dynamic and static lightweight methods, not only dynamically integrates tokens, but also uses structured pruning technology to improve the efficiency and performance of the model;
[0038] (3) The TF-ViT model can be embedded into other models as an independent module to maximize its potential and provide practical possibilities for the deployment of the model on mobile terminals. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the model framework of the present invention;
[0040] Figure 2 It is the process of gradually removing image blocks with low contribution. DETAILED DESCRIPTION
[0041] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.
[0042] In this invention, we mainly focus on large-scale visual Transformer models in the image field, such as Figure 1 As shown, the present invention comprises the following steps:
[0043] Step 1: Input image segmentation and token generation
[0044] First, the input image will be cut in the standard way of the Visual Transformer (ViT) model. The input image Input is assumed to be X×X. According to the ViT model, the image is cut into W×H (where W=H) image blocks. At this time, the number of image blocks is:
[0045]
[0046] Each image block is flattened into a one-dimensional vector and transformed through a linear layer. The linear layer maps the features of each image block through the weight matrix and bias term and converts them into corresponding tokens. These tokens contain the features of the local area of the image. After the introduction of Class Token, Class Token is a special token that represents the information of the entire image. It is input into the Transformer model together with other tokens for the final classification task. The number of tokens at this time is:
[0047]
[0048] Step 2: Extract information through the Encoder module and multi-head attention mechanism
[0049] The obtained Token enters the Encoder module of the first-layer Transformer model. The Encoder module is composed of multiple stacked layers. Each layer includes a multi-head attention mechanism and a feedforward neural network. First, the multi-head attention is used to extract the Token information and fuse it into the Class Token. By calculating multiple "attention heads" in parallel, the model can learn feature information at different levels from multiple subspaces. In each attention head, the model calculates the interaction score of each Token with other Tokens, and calculates the interaction score of each Token with other M-1 Tokens:
[0050] Attention i =[a i,1 ,a i,2 ,...,ai,i ,...,a i,M ]
[0051] where a i,M represents the interaction score between M Tokens and the i-th Token, and also represents the contribution of M-1 Tokens to the i-th Token in this single attention. And so on. i Addition:
[0052]
[0053] Where ai is the sum of the attention scores of the i-th Token on M-1 Tokens in a single attention. If a Token's attention scores on other M-1 Tokens are generally not high, it cannot be directly considered that its contribution is low. Consider that this means that in the current layer of the self-attention mechanism, the Token does not form a strong association with other Tokens. This may indicate that the information contained in the Token is not particularly important in the current context, or that other Tokens already contain enough information to make the additional information of the Token less critical. Therefore, considering multi-head attention, assuming the number of heads is N, we get:
[0054]
[0055] Step 3: Calculate the contribution of each token and merge them
[0056] Add the attention scores of the i-th Token and other M-1 Tokens from different heads and reorganize the scores to get:
[0057] Score=[s1,s2,s3,...,s M ]
[0058] where s1≥s2≥s3,...≥s M , that is, reorganize the values in the above formula and sort them in descending order. This paper uses the score value of the above formula to judge the comprehensive contribution of each Token to other M-1 Tokens in different self-attention heads. The Token with lower contribution means that its information is less important in the current context. Figure 2 As shown in the figure, by integrating low-contribution tokens, unnecessary redundant information in the network is reduced and the computational efficiency of the model is optimized.
[0059] Step 4: Complexity estimation
[0060] Assume that for the above S coreThe complexity of the model is proportional to the square of the Token. For the original ViT model, the complexity can be estimated as:
[0061]
[0062] Where k is the complexity coefficient. For the TF-ViT model, its complexity is approximately:
[0063]
[0064] Step 5: Maintain model accuracy while improving computational efficiency
[0065] Finally, ensure that the accuracy of the model is not lost while reducing redundant calculations. Comparing the above formula, we can see that the complexity of the model changes with the size of the window, and as the window becomes larger, the complexity of the TF-ViT model becomes lower, and the computational redundancy becomes lower. According to task requirements and computing resources, the size of the fusion window is dynamically adjusted through regularization technology to ensure that the performance is not significantly affected while reducing calculations.
[0066] The above-described embodiments merely express the implementation methods of the present invention, but they cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A visual Transformer lightweight method based on Token fusion, characterized in that: The following steps are involved: S1: For the TF-ViT model, the input image is segmented into multiple image blocks according to the ViT model, multiple tokens are obtained, and the Class Token is introduced; S2: Through the Encoder module in the Transformer model, the Token is input and the multi-head attention mechanism is used to extract the information of each Token and calculate the interaction score between each Token and other Tokens; S3: Calculate the contribution of each Token according to the interaction score, determine the Token with low contribution, and merge it; S4: in the fusion process, performing complexity estimation on the ViT model and the TF-ViT model; S5: While maintaining the accuracy of the model, reduce the computational redundancy of the model and improve the computational efficiency through Token fusion.
2. According to the Token fusion-based visual Transformer lightweight method of claim 1, it is characterized in that: The input image described in S1 is assumed to be X×X. According to the ViT model, the image is cut into W×H image blocks, where W=H. At this time, the number of image blocks is: After the image block passes through the linear layer, a series of tokens are obtained. The linear layer is a layer in the neural network. It performs a linear transformation on the input through a weight matrix and a bias term. At this time, the Class Token is introduced. The Class Token is the summary of the results of all Token processing. At this time, the number of tokens is:
3. According to claim 1, a visual Transformer lightweight method based on Token fusion is characterized in that: The token described in S2 enters the Encoder module of the first-layer Transformer model. First, the token information is extracted and fused into the Class Token using multi-head attention. The multi-head attention mechanism is the core component of the Transformer model. In the multi-head attention mechanism, each token is associated with other tokens, and the degree of attention of each token to other tokens is determined by calculating the attention score.
4. According to claim 1, a visual Transformer lightweight method based on Token fusion is characterized in that: The interaction score between each token and other tokens described in S2: Attention i =[a i,1 ,a i,2 ,...,a i,i ,...,a i,M ] where a i,M represents the interaction score between M Tokens and the i-th Token, and also represents the contribution of M-1 Tokens to the i-th Token in this single attention. And so on. i Addition: Where ai is the sum of the attention scores of the i-th Token on M-1 Tokens in a single attention.
5. According to claim 4, a visual Transformer lightweight method based on Token fusion is characterized in that: If a token’s attention scores for other M-1 tokens are generally not high, it cannot be directly considered that its contribution is low. Considering that this means that in the current layer of the self-attention mechanism, the token does not form a strong association with other tokens, this may indicate that the information contained in the token is not particularly important in the current context, or that other tokens already contain enough information to make the additional information of the token less critical. Therefore, considering multi-head attention, assuming the number of heads is N, we get:
6. According to claim 1, a visual Transformer lightweight method based on Token fusion is characterized in that: The interaction score calculates the contribution of each Token by adding the attention scores of the ith Token and other M-1 Tokens from different heads, and reorganizing the scores to obtain: Score=[s1,s2,s3,...,s M ] where s1≥s2≥s3,...≥s M , that is, reorganize the values in Score and sort them in descending order.
7. According to claim 6, a visual Transformer lightweight method based on Token fusion is characterized in that: The score value in Score is used to determine the comprehensive contribution of each Token to other Tokens in different self-attention heads. By integrating Tokens with low comprehensive contributions, the model complexity can be reduced and the convergence speed can be improved.
8. According to claim 1, a visual Transformer lightweight method based on Token fusion is characterized in that: The complexity estimation of the ViT model and TF-ViT model described in S4 refers to the fusion of the last B scores of Score, where B refers to the size of the fusion window. The complexity of the model is proportional to the square of the Token. For the original ViT model, the complexity can be estimated as: Where k is the coefficient of complexity, O is the approximate representation in the computer, and for the TF-ViT model, its complexity is approximately:
9. According to claim 1, a visual Transformer lightweight method based on Token fusion is characterized in that: The complexity of the TF-ViT model varies with the size of the fusion window, and as the window becomes larger, the complexity of the TF-ViT model becomes lower and the computational redundancy becomes lower.
Citation Information
Cited By
Method, device, medium, product and system for reasoning acceleration of pre-training model
CN120408126A
Inference acceleration method, device, medium, product and system for pre-trained models
CN120408126B