Hardware-friendly visual Transform compression method based on quantization and Token pruning technology
By inserting the Token compression module into the visual Transformer model and performing token pruning, combining adaptive search to optimize the pruning parameters, the problems of large hardware overhead and high resource demand caused by floating-point operations in the existing technology are solved, and the efficient end-side deployment of visual Transformer is realized.
Patent Information
- Application Number
- CN202510177935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
During the quantization and token pruning process of the existing visual Transformer model, floating-point operations have resulted in large hardware overhead and high storage and computing resources requirements, making it difficult to achieve efficient end-side deployment.
The hardware-friendly visual Transformer compression method based on quantization and token pruning technology is adopted. By inserting the Token compression module into the model, pruning the unimportant tokens, and using the adaptive search method to optimize the pruning parameters to achieve the balance of quantization accuracy and pruning rate.
The model inference computing volume and storage requirements are greatly reduced, and the fixed-point computing power of the end-side equipment is fully utilized to realize the efficient end-side deployment of visual Transformer.
Smart Images

Figure CN120106168A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model quantization and pruning technology, and in particular to a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology. Background Art
[0002] With the rapid development of artificial intelligence, visual transformers have been widely used in image classification, autonomous driving and other fields. However, due to the high demand for computing and storage resources, the deployment of visual model transformers on the client side has become a huge challenge. To solve this problem, model quantization and token pruning technologies have emerged and have been widely used. Model quantization technology has been able to quantize the weights and activation values of visual transformers to 8 bits with little loss in model accuracy. Currently, token pruning no longer requires manual setting of the token compression rate of each block of the transformer, but uses deep learning to achieve adaptive compression.
[0003] However, when quantizing the visual Transformer, the quantization parameter is a floating-point type that is not a power of 2. When performing inference, the quantization and re-quantization process involves floating-point operations, which results in high hardware overhead and inefficient deployment.
[0004] The current quantization technology uses a channel-by-channel fine-grained quantization method to quantize the Transformer weights, and each channel has a set of quantization parameters. This method involves the storage of a large number of quantization parameters, which is a big challenge for the end-side devices that store resources.
[0005] When the current method performs token pruning on the visual Transformer model, the activation values and weights of the model are both floating-point types. When deployed on the client side, only the floating-point computing power of the client side device can be used, but the fixed-point computing power of the client side device with higher computing power cannot be used, resulting in low computing performance. Summary of the invention
[0006] In order to overcome the defects of the above-mentioned prior art, the present invention provides a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology. The compression method has the characteristics of efficient hardware deployment, full utilization of the integer inference computing power of the terminal device, and significant reduction of model inference calculation amount and storage requirements.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A hardware-friendly visual Transformer compression method based on quantization and Token pruning technology, comprising the following steps;
[0009] Step 1: Select an image, perform post-training quantization and calibration on the weights and activation values of the visual Transformer model for the image classification task, and obtain a quantized model;
[0010] Step 2: Insert a Token compression module between the multi-head attention module and the feed-forward layer of each block of the quantized visual Transformer model. The Token compression module prunes unimportant tokens.
[0011] Step 3: Train the quantized model with the Token compression module inserted to learn the pruning parameters. The pruning parameters searched by the adaptive search method achieve a balance between the quantization accuracy and the pruning rate.
[0012] In step 1:
[0013] Before quantization calibration, prepare a pre-trained model and randomly select pictures from the ImageNet2012 dataset as calibration data for model activation values for post-training quantization calibration of the pre-trained model.
[0014] The step 1 is specifically as follows:
[0015] The quantization granularity of the quantization calibration process is layer-by-layer quantization, and the layer-by-layer quantization parameter s is shared by the entire weight or activation value, while each channel of the weight has a set of quantization parameters during channel-by-channel quantization;
[0016] The quantization parameter is processed while the calibration is being performed. There are two ways to process the quantization parameter s:
[0017] They are
[0018] Floor is a rounding function, ceil is a rounding function, and the quantization parameter can adaptively select an optimal processing method during the calibration process.
[0019] The calibration algorithm is specifically:
[0020] First, the calibration image tensor X, the maximum value Qmax and minimum value Qmin of the 8-bit quantization bit width, and the maximum value M of the absolute value of the tensor X are given;
[0021] Initialize the iteration number k = 0 and the iteration step length t = 0.01; according to the formula:
[0022] M'=M*(1-k*t):
[0023] Calculate the maximum value of the tensor X in the kth round; calculate the floating-point quantization parameter s that is not a power of 2 in the kth round, and perform a power of 2 processing on s according to the quantization parameter processing algorithm to obtain quantization parameters s1 and s2 in the form of powers of 2; then calculate the tensors Y1 and Y2 after the dequantization of the tensor X;
[0024] Compare the mean square error of Y1 and Y2 with the tensor X, and select the smaller mean square error as the target score for the k+1th round;
[0025] After the calibration is completed, the quantization parameter scale in the form of a power of 2 is output for subsequent quantization and requantization processes.
[0026] In step 2:
[0027] The multi-head attention of the visual Transformer i The calculation process is:
[0028]
[0029] Where Q i :N*D,K i :N*D, N is the number of tokens, D is the dimension of the token, and the complexity of its calculation is proportional to the square of the number of tokens.
[0030] Multi-head attention i After Token pruning, the amount of computation can be greatly reduced.
[0031] The token pruning process includes three parts: token importance sorting, pruning, and fusion;
[0032] First, according to the formula X c =A c V, sort the importance of Tokens, where
[0033] A c is the attention score of the class token and other tokens, q c is the query vector of the class token, K is the key matrix, and V is the value matrix;
[0034] Then, according to X c The calculation results are used to sort the tokens. The larger the calculation result, the more important the token is. p Tokens are pruned, where N p =N*α p , α p is the pruning rate;
[0035] Calculate (N m -N p ) cosine similarity between unimportant tokens and the remaining tokens, where N m =N*α m , α m It is the fusion rate. Tokens with high similarity are added together and then the average is taken to complete the fusion.
[0036] In step 3, the pruning parameter, i.e., the pruning rate α p and the fusion rate α m In the adaptive search method, a set of candidate discrete compression rates (C 1 ,C 2 ,...C N ), compression ratio Indicates that (i-1) unimportant tokens will be compressed;
[0037] Compression ratio C i There is a learnable probability parameter ρ i ,and For each block of the visual Transformer, the compression rate can be expressed as In this way, the problem of discrete compressibility is transformed into learning the probability parameter ρ i question;
[0038] To make the probability parameter ρ i It can also be derived using π i represents the probability that the i-th Token is compressed, π i =ρ N+2-i +…+ρ N-1 +ρ N , where π 1 =0, i≥2.
[0039] Furthermore, the Token is encoded with a 0 / 1 mask. If π i ≥α, then m i =0, otherwise m i =1, 1 means that the token is retained. Now, each token can be represented by a mask, and the probability parameter ρ i Can be directed.
[0040] The hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is used in the fields of image classification and autonomous driving.
[0041] Beneficial effects of the present invention:
[0042] The present invention fully considers the hardware characteristics of the end-side device at the algorithm level. The technical method is hardware-friendly and can greatly reduce the demand of the visual Transformer for the computing resources and storage resources of the end-side device, thereby realizing efficient deployment of the visual Transformer on the end-side.
[0043] The quantization parameter of the model of the present invention is in the form of a power of 2. During model reasoning, the quantization and re-quantization processes can use shifting to replace floating-point calculations, thereby reducing hardware overhead.
[0044] The coarse-grained quantization method in the present invention reduces the demand for storage resources during the deployment process. Since Token pruning is combined on the basis of the quantization model, the computational complexity of the multi-head attention and feedforward layers can be greatly reduced, while enabling the model to perform reasoning under 8-bit fixed-point data, making full use of the fixed-point computing power of the end-side device.
[0045] Since the present invention introduces quantization error loss, the searched pruning parameters achieve a balance between quantization accuracy and pruning rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 The figure is a schematic diagram of the model quantization process according to an embodiment of the present invention.
[0047] Figure 2 A schematic diagram of the model quantization granularity according to an embodiment of the present invention.
[0048] Figure 3 This is a flow chart of a model quantization calibration algorithm according to an embodiment of the present invention.
[0049] Figure 4 Schematic diagram of a joint compression framework based on model quantization and Token pruning technology according to an embodiment of the present invention.
[0050] Figure 5 Schematic diagram of the joint compression process of model quantization and Token pruning in an embodiment of the present invention.
[0051] Figure 6 The figure is a schematic diagram of Token pruning according to an embodiment of the present invention.
[0052] Figure 7 Schematic diagram of a pruning parameter learning method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0054] In the embodiment of the present invention, on the one hand, a hardware-friendly visual Transformer quantization method is provided. Specifically, the quantization process is as follows: Figure 1 ;
[0055] Before quantization calibration, prepare a pre-trained model and randomly select about 1,000 images from the ImageNet2012 dataset as calibration data. Then perform post-training quantization calibration on the pre-trained model.
[0056] The quantization granularity of the calibration process is layer-by-layer quantization. Figure 2 , the layer-by-layer quantization parameter s is shared by the entire weight or activation value, while each channel of the weight has a set of quantization parameters during channel-by-channel quantization, which increases the demand for storage resources and poses great deployment challenges.
[0057] The quantization parameters are processed during calibration. There are two ways to process the quantization parameters s. Floor is a function for rounding down, and ceil is a function for rounding up. The quantization parameter can adaptively select an optimal processing method during the calibration process, which greatly reduces the quantization error caused by the quantization parameter processing.
[0058] The calibration algorithm is shown in the attached Figure 3 , firstly, given the calibration image tensor X, the maximum value Qmax and minimum value Qmin of the 8-bit quantization bit width, and the maximum value M of the absolute value of the tensor X;
[0059] Initialize the iteration number k=0 and the iteration step length t=0.01; calculate the maximum absolute value of the tensor X in the kth round according to the formula M'=M*(1-k*t); calculate the floating-point quantization parameter s that is not a power of 2 in the kth round, and calculate the quantization parameters s1 and s2 according to the quantization parameter processing algorithm; then calculate the tensors Y1 and Y2 after the dequantization of tensor X; compare the mean square errors of Y1 and Y2 with tensor X, and select the smaller mean square error as the target score score for the k+1th round.
[0060] After completing 90 rounds of calibration, the quantization parameter scale in the form of a power of 2 is output. Using this algorithm for calibration and quantization parameter processing, the reasoning accuracy loss of the visual Transformer models used for image classification tasks such as DeiT_tiny, DeiT_small, DeiT_base, ViT_base, Swin_tiny, Swin_small, Swin_base, and Vit_large after quantization is less than 2% compared to the floating-point model.
[0061] The embodiment of the present invention, on the other hand, provides a visual Transformer joint compression method based on model quantization and Token pruning on the basis of the above quantization.
[0062] Refer to the attached Figure 4,The present invention inserts a Token compression module between the layer normalization and the feed-forward layer to perform Token pruning.
[0063] Attached Figure 4 Except for the Token compression module, the other structures are the basic components of the visual Transformer. After the compression module is pruned, the number of tokens involved in the calculation is reduced, which can greatly reduce the amount of calculation of the multi-head attention and feedforward layers in the subsequent blocks.
[0064] Refer to the attached Figure 6 ,The token pruning process includes token importance ranking, pruning, and fusion. First, according to the formula X c =A c V, sort the importance of Tokens, where A c is the attention score of the class token and other tokens, q c is the query vector of the class token, K is the key matrix, and V is the value matrix.
[0065] Then, according to X c The calculation results of Token are used to sort the tokens. The larger the calculation result, the more important the token is. p Tokens are pruned, where N p =N*α p , α p is the pruning rate.
[0066] Further, calculate (N m -N p ) cosine similarity between unimportant tokens and the remaining tokens, where N m =N*α m , α m It is the fusion rate. Tokens with high similarity are added together and then the average is taken to complete the fusion.
[0067] The process of model quantization and token pruning combined compression is shown in the attached Figure 5 ;
[0068] First, the pre-trained visual Transformer is trained according to Figure 1 Quantitative process and Figure 3 The calibration algorithm is trained and quantized. The output of the feedforward layer is quantized token by token and symmetrically. The output of the Softmax component is quantized in Log 2Quantization, quantization of weights and other activation values adopts layer-by-layer, symmetric quantization, and the corresponding quantization parameters are processed into the form of exponential powers of 2. That is, the weights and activation values of the visual Transformer model of the image classification task are quantized after Int8 training by adopting layer-by-layer or token-by-token quantization and quantization parameters are processed into powers of 2; the Token compression module prunes unimportant tokens to reduce the number of tokens; then, the Token compression module is inserted into the quantized model, the model is trained, and the Token pruning parameters are searched. The training data set uses ImageNet2012.
[0069] Pruning parameter is the pruning rate α p and the fusion rate α m , the adaptive search method is shown in the attached Figure 7 First, we manually set a set of candidate discrete compression rates (C 1 ,C 2 ,...C N ), compression ratio Indicates that (i-1) unimportant Tokens will be compressed.
[0070] Compression ratio C i There is a learnable probability parameter ρ i ,and For each block of the visual Transformer, the compression rate can be expressed as In this way, the problem of discrete compressibility can be transformed into learning the probability parameter ρ i question.
[0071] To make the probability parameter ρ i It can also be derived using π i represents the probability that the i-th Token is compressed, π i =ρ N+2-i +…+ρ N-1 +ρ N , where π 1 =0, i≥2.
[0072] Furthermore, the Token is encoded with a 0 / 1 mask. If π i ≥α, then m i =0, otherwise m i =1, 1 means that the token is retained. Now, each token can be represented by a mask, and the probability parameter ρ i Due to the introduction of quantization error loss, the searched pruning parameters achieve a balance between quantization accuracy and pruning rate.
[0073] Compared with the method of first performing pruning parameter search and then quantizing, the model inference accuracy is greatly improved, which can meet the actual deployment requirements on the terminal side.
[0074] In general, since the quantization parameters of the model are in the form of powers of 2, the quantization and requantization process during model reasoning can use shifting to replace floating-point calculations, reducing hardware overhead; the coarse-grained quantization method reduces the demand for storage resources during the deployment process. Since Token pruning is combined on the basis of the quantization model, the amount of calculation of the multi-head attention and feedforward layers can be greatly reduced, while enabling the model to perform reasoning under 8-bit fixed-point data, making full use of the fixed-point computing power of the terminal device.
Claims
1. A hardware-friendly visual Transformer compression method based on quantization and Token pruning technology, characterized in that: The steps include: Step 1: Select an image, perform post-training quantization and calibration on the weights and activation values of the visual Transformer model for the image classification task, and obtain a quantized model; Step 2: Insert a Token compression module between the multi-head attention module and the feed-forward layer of each block of the quantized visual Transformer model. The Token compression module prunes unimportant tokens. Step 3: Train the quantized model with the Token compression module inserted to learn the pruning parameters. The pruning parameters searched by the adaptive search method achieve a balance between the quantization accuracy and the pruning rate.
2. According to claim 1, a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is characterized in that: In step 1: Before quantization calibration, prepare a pre-trained model and randomly select pictures from the ImageNet2012 dataset as calibration data for model activation values for post-training quantization calibration of the pre-trained model.
3. According to claim 2, a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is characterized in that: The step 1 is specifically as follows: The quantization granularity of the quantization calibration process is layer-by-layer quantization, and the layer-by-layer quantization parameter s is shared by the entire weight or activation value, while each channel of the weight has a set of quantization parameters during channel-by-channel quantization; The quantization parameter is processed while the calibration is being performed. There are two ways to process the quantization parameter s: They are Floor is a rounding function, ceil is a rounding function, and the quantization parameter can adaptively select an optimal processing method during the calibration process.
4. According to claim 3, a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is characterized in that: The calibration algorithm is specifically: First, given the calibration image tensor X, the maximum value Qmax and minimum value Qmin of the 8-bit quantization bit width, and the maximum value M of the absolute value of the tensor X; Initialize the iteration number k = 0 and the iteration step length t = 0.01; according to the formula: M'=M*(1-k*t): Calculate the maximum value of the tensor X in the kth round; calculate the floating-point quantization parameter s that is not a power of 2 in the kth round, and perform a power of 2 processing on s according to the quantization parameter processing algorithm to obtain quantization parameters s1 and s2 in the form of powers of 2; then calculate the tensors Y1 and Y2 after the dequantization of the tensor X; Compare the mean square error of Y1 and Y2 with the tensor X, and select the smaller mean square error as the target score for the k+1th round; After the calibration is completed, the quantization parameter scale in the form of a power of 2 is output for subsequent quantization and requantization processes.
5. According to claim 4, a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is characterized in that: In step 2: The multi-head attention of the visual Transformer i The calculation process is: Where Q i :N*D,K i :N*D, N is the number of tokens, D is the dimension of the token, and the complexity of its calculation is proportional to the square of the number of tokens.
6. The hardware-friendly visual Transformer compression method based on quantization and Token pruning technology according to claim 5, characterized in that: The token pruning process includes three parts: token importance sorting, pruning, and fusion; First, according to the formula Sort the importance of Tokens, where A c is the attention score of the class token and other tokens, q c is the query vector of the class token, K is the key matrix, and V is the value matrix; Then, according to X c The calculation results are used to sort the tokens. The larger the calculation result, the more important the token is. p Tokens are pruned, where N p =N*α p , α p is the pruning rate; Calculate (N m -N p ) cosine similarity between unimportant tokens and the remaining tokens, where N m =N*α m , α m It is the fusion rate. Tokens with high similarity are added together and then the average is taken to complete the fusion.
7. The hardware-friendly visual Transformer compression method based on quantization and Token pruning technology according to claim 1, characterized in that: In step 3, the pruning parameter, i.e., the pruning rate α p and the fusion rate α m In the adaptive search method, we first manually set a set of candidate discrete compression rates (C1, C2, ... C N ), compression ratio Indicates that (i-1) unimportant tokens will be compressed; Compression ratio C i There is a learnable probability parameter ρ i ,and For each block of the visual Transformer, the compression rate can be expressed as In this way, the problem of discrete compressibility is transformed into learning the probability parameter ρ i question; Use π i represents the probability that the i-th Token is compressed, π i =ρ N+2-i +…+ρ N-1 +ρ N , where π1=0, i≥2.
8. The hardware-friendly visual Transformer compression method based on quantization and Token pruning technology according to claim 7, characterized in that: Encode the Token with a 0 / 1 mask, if π i ≥α, then m i =0, otherwise m i =1, 1 means the token is retained, probability parameter ρ i Can be directed.
9. An application of a hardware-friendly visual Transformer compression method based on quantization and Token pruning technology, characterized in that: The hardware-friendly visual Transformer compression method based on quantization and Token pruning technology is used in the fields of image classification and autonomous driving.
Citation Information
Cited By
ViT semantic segmentation progressive Token pruning method and system based on multi-scale Tsallis entropy and low-level visual feature guidance
CN120851110A
Compression method based on information density driving and adaptive quadtree division
CN121353432A
A compression method based on information density driving and adaptive quadtree partitioning
CN121353432B