A Visual Transformer Compression Method and System Based on Differentiable Quantization Training
By introducing the micro-digitized step size and bias training method in the visual Transformer model, the problem of deploying the visual Transformer model on devices with limited computing power is solved, and the efficient compression and performance maintenance of the model are achieved.
Patent Information
- Application Number
- CN202210295189.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-24
AI Technical Summary
Due to the large number of parameters and high computational volume, the existing visual Transformer model is difficult to deploy on devices with limited computing power, and the performance of traditional quantization training methods at low bits is not ideal.
Using the visual Transformer compression method based on the micro-trainable training, the step size matching degree of the quantizer and the information of the negative activation area is improved by introducing the micro-numberable step size training method and the micro-numberable bias training method, thereby maintaining low performance losses while compressing the model size and inference delay.
It greatly reduces quantization errors and maintains performance close to the full-precision model. It is suitable for equipment deployment with limited computing power.
Smart Images

Figure CN114756517B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a visual Transformer compression method and system based on differentiable quantization training. Background Art
[0002] In recent years, models based on the Transformer architecture have achieved very successful results in various natural language processing (NLP) tasks. In the field of computer vision, some works based on the vision Transformer (ViT) have also achieved results approaching or even surpassing traditional convolutional neural networks (CNNs) in various vision tasks, including classification, detection, segmentation, super-resolution, denoising, etc. However, due to the very large number of parameters of the vision Transformer and the computational complexity that grows quadratically with the input image resolution, it brings high memory occupancy and high latency during inference, making it difficult to deploy on some devices with limited computing power, such as mobile devices and autonomous driving chips. Therefore, exploring suitable compression techniques to significantly reduce the size and inference latency of the vision Transformer model while maintaining low performance loss is crucial.
[0003] Quantization, as an effective compression technique, has been widely used in convolutional neural networks. Whether it is a convolutional neural network or a vision Transformer model, their core operations are matrix multiplications. By quantizing the weights and features originally represented as 32-bit floating-point numbers in the model into low-bit fixed-point numbers, low-bit fixed-point matrix multiplication operations can be used to replace the original floating-point matrix multiplication operations, thereby accelerating inference while compressing the model size. Quantization can be divided into post-training quantization and quantization-aware training according to whether fine-tuning is performed after quantization. For vision Transformers, existing post-training quantization-based works have all caused significant performance losses. And traditional quantization training methods have unsatisfactory performance at low bits because they do not fully consider the characteristics of vision Transformers. Summary of the Invention
[0004] The present invention provides a visual Transformer compression method and system based on differentiable quantization training to solve the technical problems existing in the above background art.
[0005] The present invention adopts the following technical solution: A visual Transformer compression method based on differentiable quantization training, comprising the following steps:
[0006] Step 1: Perform block processing on the input image and convert it into a corresponding image sequence through linear mapping;
[0007] Step 2: Sequentially process the image sequence through M times of quantization alternating processing of global information and local information to obtain a compressed image sequence;
[0008] Step 3: Classify the compressed image sequence and output the predicted probability value;
[0009] When performing Steps 1 to 3, a differentiable quantization step training method is introduced, and based on the differentiable quantization step training method, the matching degree between each differentiable quantization step and the image data is improved; at the same time, in Step 2, a differentiable quantization bias training method is introduced during local information quantization, and based on the differentiable quantization bias training method, the optimal quantization interval is automatically learned to retain the information of the negative activation region.
[0010] In a further embodiment, when performing the differentiable quantization step training method and / or the differentiable quantization bias training method, it further includes quantization parameter initialization based on minimizing the mean square error.
[0011] In a further embodiment, the differentiable quantization step training method is applicable to both image feature quantization and image weight quantization;
[0012] Among them, the differentiable quantization step training method includes the following process:
[0013] Define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as:
[0014]
[0015] In the formula, clip(z,a,b) means setting the elements in matrix z greater than a to a and the elements greater than b to b; the round operation means rounding-based integer taking; α represents the differentiable quantization step, -q min ,q max respectively represent the minimum and maximum values of the quantization range;
[0016] Through the dequantization operation, the floating-point corresponding to the fixed-point weight q is calculated
[0017] In a further embodiment, the differentiable quantization bias training method includes the following process:
[0018] Define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as:
[0019]
[0020] Wherein, clip(z,a,b) means setting the elements in matrix z greater than a to a and the elements greater than b to b; the round operation represents rounding-based integer truncation; α represents the differentiable quantization step size, -q miN ,q mAx represent the minimum and maximum values of the quantization range respectively;
[0021] Through the dequantization operation, the floating-point number corresponding to the fixed-point weight q is calculated where β is the introduced differentiable quantization bias.
[0022] In a further embodiment, the -q miN ,q max take the following values: given the quantization bit number b,
[0023] For signed number quantization, there is q min = 2 b-1 ,q max = 2 b-1 - 1;
[0024] For unsigned number quantization, there is q min = 0,q max = 2 b - 1.
[0025] In a further embodiment, during the dequantization operation, the straight-through estimator is used to process the gradient: if α is updated, its gradient is divided by an additional scaling factor g, where N w represents the number of elements of the full-precision weight w.
[0026] In a further embodiment, the initialization of the quantization parameters based on minimizing the mean square error specifically includes the following process:
[0027] For a layer with only the differentiable quantization step size α and no bias, its initialization method is expressed as:
[0028]
[0029] Assume α is known, and solve to obtain
[0030] Assume q is known, and solve to obtain
[0031] Iteratively solve q and α until α converges, take its value as the initial value of α, and then update α by the gradient descent method.
[0032] In a further embodiment, for a layer additionally having a bias β, its initialization method is expressed as:
[0033]
[0034] β * = E(w - α·q)
[0035] E(z) represents the average value of all elements of vector z; q, α, and β are solved by iterative iteration until α and β converge, and the solved values are used as the initial values of α and β, and then α and β are updated using the gradient descent method.
[0036] A vision Transformer compression system based on differentiable quantization training includes:
[0037] A quantization processing layer that performs block processing on the input image and converts it into a corresponding image sequence through linear mapping;
[0038] A self-attention layer configured to perform global information quantization processing on the image sequence;
[0039] A feed-forward layer configured to perform global information quantization processing on the image sequence; the feed-forward layer includes an activation layer; wherein, the self-attention layer and the feed-forward layer are alternately arranged M times;
[0040] A classification processing layer configured to classify the compressed image sequence and output a predicted probability value;
[0041] It further includes: a differentiable quantization step training module, which is sequentially embedded in the quantization processing layer, the self-attention layer, the feed-forward layer, and the classification processing layer; the differentiable quantization step training module is configured to improve the matching degree between each differentiable quantization step and the image data;
[0042] A differentiable quantization bias training module, which is embedded in the activation layer; the differentiable quantization bias training module is configured to automatically learn the optimal quantization interval and retain the information of the negative activation region.
[0043] In a further embodiment, it further includes: a quantization parameter initialization module, which is connected to both the differentiable quantization step training module and the differentiable quantization bias training module.
[0044] Advantages of the present invention: In the compression process of the present invention, a differentiable quantization step training method is introduced, making the step size of the quantizer more matched with the data distribution, thereby significantly reducing the quantization error. At the same time, when quantizing local information, a differentiable quantization bias training method is introduced, so that the information in the negative activation region is retained. And when running the differentiable quantization step training method and the differentiable quantization bias training method, a quantization parameter initialization based on minimizing the mean square error is used, aiming to ensure the convergence speed of the model and avoid the performance of the quantized model caused by slow convergence speed. Description of the Drawings
[0045] Figure 1 It is a schematic diagram of self-attention layer quantization.
[0046] Figure 2 It is an activation comparison diagram. Detailed Embodiments
[0047] The present invention will be further described below in conjunction with the drawings in the specification and specific embodiments.
[0048] Embodiment 1
[0049] This embodiment discloses a visual Transformer compression method based on differentiable quantization training, including:
[0050] Step 1: The input image is divided into blocks and converted into a corresponding image sequence through linear mapping; in this embodiment, a performance close to that of the full-precision visual Transformer model can be achieved under 8-bit quantization (4-fold compression rate) or 4-bit quantization (8-fold compression rate).
[0051] Step 2: The image sequence is sequentially subjected to M times of alternating processing of global information and local information quantization to obtain a compressed image sequence; where M is an integer; through M times of alternating quantization processing, while increasing the performance of the quantized image, fast compression can also be ensured. In this embodiment, the value of M is 12.
[0052] Step 3: The compressed image sequence is classified to output the predicted probability value; in this embodiment, 8-bit quantization (4-fold compression rate) can be used.
[0053] When performing Steps 1 to 3, a differentiable quantization step training method is introduced, and the matching degree between each differentiable quantization step and the image data is improved based on the differentiable quantization step training method. In other words, the differentiable quantization step training method is used every time quantization processing is performed.
[0054] Meanwhile, in Step 2, a differentiable quantization bias training method is introduced during local information quantization. The optimal quantization interval is automatically learned based on the differentiable quantization bias training method, and the information in the negative activation region is retained. In other words, every time an activation layer is executed, the differentiable quantization bias training method will be applied.
[0055] In a further embodiment, the differentiable quantization step training method is applicable to both image feature quantization and image weight quantization, that is, the quantization strategies for the features and weights of the image are the same. Taking weight quantization as an example, define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as:
[0056]
[0057] In the formula, clip(z,a,b) means setting the elements in matrix z greater than a to a and the elements greater than b to b; the round operation represents rounding-based integer conversion; α represents the differentiable quantization step, -q min ,q max respectively represent the minimum and maximum values of the quantization range;
[0058] Through the dequantization operation, the floating-point number corresponding to the fixed-point weight q is calculated
[0059] In another embodiment, if the quantization process of local information uses the GELU activation function to improve the performance of the model. The GELU activation function introduces negative activation values compared to the ReLU activation function. That is, in this embodiment, as Figure 2 shown, for the quantization of the GELU activation layer, unsigned integer quantization cannot be directly used like ReLU, which will lose the information contained in the negative activation values.
[0060] Therefore, to solve the above technical problems, a differentiable quantization bias training method is introduced for the quantization process of local information. The differentiable quantization bias training method includes the following process:
[0061] Define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as:
[0062]
[0063] In the formula, clip(z,a,b) means setting the elements in matrix z greater than a to a and the elements greater than b to b; the round operation represents rounding-based integer conversion; α represents the differentiable quantization step, -q min ,q max respectively represent the minimum and maximum values of the quantization range;
[0064] Through the dequantization operation, the floating-point corresponding to the fixed-point weight q is calculated where β is the introduced differentiable quantization bias.
[0065] When performing the dequantization operation in the above step quantization training method and bias quantization training method, since the round operation encounters the problem of gradient disappearance during backpropagation, STE (Straight Through Estimator) is used to handle its gradient, that is, the round operation is ignored during backpropagation. If α is updated, its gradient is divided by an additional scaling factor g where N w represents the number of elements of the full-precision weight w.
[0066] Adopting the above technology avoids the model from not converging due to too drastic changes in α. Therefore, the step quantization training method introduced in this embodiment makes the step of the quantizer more matched to the distribution of the data compared with the traditional fixed-step quantization training method, thereby greatly reducing the quantization error.
[0067] In a further embodiment, the -q min ,q max takes the following values: Given the quantization bit b
[0068] For signed quantization of quantities, there is q min = 2 b-1 ,q max = 2 b-1 - 1;
[0069] For unsigned quantization of quantities, there is q min = 0,q max = 2 b - 1.
[0070] In another embodiment, although the differentiable quantization step and bias are learnable parameters, it is still very important to select a suitable quantization parameter initialization method. If the initialization method is not selected properly, it will lead to a slow model convergence speed, thus affecting the performance of the model obtained by quantization training.
[0071] Therefore, when performing the differentiable quantization step training method and / or the differentiable quantization bias training method, it also includes quantization parameter initialization based on minimizing the mean square error, specifically including the following process:
[0072] For a layer with only the differentiable quantization step α and no bias, its initialization method is expressed as:
[0073]
[0074] Assuming α is known, solve to obtain
[0075] Assume q is known, and solve to obtain
[0076] Iteratively solve for q and α until α converges. Use its value as the initial value of α, and then update α through gradient descent.
[0077] Similarly, during the dequantization operation, use the straight-through estimator to handle the gradient: if α is updated, divide its gradient by an additional scaling factor g, where N w represents the number of elements of the full-precision weight w.
[0078] Based on the above method, tests are conducted on DeiT-Tiny and DeiT-Small. The test dataset is ImageNet2012, and the accuracy is the test result on the test set (Validation dataset) in this dataset. As shown in Table 1-1, it shows the compression rate and classification Top-1 accuracy of different models at different bit widths. Among them, FP32 represents the model represented by 32-bit floating-point numbers, that is, the full-precision model. Int8 and Int4 respectively represent the models of 8-bit quantization and 4-bit quantization. For 8-bit quantization, fine-tuning is only performed for 1 epoch. For 4-bit quantization, 300 epochs of fine-tuning are required. It can be seen from the table that whether it is Int8 or Int4, the accuracy loss of the quantized model is within 0.5%.
[0079]
[0080] Table 1-1 Experimental Results of Visual Transformer Quantization Training
[0081] Example 2
[0082] To complete the visual Transformer compression method described in Example 1, this example discloses a visual Transformer compression system based on differentiable quantization training, including:
[0083] A quantization processing layer, which is configured to perform block processing on the input image and convert it into a corresponding image sequence through linear mapping; in this example, similar performance to the full-precision visual Transformer model can be achieved under 8-bit quantization (4-fold compression rate) or 4-bit quantization (8-fold compression rate).
[0084] The self-attention layer is configured to perform global information quantization processing on the picture sequence. By quantizing both the weights and features in the self-attention layer into fixed-point numbers, all operations in the self-attention layer are implemented using fixed-point matrix multiplication. For the attention score, since its value is always greater than 0, unsigned quantization is used in this embodiment. In addition, all weights and features are quantized using signed numbers. As Figure 1 shown, Figure 1 the English symbol explanations are shown in Table 2.
[0085] Table 2
[0086]
[0087] The feed-forward layer is configured to perform global information quantization processing on the picture sequence. The feed-forward layer includes an activation layer. Among them, the self-attention layer and the feed-forward layer are alternately arranged M times;
[0088] The classification processing layer is configured to classify the compressed picture sequence and output the predicted probability value. In this embodiment, it is not necessary to use 8-bit quantization (4 times compression ratio).
[0089] It further includes: a differentiable quantization step training module, which is sequentially embedded into the quantization processing layer, the self-attention layer, the feed-forward layer, and the classification processing layer. The differentiable quantization step training module is configured to improve the matching degree between each differentiable quantization step and the image data;
[0090] The differentiable quantization bias training module is embedded in the activation layer. The differentiable quantization bias training module is configured to automatically learn the optimal quantization interval and retain the information in the negative activation region.
Claims
1. A method for compressing vision transformers based on differentiable quantization training, characterized in that, It includes the following steps: Step 1: Block-process the input image and convert it into a corresponding image sequence through linear mapping; Step 2: Sequentially process the image sequence through M times of alternating quantization of global information and local information to obtain a compressed image sequence; where M is an integer; Step 3: Classify the compressed image sequence and output the predicted probability value; When performing Steps 1 to 3, a differentiable quantization step training method is introduced, and based on the differentiable quantization step training method, the matching degree between each differentiable quantization step and the image data is improved; At the same time, when performing local information quantization in Step 2, a differentiable quantization bias training method is introduced, and based on the differentiable quantization bias training method, the optimal quantization interval is automatically learned to retain the information of the negative activation region; The differentiable quantization step training method is applicable to both image feature quantization and image weight quantization; Among them, the differentiable quantization step training method includes the following process: Define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as: Where clip(z, a, b) means setting the elements in matrix z greater than a to a and those greater than b to b; the round operation represents rounding-based integerization; α represents the differentiable quantization step size, -q min , q max represent the minimum and maximum values of the quantization range, respectively; Through the dequantization operation, the floating-point value corresponding to the fixed-point weight q is calculated The differentiable quantization bias training method includes the following process: Define the full-precision weight as w, the quantized fixed-point weight as q, and the quantization operation is expressed as: where clip(z,a,b) means setting the elements in matrix z greater than a to a and those greater than b to b; the round operation represents rounding-based integerization; α represents the differentiable quantization step, -q min ,q max represent the minimum and maximum values of the quantization range, respectively; Through the dequantization operation, the floating point corresponding to the fixed-point weight q is calculated where β is the introduced differentiable quantization bias.
2. The method for compressing vision transformers based on differentiable quantization training according to claim 1, characterized in that, When performing the differentiable quantization step training method and / or the differentiable quantization bias training method, it also includes quantization parameter initialization based on minimizing the mean square error.
3. The method for compressing vision transformers based on differentiable quantization training according to claim 1, characterized in that, The -q min , q max takes the following values: Given the quantization bit b, For signed quantization, q min = 2 b-1 , q max = 2 b-1 - 1; For unsigned quantization, there is q min = 0, q max = 2 b - 1.
4. The method for compressing vision transformers based on differentiable quantization training according to claim 1, characterized in that, During the dequantization operation, a straight-through estimator is used to process the gradient: if α is updated, its gradient is divided by an additional scaling factor g, where N w represents the number of elements of the full-precision weight w.
5. The method for compressing vision transformers based on differentiable quantization training according to claim 2, characterized in that, The quantization parameter initialization based on minimizing the mean square error specifically includes the following process: For a layer with only differentiable quantization step α and no bias, its initialization method is expressed as: Assume that α is known and solve to obtain Assume that q is known and solve to obtain Iteratively solve q and α until α converges, take its value as the initial value of α, and then update α through the gradient descent method.
6. The method for compressing vision transformers based on differentiable quantization training according to claim 2, characterized in that, For a layer with an additional bias β, its initialization method is expressed as: β * = E(w - α·q) E(z) represents the average value of all elements of vector z; Iteratively solve q, α, and β until α and β converge, take the obtained values as the initial values of α and β, and then use the gradient descent method to update α and β.
7. A visual Transformer compression system based on differentiable quantization training, for implementing the visual Transformer compression method according to any one of claims 1 to 6, characterized in that, It includes: A quantization processing layer that block-processes the input image and converts it into a corresponding image sequence through linear mapping; A self-attention layer configured to perform global information quantization processing on the image sequence; A feed-forward layer configured to perform global information quantization processing on the image sequence; The feed-forward layer includes an activation layer; Among them, the self-attention layer and the feed-forward layer are alternately arranged M times; A classification processing layer configured to classify the compressed image sequence and output the predicted probability value; It also includes: A differentiable quantization step training module sequentially embedded in the quantization processing layer, the self-attention layer, the feed-forward layer, and the classification processing layer; The differentiable quantization step training module is configured to improve the matching degree between each differentiable quantization step and the image data; A differentiable quantization bias training module embedded in the activation layer; The differentiable quantization bias training module is configured to automatically learn the optimal quantization interval and retain the information of the negative activation region.
8. The visual Transformer compression system based on differentiable quantization training according to claim 7, characterized in that, It also includes: A quantization parameter initialization module connected to both the differentiable quantization step training module and the differentiable quantization bias training module.
Citation Information
Patent Citations
Convolutional neural network post-training quantification method and system based on activated fixed-point fitting
CN111783961A
Transformer-based mark selection and combined expression recognition method and system
CN113705541A