Image processing method of visual Transform model based on multi-scaling factor quantization
By adopting a multi-scaling factor quantization method in the visual Transformer model, combining pair quantizers and uniform quantizers, dynamically selecting the scaling factor and logarithmic cardinality, the problem that a single scaling factor quantization method in the existing technology is difficult to deal with complex activation distribution, and higher model accuracy and inference performance are achieved.
Patent Information
- Application Number
- CN202411978180.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing quantization methods for visual Transformer models usually use a single scaling factor, making it difficult to effectively deal with Transformer models with complex activation distributions, resulting in a large loss of accuracy during the quantization process.
Multi-scaling factor quantization method is adopted to construct multi-scaling factor quantizers for the post-softmax activation layer and post-GELU activation layer respectively. By combining the quantizer and the uniform quantizer, the distribution characteristics of different activation values are adapted, and the scaling factor and logarithmic cardinality are dynamically selected through the iterative grid search strategy.
The performance and accuracy of the quantized model is significantly improved, the accuracy loss during the quantization process is reduced, and the inference performance and accuracy of the model is improved in resource-constrained environments.
Smart Images

Figure CN119963970A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image processing method based on a visual Transformer model with multi-scaling factor quantization. Background Art
[0002] In recent years, the Transformer model has made significant progress in the field of image processing. The Transformer's self-attention mechanism enables it to efficiently process long-distance dependencies and global features, especially in the field of natural language processing (NLP), and has excellent performance in tasks such as text generation, summarization, translation, and question answering. However, such large-scale Transformer models rely on a large amount of computing resources, huge training data, and long training time, making them difficult to directly apply in scenarios with limited resources.
[0003] The Visual Transformer (ViT) model is a deep neural network based on the self-attention mechanism that can effectively process global features, especially in tasks such as image classification, semantic segmentation, and object detection. Compared with traditional convolutional neural networks (CNNs), ViT is good at modeling long-distance dependencies in images, which gives it an advantage in processing global contextual information. However, the computational complexity of the ViT model is high, especially in the multi-head self-attention mechanism, where the computational complexity grows with the square of the number of image blocks, making it perform well on high-performance hardware, but its application is limited on devices with limited resources.
[0004] In practical applications, especially in resource-constrained scenarios such as drones and edge devices, the high computing requirements and inference latency of the ViT model have become bottlenecks. These scenarios require the model to have fast response capabilities while ensuring high accuracy. In order to meet this challenge, quantization methods have become an important means to improve the efficiency of the ViT model.
[0005] By quantizing model parameters and activation values, especially using low-bit quantization and multi-scale quantization factor technology, the computational complexity and memory requirements of the model can be significantly reduced, making it better suited to resource-constrained devices while maintaining relatively high accuracy.
[0006] Existing quantization methods for visual Transformer models usually use a single scaling factor and make them suitable for a single scaling factor through bias reparameterization. This cumulative effect is particularly significant when dealing with Transformer models with complex activation distributions. For post-softmax activations, the output is usually highly centralized, resulting in a large loss of accuracy during quantization. In the post-GELU positive and negative activation distributions, positive activation values are mostly concentrated in the positive interval, while negative activation values are more dispersed. This asymmetric activation distribution makes it impossible to capture the features of different activations in a balanced manner when quantizing with a single scaling factor. Although these methods can provide simplified implementations in some cases, bias reparameterization introduces errors, and these errors gradually accumulate in each layer, affecting the accuracy of the quantized model, which in turn leads to reduced accuracy and increased complexity in image processing. Summary of the invention
[0007] In view of the above technical problems, the present invention proposes an image processing solution based on a visual Transformer model with multi-scaling factor quantization.
[0008] The first aspect of the present invention discloses an image processing method based on a visual Transformer model with multi-scaling factor quantization, the method comprising:
[0009] Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0010] Step S2, using the image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, perform quantization scaling training on the visual Transformer model to obtain a visual Transformer model based on multi-scaling factor quantization;
[0011] Step S3: Use the multi-scaling factor quantization-based visual Transformer model to test the image instances in the test set.
[0012] According to the method of the first aspect of the present invention, the visual Transformer model includes a normalization layer, an attention layer, a post-softmax activation layer, a linear transformation layer, a feedforward neural network layer, a post-GELU activation layer, a Transformer processing layer, and an output layer; wherein:
[0013] In step S2, a multi-scaling factor quantizer is constructed for the post-softmax activation layer and the post-GELU activation layer, and when the post-softmax activation layer and the post-GELU activation layer process image instances in a training set, the multi-scaling factor quantizer is used to perform the quantization scaling training.
[0014] According to the method of the first aspect of the present invention, in step S2, when constructing a multi-scaling factor quantizer:
[0015] The activation value distribution characteristics of the post-softmax activation layer, post-GELU positive activation layer, and post-GELU negative activation layer are analyzed respectively;
[0016] According to the analysis results of the distribution characteristics, a logarithmic quantizer is constructed for the post-softmax activation layer, a uniform quantizer is constructed for the post-GELU negative activation layer, a logarithmic quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic of 0 to 1, and a uniform quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic greater than 1;
[0017] The individual logarithmic quantizers and uniform quantizers are combined into a multi-scaling factor quantizer.
[0018] According to the method of the first aspect of the present invention, in step S2, for the logarithmic quantizer:
[0019] Given activation A, logarithmic base Model quantization bit width ratio and quantization scaling factor The quantization process of activation A is expressed as:
[0020]
[0021] The dequantization process is expressed as:
[0022]
[0023] in, is the quantized activation value, A is the original activation value, b is the adaptive logarithmic base, s is the scaling factor, and bit is the quantization bit width.
[0024] According to the method of the first aspect of the present invention, in step S2, for the uniform quantizer:
[0025] Given activation A', model quantization bit width ratio and quantization scaling factor The quantization process of activation A' is expressed as:
[0026] Quantification:
[0027] The dequantization process is expressed as:
[0028]
[0029] in, is the quantized activation value, A' is the original activation value, s' is the scaling factor, and bit is the quantization bit width.
[0030] According to the method of the first aspect of the present invention, in step S2, the scaling factor and the logarithmic base are dynamically selected by an iterative grid search strategy; specifically comprising:
[0031] The scaling factors and zero points are evenly divided in the entire search space to generate an initial grid. Each grid point represents a combination of a set of hyperparameters, and the candidate set is a set of grid points.
[0032] For each set of hyperparameter combinations in the grid, based on the image instance, the quantization loss function is calculated and the loss value of each combination is recorded. The hyperparameter combination with the smallest loss is selected as the result of the preliminary search based on the quantization loss.
[0033] Based on the optimal hyperparameter combination, generate a smaller search grid, reduce the grid interval and refine the search range to cover more hyperparameter combinations with finer granularity; use the refined grid to generate a new candidate set to cover the optimal hyperparameter neighborhood in the preliminary search;
[0034] The same quantization loss calculation and search are performed on the refined grid, and the search range is gradually narrowed through continuous iteration until the search range reaches the preset range, and the hyperparameter combination with the smallest quantization loss is selected as the final optimal solution.
[0035] According to the method of the first aspect of the present invention, the visual Transformer model includes ViT-S, DeiT-T, and Swin-S models; the image dataset includes the atlas corresponding to Image Net and COCO.
[0036] A second aspect of the present invention discloses an image processing system based on a visual Transformer model with multiple scaling factors quantized, the system comprising a processing unit, the processing unit being configured to execute:
[0037] Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0038] Using image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, the visual Transformer model is trained for quantization scaling to obtain a visual Transformer model based on multi-scaling factor quantization;
[0039] The image instances in the test set are tested using the multi-scaling factor quantization-based visual Transformer model.
[0040] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the image processing method based on the multi-scaling factor quantized visual Transformer model described in the first aspect of the present disclosure is implemented.
[0041] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image processing method based on a multi-scaling factor quantized visual Transformer model described in the first aspect of the present disclosure is implemented.
[0042] In summary, the technical solution provided by the present invention aims to improve the performance and accuracy of the model after quantization. In particular, for the processing of post-softmax activation values, a logarithmic quantizer is used to more effectively capture the characteristics of the activation value, thereby reducing the loss of accuracy in the quantization process. At the same time, the method is specially optimized for the characteristics of the post-GELU activation value. During the quantization process, the activation values of the positive and negative parts use a logarithmic quantizer and a uniform quantizer respectively, so that the advantages of the two quantization methods can be fully utilized to achieve a more accurate representation of the activation value. In order to ensure that the most appropriate scaling factor is used for each part, an iterative grid search strategy is introduced. This strategy can efficiently explore the potential scaling factor space and quickly locate the optimal scaling factor configuration. This method can significantly improve its reasoning performance and accuracy in low-bit quantization while retaining the efficient characteristics of the visual Transformer model, providing reliable technical support for the application of the model in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 Schematic diagram of multi-scaling factor quantization of the visual Transformer model. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0046] Existing quantization methods for visual Transformer models mostly use a single scaling factor to reparameterize the deviation to make it suitable for a single scaling factor, but the deviation reparameterization will have errors and accumulate layer by layer. In this regard, the present invention proposes an image processing solution for a visual Transformer model based on multi-scaling factor quantization.
[0047] The first aspect of the present invention discloses an image processing method based on a visual Transformer model with multi-scaling factor quantization, the method comprising:
[0048] Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0049] Step S2, using the image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, perform quantization scaling training on the visual Transformer model to obtain a visual Transformer model based on multi-scaling factor quantization;
[0050] Step S3: Use the multi-scaling factor quantization-based visual Transformer model to test the image instances in the test set.
[0051] According to the method of the first aspect of the present invention, the visual Transformer model includes a normalization layer, an attention layer, a post-softmax activation layer, a linear transformation layer, a feedforward neural network layer, a post-GELU activation layer, a Transformer processing layer, and an output layer; wherein:
[0052] In step S2, a multi-scaling factor quantizer is constructed for the post-softmax activation layer and the post-GELU activation layer, and when the post-softmax activation layer and the post-GELU activation layer process image instances in a training set, the multi-scaling factor quantizer is used to perform the quantization scaling training.
[0053] According to the method of the first aspect of the present invention, in step S2, when constructing a multi-scaling factor quantizer:
[0054] The activation value distribution characteristics of the post-softmax activation layer, post-GELU positive activation layer, and post-GELU negative activation layer are analyzed respectively;
[0055] According to the analysis results of the distribution characteristics, a logarithmic quantizer is constructed for the post-softmax activation layer, a uniform quantizer is constructed for the post-GELU negative activation layer, a logarithmic quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic of 0 to 1, and a uniform quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic greater than 1;
[0056] The individual logarithmic quantizers and uniform quantizers are combined into a multi-scaling factor quantizer.
[0057] According to the method of the first aspect of the present invention, in step S2, for the logarithmic quantizer:
[0058] Given activation A, logarithmic base Model quantization bit width ratio and quantization scaling factor The quantization process of activation A is expressed as:
[0059]
[0060] The dequantization process is expressed as:
[0061]
[0062] in, is the quantized activation value, A is the original activation value, b is the adaptive logarithmic base, s is the scaling factor, and bit is the quantization bit width.
[0063] According to the method of the first aspect of the present invention, in step S2, for the uniform quantizer:
[0064] Given activation A', model quantization bit width ratio and quantization scaling factor The quantization process of activation A' is expressed as:
[0065] Quantification:
[0066] The dequantization process is expressed as:
[0067]
[0068] in, is the quantized activation value, A' is the original activation value, s' is the scaling factor, and bit is the quantization bit width.
[0069] According to the method of the first aspect of the present invention, in step S2, the scaling factor and the logarithmic base are dynamically selected by an iterative grid search strategy; specifically comprising:
[0070] The scaling factors and zero points are evenly divided in the entire search space to generate an initial grid. Each grid point represents a combination of a set of hyperparameters, and the candidate set is a set of grid points.
[0071] For each set of hyperparameter combinations in the grid, based on the image instance, the quantization loss function is calculated and the loss value of each combination is recorded. The hyperparameter combination with the smallest loss is selected as the result of the preliminary search based on the quantization loss.
[0072] Based on the optimal hyperparameter combination, generate a smaller search grid, reduce the grid interval and refine the search range to cover more hyperparameter combinations with finer granularity; use the refined grid to generate a new candidate set to cover the optimal hyperparameter neighborhood in the preliminary search;
[0073] The same quantization loss calculation and search are performed on the refined grid, and the search range is gradually narrowed through continuous iteration until the search range reaches the preset range, and the hyperparameter combination with the smallest quantization loss is selected as the final optimal solution.
[0074] According to the method of the first aspect of the present invention, the visual Transformer model includes ViT-S, DeiT-T, and Swin-S models; the image dataset includes the atlas corresponding to Image Net and COCO.
[0075] First embodiment (such as Figure 1 (shown)
[0076] S1: Start by selecting a full-precision visual Transformer model that has been pre-trained on a large dataset, such as ViT-S, DeiT-T, and Swin-S, and prepare a calibration dataset (the corresponding subset of ImageNet and COCO), which usually contains hundreds to thousands of images for determining quantization parameters.
[0077] S2: The activation values in the Transformer model are analyzed, especially the activation distribution of different modules is not analyzed. That is, the softmax activation distribution after different modules is mostly concentrated near zero and between zero and one; the GELU positive activation distribution after different modules is mostly concentrated near zero and widely distributed; the GELU negative activation distribution after different modules is distributed in the interval of (-0.17,0], most of which are distributed at both ends but relatively evenly; finally, it is concluded that the gap between the post-softmax activation distribution, the post-GELU positive activation distribution, and the post-GELU negative activation distribution is large.
[0078] S3: Based on the different activation distribution characteristics, a multi-scaling factor quantization method is designed for the post-Softmax and post-GELU layers, and the positive and negative parts of the post-GELU activation are independently uniformly quantized and logarithmically quantized. Specifically, most of the post-Softmax activation distributions of different modules are concentrated near zero, and are between zero and one, and show a power law, so a logarithmic quantizer is designed for the post-Softmax activation distribution. The post-GELU negative activation distributions of different modules are all distributed in the interval (-0.17,0], so a uniform quantizer is designed for it. The post-GELU positive activation distributions of different modules are divided into two parts, one part is mostly distributed between zero and one, and the other part is greater than one. Because they are all designed logarithmic quantizers, there will be parts less than one and greater than one with the same quantized results, so logarithmic quantizers and uniform quantizers are used for the two parts respectively.
[0079] Logarithmic Quantizer: Given an activation A, a logarithmic base Quantization bit width ratio and scaling factor The quantization process of activation A is formulated as (1)(2):
[0080] Quantification:
[0081] Dequantization:
[0082] in is the quantized activation value, A is the original activation value, b is the adaptive logarithmic base, s is the scaling factor, and bit is the quantization bit width.
[0083] Uniform quantizer: Given activation A, quantization bit width ratio and scaling factor The quantization process of activation A is formulated as (3)(4):
[0084] Quantification:
[0085] Dequantization:
[0086] in is the quantized activation value, A is the original activation value, s is the scaling factor, and bit is the quantization bit width.
[0087] The logarithmic quantizer and uniform quantizer designed above are integrated into a multi-scaling factor quantizer, and then operated on the post-Softmax activation and post-GELU activation of the visual Transformer model respectively.
[0088] S4: In order to further improve the quantization accuracy and quantization speed, the scaling factor and the logarithmic base should be adjusted dynamically.
[0089] S5: Perform step S3 on the input visual Transformer model, i.e., use multi-scaling factor quantizers on the post-Softmax activation and post-GELU activation of the visual Transformer model respectively, and cooperate with the iterative grid search strategy of S4, i.e., dynamically select the scaling factor and logarithmic base, except that the post-Softmax activation and post-GELU activation will be quantized using uniform quantizers and iterative grid search strategies, and finally output the quantized visual Transformer model as the result.
[0090] The above outputs are deployed on resource-constrained devices such as drones to perform image classification tasks on the ImageNet dataset and object detection tasks on the COCO dataset, thereby achieving significant model compression and significantly improving inference speed without significantly reducing the accuracy of the above tasks.
[0091] Second embodiment (dynamic selection of scaling factor and logarithmic base for iterative grid search strategy)
[0092] Iterative grid search strategy:
[0093] Step 1: First, two hyperparameters (such as scaling factors and zero points) are evenly divided in the entire search space to generate an initial grid. Each grid point represents a combination of a set of hyperparameters, and the candidate set is the set of these grid points.
[0094] Step 2: For each set of hyperparameter combinations in the grid, calculate the quantization loss function, i.e., mean square error, based on the calibration data, and record the loss value of each combination. Based on the quantization loss, select the hyperparameter combination with the smallest loss. This combination will be used as the result of the preliminary search.
[0095] Step 3: Generate a smaller search grid around the optimal hyperparameter combination found in step 2. Specifically, reduce the spacing of the grid and refine the search range to cover more fine-grained hyperparameter combinations. Use the refined grid to generate new candidate sets to cover the neighborhood of the best hyperparameters in the preliminary search.
[0096] Step 4: Repeat steps 2 and 3, and perform the same quantization loss calculation and search steps on the refined grid. Each iteration performs a more refined search near the best hyperparameters of the previous round, gradually narrowing the search range until the search range is small enough or the preset number of iterations is reached, and the hyperparameter combination with the smallest quantization loss is selected as the final optimal solution.
[0097] A second aspect of the present invention discloses an image processing system based on a visual Transformer model with multiple scaling factors quantized, the system comprising a processing unit, the processing unit being configured to execute:
[0098] Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0099] Using image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, the visual Transformer model is trained for quantization scaling to obtain a visual Transformer model based on multi-scaling factor quantization;
[0100] The image instances in the test set are tested using the multi-scaling factor quantization-based visual Transformer model.
[0101] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the image processing method based on the multi-scaling factor quantized visual Transformer model described in the first aspect of the present disclosure is implemented.
[0102] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image processing method based on a multi-scaling factor quantized visual Transformer model described in the first aspect of the present disclosure is implemented.
[0103] In summary, the technical solution provided by the present invention aims to improve the performance and accuracy of the model after quantization. In particular, for the processing of post-softmax activation values, a logarithmic quantizer is used to more effectively capture the characteristics of the activation value, thereby reducing the loss of accuracy in the quantization process. At the same time, the method is specially optimized for the characteristics of the post-GELU activation value. During the quantization process, the activation values of the positive and negative parts use a logarithmic quantizer and a uniform quantizer respectively, so that the advantages of the two quantization methods can be fully utilized to achieve a more accurate representation of the activation value. In order to ensure that the most appropriate scaling factor is used for each part, an iterative grid search strategy is introduced. This strategy can efficiently explore the potential scaling factor space and quickly locate the optimal scaling factor configuration. This method can significantly improve its reasoning performance and accuracy in low-bit quantization while retaining the efficient characteristics of the visual Transformer model, providing reliable technical support for the application of the model in resource-constrained environments.
[0104] The present invention designs a multi-scaling factor quantization method for a visual Transformer model, which includes a multi-scaling factor quantizer, quantizing post-Softmax activation and post-GELU activation of the visual Transformer model, and an iterative grid search strategy.
[0105] The present invention designs a multi-scaling factor quantizer. Based on the fact that the post-Softmax activation and post-GELU activation distributions of the visual Transformer model are relatively special, their distributions are divided and quantized in a more detailed manner. A multi-scaling factor quantizer composed of a logarithmic quantizer and a uniform quantizer is designed, thereby improving the accuracy of the model after quantization, namely, improving the Top-1 accuracy of image classification tasks on the ImageNet dataset and improving the accuracy of object detection tasks on the COCO dataset.
[0106] The present invention designs an iterative grid search strategy to select the optimal scaling factor and zero point for the quantization part of each layer, and to improve the quantization and inference speed, that is, to improve the real-time performance of image classification tasks on the ImageNet dataset and object detection tasks on the COCO dataset.
[0107] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above-mentioned embodiments only express several implementation methods of the present application, and their descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present application, several variations and improvements can be made, which all belong to the scope of protection of the present application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.
Claims
1. An image processing method based on a visual Transformer model with multi-scaling factor quantization, characterized in that: The method comprises: Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances; Step S2, using the image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, perform quantization scaling training on the visual Transformer model to obtain a visual Transformer model based on multi-scaling factor quantization; Step S3: Use the multi-scaling factor quantization-based visual Transformer model to test the image instances in the test set.
2. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 1, characterized in that: The visual Transformer model includes a normalization layer, an attention layer, a post-softmax activation layer, a linear transformation layer, a feedforward neural network layer, a post-GELU activation layer, a Transformer processing layer, and an output layer; wherein: In step S2, a multi-scaling factor quantizer is constructed for the post-softmax activation layer and the post-GELU activation layer, and when the post-softmax activation layer and the post-GELU activation layer process image instances in a training set, the multi-scaling factor quantizer is used to perform the quantization scaling training.
3. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 2, characterized in that: In step S2, when constructing a multi-scaling factor quantizer: The activation value distribution characteristics of the post-softmax activation layer, post-GELU positive activation layer, and post-GELU negative activation layer are analyzed respectively; According to the analysis results of the distribution characteristics, a logarithmic quantizer is constructed for the post-softmax activation layer, a uniform quantizer is constructed for the post-GELU negative activation layer, a logarithmic quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic of 0 to 1, and a uniform quantizer is constructed for the post-GELU positive activation layer with a distribution characteristic greater than 1; The individual logarithmic quantizers and uniform quantizers are combined into a multi-scaling factor quantizer.
4. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 3 is characterized in that: In step S2, for the logarithmic quantizer: Given activation A, logarithmic base Model quantization bit width ratio and quantization scaling factor The quantization process of activation A is expressed as: The dequantization process is expressed as: in, is the quantized activation value, A is the original activation value, b is the adaptive logarithmic base, s is the scaling factor, and bit is the quantization bit width.
5. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 4, characterized in that: In step S2, for the uniform quantizer: Given activation A', model quantization bit width ratio and quantization scaling factor The quantization process of activation A' is expressed as: Quantification: The dequantization process is expressed as: in, is the quantized activation value, A' is the original activation value, s' is the scaling factor, and bit is the quantization bit width.
6. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 5, characterized in that: In step S2, the iterative grid search strategy dynamically selects the scaling factor and logarithmic base; specifically, it includes: The scaling factors and zero points are evenly divided in the entire search space to generate an initial grid. Each grid point represents a combination of a set of hyperparameters, and the candidate set is a set of grid points. For each set of hyperparameter combinations in the grid, based on the image instance, the quantization loss function is calculated and the loss value of each combination is recorded. The hyperparameter combination with the smallest loss is selected as the result of the preliminary search based on the quantization loss. Based on the optimal hyperparameter combination, generate a smaller search grid, reduce the grid interval and refine the search range to cover more hyperparameter combinations with finer granularity; use the refined grid to generate a new candidate set to cover the optimal hyperparameter neighborhood in the preliminary search; The same quantization loss calculation and search are performed on the refined grid, and the search range is gradually narrowed through continuous iteration until the search range reaches the preset range, and the hyperparameter combination with the smallest quantization loss is selected as the final optimal solution.
7. The image processing method based on the visual Transformer model with multi-scaling factor quantization according to claim 6, characterized in that: Visual Transformer models include ViT-S, DeiT-T, and Swin-S models; image datasets include the atlases corresponding to Image Net and COCO.
8. An image processing system based on a multi-scaling factor quantized visual Transformer model, characterized in that: The system comprises a processing unit configured to perform: Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances; Using image examples in the training set, based on the model quantization bit width and quantization scaling factor allowed by the resource-constrained device equipped with the visual Transformer model, the visual Transformer model is trained for quantization scaling to obtain a visual Transformer model based on multi-scaling factor quantization; The image instances in the test set are tested using the multi-scaling factor quantization-based visual Transformer model.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the image processing method based on a multi-scaling factor quantized visual Transformer model as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the image processing method based on the multi-scaling factor quantized visual Transformer model described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Fine-grained per-vector scaling for neural network quantization
CN114118347A
Neural network quantification method and apparatus, and electronic device
CN115329957A
Deep learning network model optimization method based on parameter quantization
CN116524173A
LSTM (Long Short Term Memory) model quantitative retraining method, system and equipment for time sequence data processing
CN116956997A
Distribution flexible subset quantification method suitable for super-division network
CN117172301A