Image processing method of visual Transform model based on grouping quantization
By adopting a group quantization-based method in the Transformer model, dynamically partitioning the activation mapping channels and optimizing the quantization parameters, the problems of insufficient generalization ability and accuracy loss of hierarchical quantization technology in the Transformer model are solved, and efficient and accurate image processing is achieved in resource-constrained environments.
Patent Information
- Application Number
- CN202411978193.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing hierarchical quantization technology lacks generalization capabilities in the Transformer model, making it difficult to achieve stable performance improvements in the diversified Transformer architecture, and low bit quantization leads to a significant decrease in model accuracy, making it difficult to operate reliably in edge devices with limited resources or real-time applications.
An image processing method based on group quantization is proposed. By dynamically dividing the activation mapping channels of each layer into multiple groups, calculating the scaling factor and zero points of each group, optimizing the quantization parameters using the expected maximization EM algorithm, and determining the optimal number of groups under the bit operation constraints, ensuring that the statistical characteristics of the activation value in the group are consistent and the quantization accuracy is maximized.
It improves the application performance of the Transformer model in resource-constrained environments, reduces the accuracy loss caused by quantization, and improves the computing efficiency and accuracy of the model in image processing scenarios.
Smart Images

Figure CN120047328A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image processing method based on a group quantization visual Transformer model. Background Art
[0002] In recent years, the Transformer model has made significant progress in the field of image processing. The Transformer's self-attention mechanism enables it to efficiently process long-distance dependencies and global features, especially in the field of natural language processing (NLP), and has excellent performance in tasks such as text generation, summarization, translation, and question answering. However, such large-scale Transformer models rely on a large amount of computing resources, huge training data, and long training time, making them difficult to directly apply in scenarios with limited resources.
[0003] In practical applications, resource-constrained devices such as embedded systems, drones, and edge devices have extremely high computational requirements for models. These devices require models to balance storage space, inference speed, and energy consumption to meet real-time response requirements. However, large Transformer models are difficult to deploy effectively on these devices due to their high computational complexity. To address this problem, researchers have developed a variety of model compression and optimization techniques, especially quantization methods, to reduce computational costs and improve operational efficiency.
[0004] Quantization technology significantly reduces storage requirements and computational complexity by compressing model parameters and activation values from high-precision floating-point numbers to low-precision integer representations. Low-bit quantization methods are particularly important, as they significantly increase inference speed and reduce power consumption by compressing the model's weights and activation values to a lower number of bits (such as 4 bits or less). At the same time, group quantization technology is also widely used, which divides weights or activation values into multiple groups and quantizes them separately, thereby further improving efficiency and reducing precision loss.
[0005] The hierarchical quantization technology in Post-Training Quantization (PTQ) has attracted widespread attention due to its advantages in reducing computational complexity and storage requirements. This technology effectively reduces the consumption of computing resources by converting model parameters and activation values from high-precision floating-point numbers to low-bit numbers. However, hierarchical quantization faces many challenges in practical applications, especially in the promotion and application of Transformer models with different structures. Its generalization ability is insufficient, and it is difficult to achieve stable performance improvement in diverse Transformer architectures.
[0006] Specifically, hierarchical quantization often exhibits inconsistent effects when processing weights and activation values at different levels. In addition, low-bit quantization can lead to a significant decrease in model accuracy, especially in visual recognition of complex tasks, where the performance of the model is more significantly impaired. This loss of accuracy is particularly prominent in resource-limited edge devices or applications with high real-time requirements, where the model must not only meet accuracy requirements but also have an efficient inference speed. However, the loss of accuracy caused by hierarchical quantization makes it difficult for the model to run reliably in these stringent application scenarios.
[0007] Therefore, how to reduce the precision loss caused by quantization while maintaining computational efficiency has become an important issue that researchers need to solve urgently. Improving quantization methods to enhance generalization and precision performance will be the key direction of future Transformer model quantization research. Summary of the invention
[0008] In response to the above technical problems, the present invention proposes an image processing solution based on a group quantized visual Transformer model.
[0009] The first aspect of the present invention discloses an image processing method based on a group quantization visual Transformer model, the method comprising:
[0010] Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0011] Step S2, performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model, to obtain a visual Transformer model based on group quantization;
[0012] Step S3: Use the group quantization-based visual Transformer model to test the image instances in the test set.
[0013] According to the method of the first aspect of the present invention, the visual Transformer model includes an input layer, a normalization layer, an attention layer, a softmax layer, a linear transformation layer, a feedforward neural network layer, an activation function layer, a Transformer processing layer, and an output layer; wherein:
[0014] In step S2, when the visual Transformer model is subjected to group quantization training, for all layers except the input layer and the output layer, the following steps are performed: for image instances in the training set, the channels in the activation map of each layer are dynamically divided into multiple groups, and the activation values in each group have the same statistical characteristics, thereby reducing errors introduced in the quantization process.
[0015] According to the method of the first aspect of the present invention, step S2 further comprises:
[0016] For each group, calculate the minimum value min(A) and the maximum value max(A) of the activation value within the group, where A represents the activation map of the current group;
[0017] According to the calculated minimum and maximum values, the scaling factor s and zero point z of each group in the quantization process are determined. The calculation formula is:
[0018]
[0019] Among them, n is the quantization bit width, which indicates the quantization level allowed to be used. The scaling factor s is used to describe the ratio between the real activation value and the quantized integer value. The zero point z is used to adjust the activation value so that the minimum activation value can be mapped to a non-negative integer value within the quantization range. The calculation formula is:
[0020]
[0021] Each activation value A(i) within a group is quantized to an integer value q(i):
[0022]
[0023] The activation values within a group are mapped within a limited quantization range.
[0024] According to the method of the first aspect of the present invention, step S2 further comprises: quantizing the scaling factor and zero point of each group using an expectation maximization EM algorithm; wherein:
[0025] The scaling factor s(0) and zero point z(0) in the initial state are determined by the min(A) and max(A) of the activation values within the group;
[0026] Quantize the activation values within the group to discrete integers q according to the current scaling factor and zero point (k) (i), the calculation formula is:
[0027]
[0028] Recalculate the optimized scaling factor and zero point, and minimize the quantization error A(i)-A′(i), where A′(i) represents the activation value after dequantization, and the calculation formula is:
[0029] A′(i)=s (k+1) *(q (k) (i)-z (k+1) )
[0030] By continuously iteratively optimizing the scaling factor s and the zero point z, the parameters that minimize the quantization error are obtained.
[0031] According to the method of the first aspect of the present invention, step S2 further comprises:
[0032] Given the constraints on the bit operations of the processor carrying the visual Transformer model, the optimal number of groups is determined for each layer. The bit operations are calculated as:
[0033] BOP=G*C*H*W*n
[0034] Where C, H, and W are the number of channels, height, and width of the image instance, respectively, and n is the bit width of each group;
[0035] The optimal number of groups G* is solved through an iterative algorithm, so that the statistical characteristics of the activation values in each group remain consistent under the constraints of bit operations and the quantization accuracy is maximized. The constraints of bit operations are characterized as: BOP≤BOP max , BOP max is the maximum bit operation allowed.
[0036] According to the method of the first aspect of the present invention, in step S3, the image instances in the test set are tested using the group quantization-based visual Transformer model to perform an image classification task or an image detection task.
[0037] According to the method of the first aspect of the present invention, the visual Transformer model includes ViT-S, DeiT-T, Swin-S, GPT-3, and BERT models; the image dataset includes the atlases corresponding to ImageNet, COCO, GLUE, and SQuAD.
[0038] A second aspect of the present invention discloses an image processing system based on a group quantization visual Transformer model, the system comprising a processing unit, wherein the processing unit is configured to execute:
[0039] Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0040] Performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model to obtain a group quantization-based visual Transformer model;
[0041] The image instances in the test set are tested using the group quantization-based visual Transformer model.
[0042] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the image processing method based on the visual Transformer model of group quantization described in the first aspect of the present disclosure is implemented.
[0043] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image processing method based on the group quantization visual Transformer model described in the first aspect of the present disclosure is implemented.
[0044] In summary, in the technical solution proposed in the present invention: first, perform instance-based grouping. In this stage, by analyzing the internal features of the model and its performance in specific tasks, the parameters are divided into different groups for more effective quantization; next, calculate the quantization parameters and use statistical methods to determine the quantization accuracy of each group to optimize model performance and resource utilization; finally, implement a group size allocation strategy to dynamically adjust the size of each group according to actual needs and system limitations to ensure that the accuracy and efficiency of the model are maintained while reducing the computing cost. The present invention can improve the application performance of the Transformer model in resource-constrained environments in image processing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 Schematic diagram of the process of image processing based on group quantization visual Transformer model. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0048] The first aspect of the present invention discloses an image processing method based on a group quantization visual Transformer model, the method comprising:
[0049] Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0050] Step S2, performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model, to obtain a visual Transformer model based on group quantization;
[0051] Step S3: Use the group quantization-based visual Transformer model to test the image instances in the test set.
[0052] According to the method of the first aspect of the present invention, the visual Transformer model includes an input layer, a normalization layer, an attention layer, a softmax layer, a linear transformation layer, a feedforward neural network layer, an activation function layer, a Transformer processing layer, and an output layer; wherein:
[0053] In step S2, when the visual Transformer model is subjected to group quantization training, for all layers except the input layer and the output layer, the following steps are performed: for image instances in the training set, the channels in the activation map of each layer are dynamically divided into multiple groups, and the activation values in each group have the same statistical characteristics, thereby reducing errors introduced in the quantization process.
[0054] According to the method of the first aspect of the present invention, step S2 further comprises:
[0055] For each group, calculate the minimum value min(A) and the maximum value max(A) of the activation value within the group, where A represents the activation map of the current group;
[0056] According to the calculated minimum and maximum values, the scaling factor s and zero point z of each group in the quantization process are determined. The calculation formula is:
[0057]
[0058] Among them, n is the quantization bit width, which indicates the quantization level allowed to be used. The scaling factor s is used to describe the ratio between the real activation value and the quantized integer value. The zero point z is used to adjust the activation value so that the minimum activation value can be mapped to a non-negative integer value within the quantization range. The calculation formula is:
[0059]
[0060] Each activation value A(i) within a group is quantized to an integer value q(i):
[0061]
[0062] The activation values within a group are mapped within a limited quantization range.
[0063] According to the method of the first aspect of the present invention, step S2 further comprises: quantizing the scaling factor and zero point of each group using an expectation maximization EM algorithm; wherein:
[0064] The scaling factor s(0) and zero point z(0) in the initial state are determined by the min(A) and max(A) of the activation values within the group;
[0065] Quantize the activation values within the group to discrete integers q according to the current scaling factor and zero point (k) (i), the calculation formula is:
[0066]
[0067] Recalculate the optimized scaling factor and zero point, and minimize the quantization error A(i)-A'(i), where A'(i) represents the activation value after dequantization, and the calculation formula is:
[0068] A'(i)=s (k+1) *(q (k) (i)-z (k+1) )
[0069] By continuously iteratively optimizing the scaling factor s and the zero point z, the parameters that minimize the quantization error are obtained.
[0070] According to the method of the first aspect of the present invention, step S2 further comprises:
[0071] Given the constraints on the bit operations of the processor carrying the visual Transformer model, the optimal number of groups is determined for each layer. The bit operations are calculated as:
[0072] BOP=G*C*H*W*n
[0073] Where C, H, and W are the number of channels, height, and width of the image instance, respectively, and n is the bit width of each group;
[0074] The optimal number of groups G* is solved through an iterative algorithm, so that the statistical characteristics of the activation values in each group remain consistent under the constraints of bit operations and the quantization accuracy is maximized. The constraints of bit operations are characterized as: BOP≤BOP max , BOP max is the maximum bit operation allowed.
[0075] According to the method of the first aspect of the present invention, in step S3, the image instances in the test set are tested using the group quantization-based visual Transformer model to perform an image classification task or an image detection task.
[0076] According to the method of the first aspect of the present invention, the visual Transformer model includes ViT-S, DeiT-T, Swin-S, GPT-3, and BERT models; the image dataset includes the atlases corresponding to ImageNet, COCO, GLUE, and SQuAD.
[0077] First embodiment
[0078] like Figure 1 As shown in the figure, the instance-based group quantization for the Transformer model specifically includes: first, instance-based grouping is performed. In this stage, the parameters are divided into different groups by analyzing the internal characteristics of the model and its performance in specific tasks for more efficient quantization. Next, the quantization parameters are calculated and the quantization accuracy of each group is determined using statistical methods to optimize model performance and resource utilization. Finally, the group size allocation strategy is implemented to dynamically adjust the size of each group according to actual needs and system limitations to ensure that the accuracy and efficiency of the model are maintained while reducing the computing cost.
[0079] Second embodiment
[0080] The existing Transformer model quantization method mainly adopts a hierarchical quantization strategy to complete quantization by accumulating layer by layer. However, this method has cumulative errors, especially in deep networks. In addition, the existing group quantization method usually divides continuous channels into multiple groups evenly, while ignoring the dynamic range differences of each channel, resulting in a decrease in quantization accuracy. In response to these problems, the present invention proposes a multi-scaling factor quantization method for a visual Transformer model; applied to image processing.
[0081] S1 Preparation stage: Start by selecting a full-precision Transformer model that has been pre-trained on a large dataset, such as ViT-S, DeiT-T, Swin-S, GPT-3, BERT, etc., and prepare a calibration dataset (a subset corresponding to Image Net, COCO, GLUE, SQuAD). This dataset usually contains hundreds to thousands of images for determining quantization parameters.
[0082] S2 is based on instance grouping: For each input instance in the calibration dataset, the channels in the activation map of each layer are dynamically divided into multiple groups to ensure that the activation values within the group have similar statistical characteristics. This step is to minimize the errors that may be introduced during the quantization process. By analyzing the statistical characteristics of each channel (such as mean, variance, etc.), it is ensured that the activation values within the group have a high similarity in distribution. In this way, the same quantization parameters can be applied to each group, thereby improving the overall computational efficiency and reducing the model complexity while ensuring the quantization accuracy.
[0083] S3 calculates the quantization parameters: First, calculate the minimum value min(A) and the maximum value max(A) of the activation value in the group. A represents the activation map of the current group. Through these boundary values, we can determine the dynamic range of the activation value of the group. Then, according to the calculated minimum and maximum values, determine the scaling factor s and zero point z of each group. The scaling factor s represents the actual activation value range of each quantization step, and the calculation formula is:
[0084]
[0085] Among them, n is the quantization bit width (such as 8 bits, 4 bits, etc.), indicating the quantization level that can be used. The scaling factor determines the ratio between the real activation value and the quantized integer value. And the zero point z is used to adjust the activation value so that the minimum activation value can be mapped to a non-negative integer value within the quantization range. The formula is:
[0086]
[0087] Zero point adjustment is used to ensure that negative activation values can also be properly represented in the quantization process without losing information. Finally, each activation value A(i) in the group is quantized to an integer value q(i):
[0088]
[0089] This ensures that the activation values within the group can be accurately mapped within a limited quantization range, reducing quantization errors and ensuring improved computational efficiency.
[0090] S4 Optimize quantization parameters: During the quantization process, the scaling factor and zero point of each group can be optimized using the expectation maximization (EM) algorithm to reduce the quantization error. The specific steps are as follows:
[0091] Step 1 Initialization: The initialization scaling factor s(0) and the zero point z(0) are determined by the min(A) and max(A) of the activation values within the group.
[0092] Step 2: Expected step: According to the current scaling factor and zero point, the activation value in the group is quantized into a discrete integer q(k)(i), the formula is:
[0093]
[0094] Step 3: Maximization: In the maximization step, the optimized scaling factors and zero points are recalculated to minimize the quantization error A(i)-A′(i), where A′(i) is the dequantized activation value, and the formula is:
[0095] A'(i)=s (k+1) *(q (k) (i)-z (k+1) ) (5)
[0096] By continuously iterating the process of step 2 and step 3, the scaling factor s and the zero point z are optimized, and finally the parameters that minimize the quantization error are obtained.
[0097] S5 Group size allocation: Under the given bit-operation (BOP) constraint, the optimal number of groups is determined for each layer. The accuracy and computational efficiency can be balanced by optimizing the number of groups G. Specifically, BOP can be expressed as:
[0098] BOP=G*C*H*W*n (6)
[0099] Among them, C, H, and W are the number of channels and spatial dimensions respectively, and n is the bit width of each group. In order to minimize the quantization error while satisfying the BOP constraint, the number of groups G needs to be optimized.
[0100] Through the iterative algorithm, the optimal G* is solved so that BOP≤BOP max In this case, the statistical characteristics of the activation values in each group remain consistent, thus ensuring maximum quantization accuracy.
[0101] S6 exports the quantized model: by applying steps 2 to 5 to different layers of different Transformer models, finally, exporting the quantized model as output.
[0102] The above outputs are deployed on resource-constrained devices such as drones, and the visual Transformer model is used to perform image classification tasks on the ImageNet dataset, object detection tasks on the COCO dataset, and classification and reading comprehension tasks on the large language model on the GLUE and SQuAD datasets. This allows the instance-based grouping quantization method to quantize Transformer models of different modalities, thereby improving the generalization of the quantization method, achieving good compression effects, and significantly improving the inference speed.
[0103] A second aspect of the present invention discloses an image processing system based on a group quantization visual Transformer model, the system comprising a processing unit, wherein the processing unit is configured to execute:
[0104] Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances;
[0105] Performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model to obtain a group quantization-based visual Transformer model;
[0106] The image instances in the test set are tested using the group quantization-based visual Transformer model.
[0107] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the image processing method based on the visual Transformer model of group quantization described in the first aspect of the present disclosure is implemented.
[0108] The fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image processing method based on the group quantization visual Transformer model described in the first aspect of the present disclosure is implemented.
[0109] In summary, the present invention designs an instance-based grouping quantization method for the Transformer model, which includes instance-based grouping, calculation and optimization of quantization parameters, and group size allocation. The scheme of the present invention designs an instance-based grouping strategy. For each input instance in the calibration data set, the channels in the activation map of each layer are dynamically divided into multiple groups to ensure that the activation values in the group have similar statistical characteristics, thereby reducing the accumulation of subsequent quantization errors and enhancing the generalization of this quantization method. Therefore, Transformer models of different modes can use this quantization method to enable them to be deployed in devices with limited resources. The present invention designs a group size allocation strategy. Under the constraints of positioning operations, the optimal group size and number of groups are determined for each layer. The accuracy and computational efficiency can be balanced by optimizing the number of groups, and the quantization accuracy is improved while the computational efficiency is reduced as little as possible, that is, the quantized Transformer model has good real-time performance for different tasks.
[0110] In the technical solution proposed in the present invention: first, perform instance-based grouping. In this stage, by analyzing the internal features of the model and its performance in specific tasks, the parameters are divided into different groups for more effective quantization; next, calculate the quantization parameters and use statistical methods to determine the quantization accuracy of each group to optimize model performance and resource utilization; finally, implement a group size allocation strategy to dynamically adjust the size of each group according to actual needs and system limitations to ensure that the accuracy and efficiency of the model are maintained while reducing the computing cost. The present invention can improve the application performance of the Transformer model in resource-constrained environments in image processing scenarios.
[0111] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above-mentioned embodiments only express several implementation methods of the present application, and their descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present application, several variations and improvements can be made, which all belong to the scope of protection of the present application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.
Claims
1. An image processing method based on a group quantized visual Transformer model, characterized in that: The method comprises: Step S1, generating an image data set, and dividing the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances; Step S2, performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model, to obtain a visual Transformer model based on group quantization; Step S3: Use the group quantization-based visual Transformer model to test the image instances in the test set.
2. The image processing method based on the visual Transformer model of group quantization according to claim 1 is characterized in that: The visual Transformer model includes an input layer, a normalization layer, an attention layer, a softmax layer, a linear transformation layer, a feedforward neural network layer, an activation function layer, a Transformer processing layer, and an output layer; wherein: In step S2, when the visual Transformer model is subjected to group quantization training, for all layers except the input layer and the output layer, the following steps are performed: for image instances in the training set, the channels in the activation map of each layer are dynamically divided into multiple groups, and the activation values in each group have the same statistical characteristics, thereby reducing errors introduced in the quantization process.
3. The image processing method based on the group quantization visual Transformer model according to claim 2 is characterized in that: Step S2 also includes: For each group, calculate the minimum value min(A) and the maximum value max(A) of the activation value within the group, where A represents the activation map of the current group; According to the calculated minimum and maximum values, the scaling factor s and zero point z of each group in the quantization process are determined. The calculation formula is: Among them, n is the quantization bit width, which indicates the quantization level allowed to be used. The scaling factor s is used to describe the ratio between the real activation value and the quantized integer value. The zero point z is used to adjust the activation value so that the minimum activation value can be mapped to a non-negative integer value within the quantization range. The calculation formula is: Each activation value A(i) within a group is quantized to an integer value q(i): The activation values within a group are mapped within a limited quantization range.
4. The image processing method of the visual Transformer model based on group quantization according to claim 3 is characterized in that: Step S2 further includes: quantizing the scaling factor and zero point of each group using an expectation maximization EM algorithm; wherein: The scaling factor s(0) and zero point z(0) in the initial state are determined by the min(A) and max(A) of the activation values within the group; Quantize the activation values within the group to discrete integers q according to the current scaling factor and zero point (k) (i), the calculation formula is: Recalculate the optimized scaling factor and zero point, and minimize the quantization error A(i)-A′(i), where A′(i) represents the activation value after dequantization, and the calculation formula is: A′(i)=s (k+1) *(q (k) (i)-z (k+1) ) By continuously iteratively optimizing the scaling factor s and the zero point z, the parameters that minimize the quantization error are obtained.
5. The image processing method based on the group quantization visual Transformer model according to claim 4 is characterized in that: Step S2 also includes: Given the constraints on the bit operations of the processor carrying the visual Transformer model, the optimal number of groups is determined for each layer. The bit operations are calculated as: BOP=G*C*H*W*n Where C, H, and W are the number of channels, height, and width of the image instance, respectively, and n is the bit width of each group; The optimal number of groups G* is solved through an iterative algorithm, so that the statistical characteristics of the activation values in each group remain consistent under the constraints of bit operations and the quantization accuracy is maximized. The constraints of bit operations are characterized as: BOP≤BOP max , BOP max is the maximum bit operation allowed.
6. The image processing method based on the group quantization visual Transformer model according to claim 5, characterized in that: In step S3, the image instances in the test set are tested using the group quantization-based visual Transformer model to perform an image classification task or an image detection task.
7. The image processing method based on the group quantization visual Transformer model according to claim 6, characterized in that: Visual Transformer models include ViT-S, DeiT-T, Swin-S, GPT-3, and BERT models; image datasets include atlases corresponding to ImageNet, COCO, GLUE, and SQuAD.
8. An image processing system based on a group quantized visual Transformer model, characterized in that: The system comprises a processing unit configured to perform: Generate an image data set, and divide the image data set into a training set and a test set; wherein the image data set includes a plurality of image instances; Performing group quantization training on the visual Transformer model using image instances in a training set and bit operations of a processor equipped with the visual Transformer model to obtain a group quantization-based visual Transformer model; The image instances in the test set are tested using the group quantization-based visual Transformer model.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the image processing method based on the group quantization visual Transformer model described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the image processing method based on the group quantization visual Transformer model described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for applying space-time vision Transform in dynamic ultrasonic instance segmentation
CN117541529A
Unsupervised learning image smoothing method and system based on Swin Transform V2
CN118297838A
Method for personalisation of ASR models
US20240257800A1