Mixing precision quantification method based on image confrontation sample perception and gradient optimization

By introducing methods of image adversarial sample perception and gradient optimization in the hybrid precision quantization framework, dynamically adjusting the bit width combination to deal with adversarial samples, solving the problem of insufficient robustness in the prior art, achieving a better balance of robustness and classification accuracy.

CN120197654APending Publication Date: 2025-06-24NORTHEASTERN UNIV CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444803.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing hybrid precision quantization framework is not robust enough when facing image adversarial samples, making it difficult to optimize computational complexity while improving robustness and maintaining image classification accuracy.

Method used

A hybrid precision quantization method based on image adversarial sample perception and gradient optimization is adopted to generate adversarial samples through adversarial attacks, and the network weight and quantization parameters of the quantizer are optimized in combination with gradient descent algorithms, and the bit width combination is dynamically adjusted to improve robustness.

Benefits of technology

The robustness of the convolutional neural network hybrid precision quantization framework based on gradient optimization is significantly improved, while maintaining the balance of image classification accuracy and computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197654A_ABST
    Figure CN120197654A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed precision quantification method based on image confrontation sample perception and gradient optimization, and relates to the field of mixed precision quantification. The method comprises the following steps: inputting an original RGB image sample and fixing bit width combination, and optimizing the network weight of MPQ-CNN; generating an optimal disturbance score corresponding to each RGB image original sample, and generating an RGB image confrontation sample; mPQ-CNN is used to predict RGB image confrontation sample categories, and cross entropy loss generated by probability distribution of prediction categories and one-hot coding vectors of real categories is used to guide quantization parameter optimization of an MPQ-CNN quantizer, so that bit width combination of the MPQ-CNN is updated; repeating the process for multiple times to complete a round of MPQ-CNN training process; after multiple rounds are repeated, the quantization parameter and the network weight of the optimal quantizer of the MPQ-CNN searched in each round are evaluated on the verification set, the network weight of the MPQ-CNN and the quantization parameter of the quantizer which enable the cross entropy loss of the verification set to be minimum are found, and then the bit width combination of the optimal MPQ-CNN is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mixed-precision quantization, and particularly to a mixed-precision quantization method based on image adversarial sample perception and gradient optimization. Background Art

[0002] In the field of deep learning, neural network quantization is a model compression technology that maps high-precision floating-point numbers to low-precision fixed-point numbers. In a neural network quantization framework, the bit width is usually used to represent the number of binary digits occupied by floating-point / fixed-point numbers in computer hardware. In computer hardware, fixed-point number operations are simpler than floating-point number operations, and the storage overhead of low-precision floating-point / fixed-point numbers is smaller than that of high-precision floating-point / fixed-point numbers. Therefore, the neural network quantization framework can improve the computing speed of the neural network and reduce its storage overhead. The quantization-aware training (QAT) process is a commonly used implementation method in the neural network quantization framework and has been proven by researchers to be one of the most effective methods for reducing neural network quantization errors.

[0003] Neural network mixed-precision quantization (MPQ) is an important branch of neural network quantization. In contrast, neural network fixed-precision quantization (FPQ) is a different neural network quantization technology. In the neural network FPQ framework, the high-precision floating-point numbers of the network weights and activation function outputs of the neural network are uniformly mapped to low-precision fixed-point numbers with exactly the same bit width; while in the neural network MPQ framework, the high-precision floating-point numbers of the network weights and activation function outputs of the neural network are differentially mapped to low-precision fixed-point numbers with not exactly the same bit width. Currently, with the continuous increase in the depth of neural networks, the number of parameters and computational complexity of neural networks have increased significantly, which poses higher requirements for the neural network MPQ framework. In this context, the neural network MPQ framework based on gradient optimization learns the quantization parameters of the weight quantizer of the network weights and the activation quantizer of the activation function outputs through the existing gradient descent algorithm, and has become the mainstream solution for neural network MPQ tasks. The specific implementation method of the neural network MPQ framework is the QAT process, as Figure 1As shown, in the structure of the K-th layer of the Mixed-Precision Quantization Framework for Convolutional Neural Network (MPQ-CNN), independent quantizers are respectively deployed for the network weights and the output of the activation function of the Convolutional Neural Network (CNN). The quantization parameters of the quantizer include the upper bound r max and the lower bound r min of the numerical range of high-precision floating-point numbers, as well as the upper bound q max and the lower bound q min of the numerical range of low-precision fixed-point numbers.

[0004] In image classification tasks, researchers have demonstrated that the classification accuracy of using MPQ-CNN is basically the same as that of using full-precision CNN. Therefore, in resource-constrained scenarios, such as mobile phone applications, using the MPQ framework to compress the CNN model is an effective model optimization method. However, the existing MPQ framework lacks research on the robustness issues of MPQ-CNN based on gradient optimization. In the classification task of visible light images (i.e., RGB images), the robustness issues of MPQ-CNN often manifest as follows: adding tiny perturbations that are difficult for the human visual system to detect to the original RGB image to generate adversarial samples to mislead the MPQ-CNN to output incorrect classification results. Such perturbations are usually constructed specifically through existing adversarial attack algorithms such as the Fast Gradient Sign Method (FGSM). Currently, improving the robustness of neural networks under the MPQ framework is still a very novel research topic, and only a few researchers have conducted research in this area: 1. The Ensembles of Mixed-Precision Deep Networks for Increased Robustness Against Adversarial Attacks (EMPIR) proposed by Sen et al. enhances the robustness against adversarial attacks by integrating full-precision and low-precision neural networks with the same neural network topology, leveraging quantization non-linearity and integration diversity. 2. The A Mixed-Precision Quantization Framework for Accurate and Certifiably Robust DNNs (ARQ) proposed by Yang et al. optimizes the MPQ framework based on deep reinforcement learning and incorporates robustness into the optimization objective for the first time. The above-mentioned robust MPQ-CNNs for RGB image classification tasks have the following problems respectively: 1. EMPIR is just an integration of multiple full-precision and low-precision CNNs and does not delve into the internal topology of the CNN model. For example, there are multiple convolutional layers in a CNN, and different convolutional layers have different sensitivities to adversarial samples of the same RGB image. Therefore, bit widths should be allocated differentially for CNN modules such as convolutional layers. 2. The image classification accuracy and robustness of ARQ highly depend on the full-precision CNN enhanced with Gaussian noise. If the pre-trained full-precision CNN has poor robustness, the optimization effect of ARQ will decline.

[0005] Currently, researchers have not conducted a systematic robustness study on the MPQ-CNN optimized based on the gradient descent algorithm. At the same time, the research by Etmann et al. (On the connection between adversarial robustness and saliency map interpretability) established the connection between the robustness and interpretability of existing neural network models. In image classification tasks, traditional adversarial training methods improve the robustness of CNNs by minimizing the cross-entropy loss of the joint distribution of the original image samples and the adversarial samples generated from the original images. However, this process is usually accompanied by a significant decrease in the classification accuracy of CNNs on the original image samples. In addition, for the existing MPQ-CNN based on gradient optimization, its QAT process requires synchronously optimizing the network weights of the MPQ-CNN and the quantization parameters of the quantizer. Therefore, the MPQ-CNN using the existing QAT process cannot effectively balance the multi-objective optimization among robustness, image classification accuracy, and computational complexity in traditional adversarial training methods. Summary of the Invention

[0006] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a mixed-precision quantization method based on image adversarial sample perception and gradient optimization, aiming to improve the robustness of the mixed-precision quantization framework of the convolutional neural network optimized based on gradients, and at the same time achieve the balanced optimization of image classification accuracy and computational complexity.

[0007] The technical solution of the present invention is as follows:

[0008] A mixed-precision quantization method based on image adversarial sample perception and gradient optimization, the method comprising the following steps:

[0009] S1. Input a batch of original RGB image samples and fix the bit-width combination of the MPQ-CNN, and use the gradient descent algorithm to optimize the network weights W of the MPQ-CNN optimized in the previous iteration t-1 , to obtain the network weights W optimized in the current iteration t ;

[0010] S2. Based on the original RGB image samples and the network weights W of the current batch t , use the adversarial attack method to generate the optimal perturbation scores corresponding to each original RGB image sample in the current batch, and then generate the adversarial samples of the original RGB image samples according to the optimal perturbation scores;

[0011] S3. Use MPQ-CNN to predict the class of RGB image adversarial samples, and use the cross-entropy loss generated by the probability distribution of the predicted class and the one-hot encoded vector of the true class to guide the optimization of the quantization parameters of the MPQ-CNN quantizer. Update the bit-width combination of MPQ-CNN using the optimized quantization parameters to obtain the current iteratively updated bit-width combination B t ;

[0012] S4. Repeat steps S1 to S3 until the preset maximum number of iterations is reached to complete one round of the training process of MPQ-CNN;

[0013] S5. Repeat steps S1 to S4 until the preset maximum number of rounds is reached;

[0014] S6. Evaluate on the validation dataset according to the quantization parameters and network weights of the optimal quantizer of MPQ-CNN searched in each round of Epoch. Finally, find the network weights of MPQ-CNN and the quantization parameters of the quantizer that minimize the cross-entropy loss of the validation dataset, and then determine the optimal bit-width combination of MPQ-CNN.

[0015] According to the hybrid precision quantization method described above, step S1 includes the following steps:

[0016] Step S1.1: Input a batch of original sample tensors of RGB images into MPQ-CNN for neural network forward propagation; the dimension of a batch of original sample tensors of RGB images is represented as [N, C, H, W], where N is the total number of image samples in the current batch, C represents the channels of the RGB image, and H and W are the height and width of the RGB image respectively;

[0017] Step S1.2: Fix the bit-width combination of MPQ-CNN and use the gradient descent algorithm to optimize the network weight W of MPQ-CNN optimized in the previous iteration t-1 , to obtain the network weight W optimized in this iteration t .

[0018] According to the hybrid precision quantization method, in step S1.2, MPQ-CNN is based on the following formula to perform gradient optimization on the current network weight of MPQ-CNN to obtain the trained network weight under the fixed bit-width combination ;

[0019]

[0020] where D tr represents the original RGB image sample tensor of the current batch; W represents the network weight to be optimized in MPQ-CNN; represents that the bit-width combination B of MPQ-CNN is fixed; Denote the loss function relied on when optimizing the weights \(W\) of the MPQ-CNN network using the gradient descent algorithm. The left side of the semicolon represents the input tensor relied on when solving this loss function; the right side represents the neural network parameters of the MPQ-CNN relied on when solving this loss function, that is, the network weights and the quantization parameters of the quantizer. Denote the network weights that minimize the output of the loss function in the MPQ-CNN.

[0021] Assign \(W\) to \(W\) t-1 , Assign Then the loss function is:

[0022]

[0023] where \(N\) represents the total number of original RGB image samples in the current batch, \(x\) i represents the tensor of dimension \([1, C, H, W]\) corresponding to the \(i\)-th original RGB image sample in the training dataset; \(W\) t-1 represents the optimized network weights in the \((t - 1)\)-th iteration; \(B\) t-1 represents the optimized bit-width combination in the \((t - 1)\)-th iteration; represents the probability distribution of the image class predicted by the MPQ-CNN for the \(i\)-th original RGB image sample in the training dataset under the optimized network weights \(W\) t-1 and the bit-width combination \(B\) t-1 in the \((t - 1)\)-th iteration, fixing \(B\) t-1 ; \(t\) i is the one-hot encoded vector of the true class of the original RGB image sample in the training dataset; \(L\) BC represents the bit-width combination complexity function of the MPQ-CNN; \(\beta\) is a preset hyperparameter ratio.

[0024] According to the mixed-precision quantization method described above, the adversarial attack method is the projected gradient descent algorithm.

[0025] According to the mixed-precision quantization method described above, the quantization parameters of the MPQ-CNN quantizer are the quantization parameters of all the weight quantizers and activation function quantizers that determine the bit-width combination of the MPQ-CNN.

[0026] According to the mixed-precision quantization method described above, the bit-width combination of the MPQ-CNN is jointly composed of the bit-width combination of the network weights in the MPQ-CNN and the bit-width combination of the activation function output.

[0027] According to the mixed-precision quantization method described above, step S3 includes the following steps:

[0028] Step S3.1: Input the adversarial sample tensor into MPQ-CNN, and perform forward propagation of the neural network on the adversarial sample tensor of the current batch;

[0029] Step S3.2: Fix the network weights of MPQ-CNN, and use the gradient descent algorithm to optimize the quantization parameters of the quantizer that determines the current MPQ-CNN bitwidth combination;

[0030] Step S3.3: Update the bitwidth combination B t-1 to B t .

[0031] According to the described mixed-precision quantization method, the process of Step S3.2 is expressed as the following formula:

[0032]

[0033] In the j-th round, denotes the network weights W of MPQ-CNN optimized at the t-th iteration in the fixed state; B t ; B t-1 denotes the unoptimized bitwidth combination B of MPQ-CNN at the t-th iteration; t-1 ; denotes the loss function on which the optimization of the bitwidth combination B of MPQ-CNN depends using the gradient descent algorithm; t-1 ; denotes the input tensor on which the solution of this loss function depends; and B t-1 denote the neural network parameters of MPQ-CNN on which the output of the solution of this loss function depends, that is, the network weights and the quantization parameters of the quantizer; denotes the solution of the bitwidth combination B that minimizes the output of the loss function in MPQ-CNN; denotes the process of determining the quantization parameters of the quantizer for the bitwidth combination of MPQ-CNN through gradient optimization by the formula 18 under the fixed network weights W t of MPQ-CNN; and there is

[0034]

[0035] where, denotes the probability distribution of the predicted classes of the RGB image adversarial samples corresponding to the input tensor t under the current network weights W t-1 optimized at the t-th iteration and the bitwidth combination B t optimized at the (t - 1)-th iteration of MPQ-CNN. After fixing W ; t iis the one-hot encoded vector of the true class of the original sample of the RGB image in the training dataset; both α and β are hyperparameter ratios; L BC (B t-1 ) is the bit-width combination complexity function; is the sum of the outputs of the cross-entropy loss function.

[0036] According to the mixed-precision quantization method described above, the gradient descent algorithm is the stochastic gradient descent algorithm.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) The present invention proposes an adversarial training method for MPQ-CNN based on gradient optimization. Through the three-stage cascaded optimization QAT process proposed by the present invention, the robustness of MPQ-CNN based on gradient optimization is improved. The Gradient-weighted Class Activation Mapping (Grad-Cam) method is used to perform visual analysis on the RGB image samples of the MPQ-CNN trained by the present invention. Experiments show that the MPQ-CNN optimized by the present invention has better interpretability.

[0039] (2) The present invention does not need to use pre-trained full-precision CNN to statically pre-generate RGB image adversarial samples, but dynamically uses the continuously optimized perturbation scores in the three-stage cascaded optimization QAT process to generate better adversarial samples to optimize the quantization parameters of the MPQ-CNN quantizer, and improves the robustness of MPQ-CNN from the perspective of bit-width combination through adversarial sample perception.

[0040] (3) During the training process of MPQ-CNN of the present invention, the network weights and the quantization parameters of the quantizer are optimized independently, so it does not affect the classification accuracy after MPQ-CNN training. After training, the present invention obtains an MPQ-CNN that achieves an optimal trade-off among the classification accuracy, robustness, and BitOps metric of bit-width combination complexity in the RGB image classification task. Description of the Drawings

[0041] Figure 1 is a schematic diagram of the structure of the K-th layer of the convolutional neural network mixed-precision quantization framework;

[0042] Figure 2 is a flowchart of the mixed-precision quantization method based on image adversarial sample perception and gradient optimization in this embodiment;

[0043] Figure 3 is an architecture diagram of the mixed-precision quantization method based on image adversarial sample perception and gradient optimization in this embodiment;

[0044] Figure 4 This is the flowchart for generating image adversarial examples by the PGD algorithm in the present invention. Specific implementation manner

[0045] To facilitate the understanding of this application, the following will provide a more comprehensive description of this application with reference to the relevant attached drawings.

[0046] As Figure 2 and Figure 3 shown, this implementation manner is based on a mixed-precision quantization method for image adversarial example perception and gradient optimization, including the following steps:

[0047] Step 1: Establish a training data set and a validation data set composed of RGB images.

[0048] In this implementation manner, the RGB images for establishing the training data set and the validation data set are provided by the ImageNet data set of the ILSVRC 2012 version. Specifically, the original RGB image samples are extracted from the ImageNet data set and divided into a training data set and a validation data set in a ratio of 8:2.

[0049] Step 2: Initialize the iteration number t = 0 and the training epoch number j = 0.

[0050] Step 3: Extract a batch of original RGB image samples from the training data set and the validation data set for standard preprocessing to obtain the original RGB image sample tensors.

[0051] In this implementation manner, all pixel values in the original RGB image samples of the training data set and the validation data set are linearly normalized from [0, 255] to the [0, 1] interval and converted into a tensor format with the [N, C, H, W] dimension in batches. Among them, H is the total number of original RGB image samples in one batch, C = 3 represents the channels of the RGB image, and H and W are the height and width of the RGB image respectively.

[0052] Step 4: Input the original RGB image sample tensors of a batch of preprocessed training data set into the MPQ-CNN.

[0053] For example, when the training data set contains 100 original RGB image samples and the maximum iteration number, i.e., the batch number, is set to 10, then the total number N of original RGB image samples in each batch is 10 (100 / 10 = 10).

[0054] Step 5: In the MPQ-CNN, perform a neural network forward propagation on the original RGB image sample tensors of the current batch.

[0055] Step 6: Fix the bit-width combination B of MPQ-CNN, that is, do not optimize the quantization parameters of the weight quantizer and activation quantizer in MPQ-CNN using gradients, and use the existing gradient descent algorithm to optimize the network weights of the current MPQ-CNN.

[0056] The process of Step 6 is shown in Equation 1:

[0057]

[0058] In Equation 1, D tr represents the original sample tensor of the RGB image of the current iterative training dataset. W represents the network weights to be optimized in MPQ-CNN; indicates that the bit-width combination B of MPQ-CNN is fixed, that is, the quantization parameters of the quantizer in MPQ-CNN do not participate in the gradient calculation. represents the loss function on which the optimization of the MPQ-CNN network weights W depends when using the existing gradient descent algorithm such as Stochastic Gradient Descent (SGD). Specifically, the left side of the semicolon represents the input tensor on which the solution of this loss function depends; the right side of the semicolon represents the neural network parameters of MPQ-CNN on which the solution of this loss function depends, that is, the network weights and the quantization parameters of the quantizer. represents the solution of the network weights W in MPQ-CNN that minimizes the output of the loss function represents the process of obtaining the trained network weights after gradient optimization in Equation 1 when MPQ-CNN is under the fixed bit-width combination

[0059] As Figure 1 shown, the network weights of the convolutional layer in MPQ-CNN are deployed with a weight quantizer, and the output of the activation function of the convolutional layer is deployed with an activation quantizer. Use B W to represent the bit-width combinations of different network weights in MPQ-CNN, and B A to represent the bit-width combinations of different activation function outputs in MPQ-CNN. Then, the bit-width combination B of MPQ-CNN can be represented by the set {B A , B W}. Only the bit-widths of the network weights deployed with the weight quantizer and the activation function outputs deployed with the activation quantizer are the targets for gradient optimization in MPQ-CNN. In the existing QAT process, during the forward propagation of the neural network, the output of the quantizer is used to simulate quantization noise, and the quantization parameters of the quantizer are optimized using the backpropagation of the neural network and gradients, thereby determining the bit-widths of the network weights and activation function outputs to be optimized in the computer hardware.

[0060] In this embodiment, use​​ denotes the output of the network weight W passing through the weight quantizer, denoted by denotes the output of the activation function output y passing through the activation quantizer. The quantization parameters of the existing quantizer map the floating-point numbers W and y to fixed-point numbers and Taking the existing linear quantizer as an example, the quantization factor s calculated by its quantization parameters linearly maps the floating-point numbers W and y to fixed-point numbers and The quantization factor s of the linear quantizer is equal to For the same linear quantizer, when the input is x, the calculation method of the output of the linear quantizer is shown in Equation 2:

[0061]

[0062] where q max and q min respectively represent the maximum and minimum values of the fixed-point number range; r max and r min respectively represent the maximum and minimum values of the floating-point number range; z is the quantization zero point. represents the rounding operation. The input x of the linear quantizer is the floating-point network weight W or the activation function output y.

[0063] It should be noted that the quantization parameters of different quantizers can be different, which reflects the characteristics of the mixed-precision quantization task. In computer hardware, the bit widths of the network weight and the activation function output are directly determined by the quantization parameters q max and q min of the weight quantizer and the activation quantizer respectively. During the forward propagation of the neural network and the backpropagation of the neural network in MPQ-CNN, the outputs of the weight quantizer and the activation quantizer are required. Therefore, during the process of optimizing the quantization parameters of the MPQ-CNN quantizer through gradients, it depends on the fixed-point number output calculated by the quantization factor and to perform the forward propagation of the neural network. Therefore, optimizing the bit width combination of MPQ-CNN using gradients depends on the quantization parameters of all quantizers in MPQ-CNN.

[0064] Step 6-1: MPQ-CNN performs the backpropagation of the neural network, and uses the existing gradient descent algorithm to calculate the gradient of the network weight of MPQ-CNN.

[0065] Step 6-1-1: Assign W to W t-1 , assign to Calculate the loss function on which the gradient optimization of the network weight W t-1 depends.

[0066] The loss function in Formula 1 is equivalent to as shown in

[0067] Formula 3:

[0068]

[0069] where N represents the total number of original RGB image samples in the current batch. x i represents a tensor of dimension [1, C, H, W] corresponding to the i-th original RGB image sample in the training dataset.

[0070] L CE (·,·) represents the cross-entropy loss function, and its calculation method is a and b are vectors of length K respectively. K represents the total number of image categories in the image classification task. In the RGB image classification task of the ImageNet dataset, a represents the probability distribution of the image categories predicted by the MPQ-CNN for the original RGB image samples, and b represents the one-hot encoded vector of the true category of the original RGB image samples or the probability distribution of the image categories predicted by the MPQ-CNN for the original RGB image samples. In Formula 3, t i is the one-hot encoded vector of the true category of the original RGB image samples in the training dataset, and L CE (a, b) can be directly calculated by . Specifically, in the function, the left side of the semicolon represents the input tensor, and the right side represents the neural network parameters of the current MPQ-CNN, that is, the network weights and the quantization parameters of the quantizer. represents the network weights W t-1 and the bit-width combination B t-1 trained in the (t - 1)-th iteration. Fixing B t-1 and then the probability distribution of the MPQ-CNN predicting the image category of the i-th original RGB image sample in the training dataset. Among them, before starting the j-th round of training, the optimal network weights obtained after the last iteration of the (j - 1)-th round and the quantization parameters of the quantizer that determine the optimal bit-width combination are used to initialize the network weights W0 of the first iteration in the j-th round and the quantization parameters of the quantizer that determine the bit-width combination B0 of the first iteration respectively. Importantly, one round contains multiple iterations. In the present invention, one round (Epoch) means that all the original RGB image sample tensors in the training dataset or the validation dataset have completed one pass from step 3 to step 14 in the MPQ-CNN; one iteration (Iteration) means that the original RGB image sample tensors in the training dataset or the validation dataset have completed one pass from step 3 to step 13 in the MPQ-CNN.

[0071] L BC represents the bit-width combination complexity function of MPQ-CNN, which is used to simulate the computational complexity of mixed-precision quantization in computer hardware. In MPQ-CNN, a commonly used measure of bit-width combination complexity is based on the number of floating-point operations (FLOPs) of the convolutional kernel f in the convolutional layer. The calculation of FLOPs is shown in Equation 4:

[0072]

[0073] where |·| represents the number of parameters of the convolutional kernel f. w x and h x are the width and height of the input tensor x of the convolutional kernel f respectively, and s is the stride of the convolutional kernel. The stride is the step size at which the convolutional kernel f moves on the image.

[0074] When the bit-widths of the convolutional kernel f and the activation function a are low, the bit-width complexity of the convolutional kernel f can be represented by the number of bit operations (BitOps). The calculation of BitOps is shown in Equation 5:

[0075]

[0076] where b f and b a represent the bit-widths of the network weights of the CNN and the output of the activation function respectively.

[0077] In the loss function of Equation 1, the bit-width combination complexity function of MPQ-CNN is multiplied by the pre-set hyperparameter ratio β and added to the sum of the outputs of the cross-entropy loss function . In Figure 3 , the original cross-entropy loss is the sum of the outputs of the cross-entropy loss function in Equation 3, and the complexity loss is the output of the bit-width combination complexity function in Equation 3.

[0078] Step 6-1-2: Calculate the gradient of the MPQ-CNN network weights with respect to the loss function according to the existing SGD algorithm, as shown in Equation 6:

[0079]

[0080] where is the gradient symbol, and the subscript indicates the derivative with respect to W t-1 . The gradient is equivalent to the loss function The partial derivative of the network weights W t-1 , that is

[0081] Step 6-2: Update the network weights W of the MPQ-CNN to be optimized in the current t-th iteration t-1 , as shown in Equation 7:

[0082]

[0083] where η W represents the learning rate used to optimize the network weights in the existing gradient descent optimizer. In the current iteration, optimize the MPQ-CNN network weights W along the gradient descent direction that minimizes the loss function t-1 , and obtain the optimized W t . W t is the network weights to be optimized in the next iteration.

[0084] Step 7: Based on the currently optimized MPQ-CNN network weights W t , use the existing adversarial attack method such as the Projected Gradient Descent (PGD) algorithm to generate the optimal perturbation score i corresponding to each original RGB image sample tensor x in the training dataset of the current iteration, as Figure 4 and shown in Equation 8:

[0085]

[0086] where δ i is the perturbation score tensor added to the i-th original sample tensor of the RGB images in the training dataset through the existing PGD algorithm. The tensor dimension of δ i is the same as that of x i . ε is the perturbation threshold, and ||·|| ∞ represents the infinity norm of the tensor. is the adversarial sample corresponding to x i generated by δ i , and the tensor dimension of is also the same as that of x i . Fix the network weights W t optimized in the t-th iteration and the bit-width combination B t-1 optimized in the (t - 1)-th iteration, represents that the input tensor is x i and the MPQ-CNN solves for the perturbation score tensor δ that maximizes the sum of the outputs of N loss functions . Among them,​ The optimal perturbation score tensor representing the original sample tensors of N RGB images in the training dataset , with the constraint that the infinity norm of is less than the perturbation threshold ε.

[0087] Step 7-1: Initialize the counter cnt to 3 according to experience.

[0088] Step 7-2: Initialize the adversarial sample tensor and the perturbation score tensor δ i of the i-th RGB image in the training dataset. As shown in Equation 9:

[0089]

[0090] where the perturbation score tensor δ i has element values in the range [-∈, ∈].

[0091] Step 7-3: Calculate and the output of the cross-entropy loss function i of x . As shown in Equation 10:

[0092]

[0093] where and represent the probability distributions of the predicted classes of the adversarial samples and the original samples of the RGB images in the training dataset corresponding to the input tensors t and x t-1 under the network weights W t-1 optimized at the current t-th iteration and the bit-width combination B t optimized at the (t-1)-th iteration. After fixing B and W i , the MPQ-CNN respectively calculates the probability distributions of the predicted classes of the adversarial samples and the original samples of the RGB images in the training dataset corresponding to the input tensors . Different from the calculation method of the cross-entropy loss function in Equation 3 , in the cross-entropy loss function of Equation 10, b uses the probability distribution of the predicted class of the adversarial samples of the RGB images in the training dataset to replace the one-hot encoded vector of the true class of the original samples of the RGB images in the training dataset, and uses

[0094] Step 7-4: Calculate the gradient g of the output of the loss function in Step 7-3 with respect to the adversarial samples i of the RGB images in the training dataset. Let As shown in Equation 11:

[0095]

[0096] Among them, is the gradient symbol, and the subscript indicates taking the derivative with respect to . The gradient g i is equivalent to the partial derivative of the loss function

[0097] with respect to the adversarial sample tensor of the RGB images in the training dataset, that is

[0098] Step 7-5: Update the adversarial sample tensor i along the ascending direction of the gradient g as shown in Equation 12:

[0099]

[0100] where sign is the sign function and α is the perturbation amplitude. The objective of Equation 12 is to perturb the original samples of the RGB images in the training dataset by the perturbation score, so as to maximize the impact on the prediction results of the MPQ-CNN.

[0101] Step 7-6: Calculate the perturbation score tensor δ i . As shown in Equation 13:

[0102]

[0103] where represents the element-wise subtraction of the adversarial sample tensor of the RGB images in the training dataset and its corresponding original sample tensor x i .

[0104] Step 7-7: Truncate the perturbation score tensor δ i to the perturbation interval [-ε, +ε]. As shown in Equation 14:

[0105]

[0106] where represents all the element values in the perturbation score tensor δ i with dimension [1, C, H, W].

[0107] Step 7-8: Update the adversarial sample tensor as shown in Equation 15:

[0108]

[0109] where x i + δ i represents the original sample tensor x i of the RGB images in the training dataset and its corresponding perturbation score tensor δ​i Add the values of each element together.

[0110] Steps 7 - 9: Truncate all the element values in the adversarial sample tensor to the range [0, 1]. As shown in Equation 16:

[0111]

[0112] where, represents all the element values in the adversarial sample tensor of the RGB images in the training dataset with dimensions [1, C, H, W]

[0113] Step 7 - 10: Let cnt = cnt - 1.

[0114] Step 7 - 11: At the same time, for the N original sample tensors x of the RGB images in the training dataset i Repeat Steps 7 - 3 to 7 - 10 until cnt = 0.

[0115] Step 8: When Step 7 - 11 ends, obtain the optimal perturbation score tensor from Step 7 - 7 Add it element - by - element with the original sample tensor of the RGB images in the training dataset using to generate the optimal RGB image adversarial sample tensor

[0116]

[0117] where, represents the element - by - element addition of the original sample tensor x of the RGB images in the training dataset i and the optimal perturbation score tensor

[0118] Step 9: Input the adversarial sample tensor into MPQ - CNN.

[0119] Step 10: In MPQ - CNN, perform the forward propagation of the neural network for the adversarial sample tensor

[0120] Step 11: Fix the network weights of MPQ - CNN, that is, do not optimize the weights W of MPQ - CNN trained in the t - th iteration using gradients t . Optimize the bit - width combination of the current MPQ - CNN, that is, use the SGD algorithm to optimize the quantization parameters of all the weight quantizers and activation function quantizers that determine the bit - width combination of MPQ - CNN.

[0121] The process of Step 11 is shown in Equation 18:

[0122] ​​​

[0123] In the j-th round, denotes the network weights W of the fixed MPQ-CNN optimized at the t-th iteration, t , that is, the network weights of the MPQ-CNN do not participate in gradient optimization in step 10; B t-1 denotes the bit-width combination to be optimized by the MPQ-CNN at the t-th iteration. denotes the loss function relied on when optimizing the bit-width combination B of the MPQ-CNN using the existing SGD algorithm. Specifically, the t-1 left side of the semicolon denotes the input tensor relied on when solving this loss function; the and B t-1 denote the neural network parameters of the MPQ-CNN relied on when solving the output of this loss function, that is, the network weights and the quantization parameters of the quantizer. denotes the bit-width combination B that minimizes the output of the loss function in the MPQ-CNN. denotes the process of determining the quantization parameters of the quantizer for the bit-width combination of the MPQ-CNN through gradient optimization by formula 18 under the fixed network weights W t of the MPQ-CNN.

[0124] Step 11-1: The MPQ-CNN performs neural network backpropagation and calculates the gradient of the quantization parameters of the quantizer of the MPQ-CNN using the existing gradient descent algorithm.

[0125] Step 11-1-1: Calculate the output of the loss function relied on for gradient-optimizing the bit-width combination B t-1 As shown in formula 19:

[0126]

[0127] where denotes the probability distribution of the predicted class of the adversarial sample of the RGB image corresponding to the input tensor t when the MPQ-CNN fixes W t-1 under the network weights W optimized at the current t-th iteration and the bit-width combination B optimized at the (t - 1)-th iteration. t t is the one-hot encoded vector of the true class of the original sample of the RGB image in the training dataset. The loss function in formulas 18 and 19 i is equal to the sum of the bit-width combination complexity function L multiplied by the hyperparameter ratio β BC (B t-1 ) and the output sum of the cross-entropy loss function multiplied by the hyperparameter ratio α Addition.

[0128] Step 11-1-2: Calculate the quantization parameters of the MPQ-CNN weight quantizer and activation quantizer for the loss function gradient according to the existing SGD algorithm.

[0129] Taking the existing linear quantizer as an example, the bit widths of the MPQ-CNN network weights W and the activation function output y are directly determined by the quantization factors q max and q min of the linear quantizer.

[0130] Step 11-1-2-1: Calculate the gradient of the maximum value q max of the fixed-point number range. As shown in Equation 20:

[0131]

[0132] where is the gradient symbol, and the subscript indicates the derivative with respect to q max . The gradient is equivalent to the partial derivative of the loss function with respect to the quantization parameter q max , that is

[0133] Step 11-1-2-2: Calculate the gradient of the minimum value q min of the fixed-point number range. As shown in Equation 21:

[0134]

[0135] where is the gradient symbol, and the subscript indicates the derivative with respect to q min . The gradient is equivalent to the partial derivative of the loss function with respect to the quantization parameter q min , that is

[0136] Step 11-1-2-3: Let r = r max - r min , and calculate the gradient of the floating-point number interval length r. As shown in Equation 22:

[0137]

[0138] where is the gradient symbol, and the subscript indicates the derivative with respect to r. The gradient is equivalent to the partial derivative of the loss function with respect to the quantization parameter r, that is

[0139] Step 11-2: Optimize the quantization parameters of the MPQ-CNN linear quantizer in the t-th iteration using the existing SGD algorithm.

[0140] Step 11-2-1: Update the maximum value of the fixed-point number range of the quantization parameters of the linear quantizer to be optimized in the current t-th iteration As shown in Equation 23:

[0141]

[0142] where η B represents the learning rate used in the existing gradient descent optimizer to optimize the quantization parameters of the quantizer. In the current iteration, optimize the quantization parameters of the linear quantizer along the gradient descent direction that minimizes the loss function to obtain the optimized which is the maximum value of the fixed-point number range to be optimized in the next iteration.

[0143] Step 11-2-2: Update the minimum value of the fixed-point number range of the quantization parameters of the linear quantizer to be optimized in the current t-th iteration As shown in Equation 24:

[0144]

[0145] where η B represents the learning rate used in the existing gradient descent optimizer to optimize the quantization parameters of the quantizer. In the current iteration, optimize the quantization parameters of the linear quantizer along the gradient descent direction that minimizes the loss function to obtain the optimized which is the minimum value of the fixed-point number range to be optimized in the next iteration.

[0146] Step 11-2-3: Update the floating-point interval length r of the quantization parameters of the linear quantizer to be optimized in the current t-th iteration t-1 . As shown in Equation 25:

[0147]

[0148] where η B represents the learning rate used in the existing gradient descent optimizer to optimize the quantization parameters of the quantizer. In the current iteration, optimize the quantization parameter r of the linear quantizer along the gradient descent direction that minimizes the loss function to obtain the optimized r t-1 . r t . r t is the floating-point interval length to be optimized in the next iteration.​​

[0149] Step 12: According to the existing quantization formula, use the optimized quantization parameters and to update the bit-width combination B t-1 to B t .

[0150] The operation of updating the bit-width combination of MPQ-CNN means that according to different quantizer quantization parameters and update the bit-width of all network weights deployed with weight quantizers and the bit-width of the output of the activation function deployed with activation quantizers in MPQ-CNN. The bit-width combination of network weights and the bit-width combination of the output of the activation function in MPQ-CNN together constitute the bit-width combination of MPQ-CNN. Taking the symmetric quantization of the linear quantizer as an example. Bit If then b = 8. If 8 is the bit-width of the network weight or the output of the activation function, it means that the output of the network weight or the activation function is stored and operated in 8 bits in the computer hardware.

[0151] Step 13: Determine whether t is equal to the preset maximum number of iterations Iteration max , if not, let t = t + 1, and return to Step 3; if so, execute Step 14.

[0152] Step 14: Save the optimal bit-width combination max updated in the Iteration th iteration and the optimal network weights Let j = j + 1.

[0153] In this embodiment, while saving the optimal bit-width combination max updated in the Iteration of MPQ-CNN, it is also necessary to save all the quantization parameters of the weight quantizers and activation quantizers that determine for the forward propagation and backward propagation of the neural network on the training dataset and the forward propagation of the neural network on the validation dataset.

[0154] Step 15: Repeat Step 3 to Step 14 until j is equal to the preset maximum number of epochs Epoch max .

[0155] Step 16: According to the optimal bit-width combination and network weights searched in each round saved in Step 14Evaluate on the validation dataset, that is, select the optimal bitwidth combination B by minimizing the cross-entropy loss of the validation dataset * 。

[0156] The specific implementation of step 16 is shown in Equation 26:

[0157]

[0158] where D val represents all the original sample tensors of RGB images after preprocessing the validation dataset for the current iteration. Epoch max represents the maximum number of training epochs during the training of MPQ-CNN. represents the set composed of the optimal bitwidth combinations obtained in each epoch from 1 to Epoch max rounds, and select the bitwidth combination B that minimizes the cross-entropy loss of the validation dataset from the set . The loss function of the validation dataset * . The D on the left side of the semicolon in the loss function represents the original sample tensor of the RGB image of the input validation dataset, and the right side of the semicolon indicates that the calculation of the loss function of the validation dataset depends on the optimal network weights obtained in this round val and the optimal bitwidth combination optimized using and during the calculation process without optimization and and represents the process of MPQ-CNN selecting the optimal bitwidth combination by Equation 26 under the fixed optimal network weights and bitwidth combination.

[0159] Step 16-1: Input the original sample tensor of the RGB image of the validation dataset preprocessed in one iteration into MPQ-CNN.

[0160] Step 16-2: In the MPQ-CNN of and , perform the forward propagation of the neural network on the original sample tensor of the RGB image of the validation dataset for the current iteration.

[0161] Step 16-3: Calculate the output of the loss function . As shown in Equation 27:

[0162]

[0163] where N represents the total number of original samples of RGB images in a batch, that is, in a corresponding validation dataset in one iteration, represents the tensor of each original sample of the RGB image in the validation dataset. Denote the network weights optimized in the current j-th round and bit-width combination Under the condition of fixing and For the input tensor The probability distribution of the predicted categories of the original samples of the RGB images in the corresponding validation dataset Is the one-hot encoded vector of the true categories of the original samples of the RGB images in the validation dataset Denote and The output of the cross-entropy loss function jointly calculated

[0164] Step 16-4: Repeat Step 16-1 to Step 16-3 until all the original samples of the RGB images in the validation dataset have been completely trained by MPQ-CNN once, and obtain the cross-entropy loss of the validation dataset

[0165] Step 16-5: Repeat Epoch max times of Step 16-4. Until Epoch max different and The combinations of MPQ-CNN have all been evaluated by the validation dataset, and finally find the and combinations of MPQ-CNN that minimize the cross-entropy loss of the validation dataset, and obtain the optimal bit-width combination B of MPQ-CNN * .

[0166] The optimal bit-width combination B * Is the bit-width combination adopted when deploying MPQ-CNN on computer hardware

[0167] It should be understood that those skilled in the art, inspired by the technical concept of the present invention and without departing from the content of the present invention, can also make various improvements or transformations according to the above content, and this still falls within the protection scope of the present invention

Claims

1. A mixed precision quantization method based on image adversarial sample perception and gradient optimization, characterized in that: The method comprises the following steps: S1. Input a batch of RGB image original samples and fix the bit width combination of MPQ-CNN, and use the gradient descent algorithm to optimize the network weight W of MPQ-CNN optimized in the previous iteration. t-1 , get the network weight W optimized for the current iteration t ; S2, based on the original RGB image samples and network weights W of the current batch t , use the adversarial attack method to generate the optimal perturbation score corresponding to each RGB image original sample in the current batch, and then generate adversarial samples of the RGB image original samples based on the optimal perturbation score; S3. Use MPQ-CNN to predict the category of the RGB image adversarial sample, and use the cross entropy loss generated by the probability distribution of the predicted category and the one-hot encoding vector of the true category to guide the optimization of the quantization parameters of the MPQ-CNN quantizer. Use the optimized quantization parameters to update the bit width combination of MPQ-CNN to obtain the bit width combination B updated in the current iteration. t ; S4, repeating steps S1 to S3 until the preset maximum number of iterations is reached, completing a round of MPQ-CNN training process; S5, repeating steps S1 to S4 until a preset maximum number of rounds is reached; S6. According to the quantization parameters and network weights of the optimal quantizer of MPQ-CNN searched in each round, an evaluation is performed on the validation data set, and finally the network weights of MPQ-CNN and the quantization parameters of the quantizer that minimize the cross entropy loss of the validation data set are found, and then the optimal bit width combination of MPQ-CNN is determined.

2. The mixed precision quantization method according to claim 1, characterized in that: Step S1 includes the following steps: Step S1.1: Input a batch of RGB image original sample tensors into MPQ-CNN for neural network forward propagation; the dimension of a batch of RGB image original sample tensors is expressed as [N, C, H, W], where N is the total number of image samples in the current batch, C represents the channel of the RGB image, and H and W are the height and width of the RGB image respectively; Step S1.2: Fix the bit width combination of MPQ-CNN and use the gradient descent algorithm to optimize the network weight W of MPQ-CNN optimized in the previous iteration t-1 , and get the network weight W optimized for this iteration t .

3. The mixed precision quantization method according to claim 2, characterized in that: In step S1.2, MPQ-CNN combines Next, the current MPQ-CNN network weights are gradient optimized based on the following formula to obtain the trained network weights; Where D tr Represents the original RGB image sample tensor of the current batch; W represents the network weight to be optimized in MPQ-CNN; It means that the bit width combination B of MPQ-CNN is fixed; It represents the loss function that is relied on when optimizing the MPQ-CNN network weight W using the gradient descent algorithm. The left side of the semicolon represents the input tensor that is relied on when solving the loss function; the right side of the semicolon represents the neural network parameters of the MPQ-CNN that are relied on when solving the loss function, that is, the network weights and the quantization parameters of the quantizer; Indicates the loss function in solving MPQ-CNN The network weight with the smallest output; Assign W to W t-1 , Assigned to Then the loss function for: Where N represents the total number of original RGB image samples in the current batch, x i Represents a tensor with dimensions [1, C, H, W] corresponding to the original sample of the i-th RGB image in the training dataset; W t-1 represents the optimized network weight in the t-1th iteration; B t-1 Indicates the optimized bit width combination in the t-1th iteration; Indicates the network weight W optimized at the t-1th iteration t-1 And bit width combination B t-1 Next, fix B t-1 The MPQ-CNN predicts the probability distribution of the image category of the i-th RGB image original sample in the training dataset; t i is the one-hot encoding vector of the true category of the original sample of the RGB image in the training dataset; L BC represents the bit width combination complexity function of MPQ-CNN; β is the preset hyperparameter ratio.

4. The mixed precision quantization method according to claim 1, characterized in that: The adversarial attack method is a projected gradient descent algorithm.

5. The mixed precision quantization method according to claim 1, characterized in that: The quantization parameters of the MPQ-CNN quantizer are quantization parameters of all weight quantizers and activation function quantizers that determine the bit width combination of the MPQ-CNN.

6. The mixed precision quantization method according to claim 1, characterized in that: The bit width combination of the MPQ-CNN is composed of the bit width combination of the network weights in the MPQ-CNN and the bit width combination of the activation function output.

7. The mixed precision quantization method according to claim 1, characterized in that: Step S3 includes the following steps: Step S3.1: Input the adversarial sample tensor into MPQ-CNN and perform neural network forward propagation on the adversarial sample tensor of the current batch; Step S3.2: fix the network weights of the MPQ-CNN, and use the gradient descent algorithm to optimize the gradient to determine the quantization parameters of the quantizer of the current MPQ-CNN bit width combination; Step S3.3: Use the optimized quantizer parameters to combine the bit width B t-1 Update to B t .

8. The mixed precision quantization method according to claim 7, characterized in that: The process of step S3.2 is expressed as follows: In round j, represents the network weight W optimized by the fixed MPQ-CNN at the tth iteration t ; B t-1 represents the bit width combination B that is not optimized by MPQ-CNN in the tth iteration t-1 ; Indicates the use of gradient descent algorithm to optimize the bit width combination B of MPQ-CNN t-1 The loss function that θ depends on; Represents the input tensor that the loss function depends on when solving it; and B t-1 Represents the neural network parameters of MPQ-CNN that the output of the loss function depends on, namely, the network weights and the quantization parameters of the quantizer; Indicates the loss function in solving MPQ-CNN The output with the smallest bit width combination B; Represents MPQ-CNN with fixed network weights W t The process of determining the quantization parameters of the quantizer of the MPQ-CNN bit width combination is performed by gradient optimization of formula 18; and there is in, Represents the network weight W optimized by MPQ-CNN at the current tth iteration t and the bit width combination B optimized in the t-1th iteration t-1 Next, fix W t After MPQ-CNN, the input tensor The probability distribution of the corresponding RGB image adversarial sample prediction category; t i is the one-hot encoding vector of the true category of the original sample of the RGB image in the training dataset; α and β are both hyperparameter ratios; L BC (B t-1 ) is the bit width combination complexity function; Outputs the sum for the cross entropy loss function.

9. The mixed precision quantization method according to any one of claims 1, 2, 7 and 8, characterized in that: The gradient descent algorithm is a stochastic gradient descent algorithm.