Model quantification method and device

By determining the quantization sensitivity and bit width probability information of the convolutional layer, the model quantization process is optimized, and the problem of difficulty in balancing accuracy and memory usage in the existing technology is solved, efficient model quantization is achieved, and deployment efficiency on edge devices is improved.

CN120278205APending Publication Date: 2025-07-08SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146933.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing model quantization methods are difficult to achieve a fine-grained balance between accuracy and memory footprint, resulting in low precision and high memory footprint of quantized models, which cannot effectively improve the deployment efficiency of models on edge devices with limited computing power and space storage capabilities.

Method used

By determining the quantization sensitivity of the convolutional layer of the model to be quantized, the quantization bit width probability information of each convolutional layer is calculated, and the model is quantized based on these probability information, and the quantitative model that meets the preset conditions is selected. The quantization sensitivity estimation value (QSE) and mixed precision quantization search method are used, and the model quantization process is optimized in combination with the quantitative memory usage limit.

Benefits of technology

The effect and efficiency of model quantization are improved, the time-consuming search of quantitative models is reduced, the verification accuracy of the model is maintained, and the adaptability is strong, and it has good generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278205A_ABST
    Figure CN120278205A_ABST
Patent Text Reader

Abstract

The invention provides a model quantification method and device. The model quantification method comprises the following steps: determining the quantification sensitivity of at least one convolutional layer of a to-be-quantified model; according to the quantization sensitivity of each convolution layer in the at least one convolution layer, respectively determining quantization bit width probability information of each convolution layer in the at least one convolution layer, the quantization bit width probability information of each of the at least one convolution layer represents the probability that each of the at least one convolution layer selects each of a plurality of candidate quantization bit widths; quantizing the to-be-quantized model based on the quantization bit width probability information of each convolution layer in the at least one convolution layer to obtain a preset number of quantization candidate models; and selecting a quantization model meeting a preset condition from the preset number of quantization candidate models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology. More specifically, the present disclosure relates to a model quantization method and apparatus. Background Art

[0002] Neural networks are increasingly being used to solve complex tasks in fields such as computer vision and natural language processing. In order to better deploy these neural networks on edge devices with limited computing power and spatial storage capacity and improve the model inference speed, it is usually necessary to quantize the model to minimize the model complexity and storage space occupied.

[0003] In the deployment work in the field of deep learning, model quantization is indispensable. The model quantization method mainly realizes the compression of the model size and the acceleration of the operation by using fixed-point numbers to represent the weight values in the model and the activation values during operation. The traditional model quantization method quantizes the weight values and activation values of the entire model to a certain fixed bit width. High-bit quantization can ensure high precision but also has a larger memory occupancy and computational latency, while low-bit quantization has lower precision but smaller memory occupancy and computational latency. Therefore, quantization with a fixed bit width can never achieve a fine-grained balance between precision and memory occupancy, resulting in a lower precision and higher memory occupancy of the quantized model, and thus the effect and efficiency of the model quantization method are not good. Summary of the Invention

[0004] According to an exemplary embodiment of the present disclosure, there is provided a model quantization method, including: determining the quantization sensitivity of at least one convolutional layer of a model to be quantized; respectively determining the quantization bit width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, where the quantization bit width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit width from a plurality of candidate quantization bit widths; quantizing the model to be quantized based on the quantization bit width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantized candidate models; and selecting a quantized model that meets a preset condition from the preset number of quantized candidate models.

[0005] Optionally, the step of determining the quantization sensitivity of at least one convolutional layer of the model to be quantized may include: parsing the model to be quantized to obtain the convolutional layer information of each convolutional layer in the at least one convolutional layer, where the convolutional layer information includes the number of input channels, the number of output channels, and the sampling matrix; and respectively calculating the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-training parameters of each convolutional layer in the at least one convolutional layer.

[0006] Optionally, the step of calculating the quantization sensitivity of each convolutional layer in the to-be-quantized model based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, calculating the Hessian matrix of the pre-trained parameters of each convolutional layer; and calculating the quantization sensitivity of each convolutional layer based on the Hessian matrix, the number of input channels, the number of output channels, and the sampling matrix.

[0007] Optionally, the step of respectively determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, determining the quantization sensitivity data distribution of each convolutional layer based on the quantization sensitivity of each convolutional layer; determining the quantization sensitivity envelope points from the quantization sensitivity data distribution of each convolutional layer; and determining the quantization bit-width probability information of each convolutional layer based on the quantization sensitivity envelope points of each convolutional layer.

[0008] Optionally, the step of quantizing the to-be-quantized model based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer may include: respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer to determine a quantization candidate model of the to-be-quantized model; and determining the preset number of quantization candidate models from the quantization candidate models of the to-be-quantized model based on a preset quantization memory occupancy limit, wherein between every two quantization candidate models among the preset number of quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit-widths.

[0009] Optionally, the step of determining a quantization candidate model of the to-be-quantized model by respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, sequentially determining the quantization bit-width of each convolutional layer as the candidate quantization bit-widths in the order from high to low of the quantization bit-width probabilities in the quantization bit-width probability information of each convolutional layer to quantize each convolutional layer, thereby obtaining the quantization candidate model.

[0010] Optionally, the step of determining the preset number of quantization candidate models from the quantization candidate models of the to-be-quantized model based on a preset quantization memory occupancy limit includes: determining whether the quantization candidate models of the to-be-quantized model satisfy the preset quantization memory occupancy limit; based on determining that the quantization candidate models of the to-be-quantized model satisfy the preset quantization memory occupancy limit, taking the quantization candidate models that satisfy the preset quantization memory occupancy limit as the first quantization candidate models until the number of the first quantization candidate models reaches a first preset number; performing a repeatability determination on the first preset number of first quantization candidate models, and selecting non-repeating first quantization candidate models among the first preset number of first quantization candidate models as the second preset number of second quantization candidate models, wherein, between every two of the second preset number of second quantization candidate models, one or more of the at least one convolutional layer have different candidate quantization bit widths; selecting the preset number of second quantization candidate models that satisfy the preset quantization memory occupancy limit from the second preset number of second quantization candidate models as the preset number of quantization candidate models.

[0011] Optionally, the step of selecting a quantization model that meets a preset condition from the preset number of quantization candidate models may include: performing accuracy verification on the preset number of quantization candidate models, and taking the quantization candidate model with the highest accuracy among the preset number of quantization candidate models as the quantization model that meets the preset condition.

[0012] Optionally, the at least one convolutional layer may not include the first convolutional layer in the to-be-quantized model.

[0013] According to an exemplary embodiment of the present disclosure, there is provided a model quantization apparatus, including: a sensitivity determination unit configured to determine the quantization sensitivity of at least one convolutional layer of a to-be-quantized model; a probability information determination unit configured to respectively determine the quantization bit width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit width among a plurality of candidate quantization bit widths; a model quantization unit configured to quantize the to-be-quantized model based on the quantization bit width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantization candidate models; and an accuracy selection unit configured to select a quantization model that meets a preset condition from the preset number of quantization candidate models.

[0014] Optionally, the sensitivity determination unit may be configured to: parse the model to be quantized to obtain convolution layer information of each convolution layer in the at least one convolution layer; calculate the quantization sensitivity of each convolution layer in the at least one convolution layer of the model to be quantized respectively based on the convolution layer information of each convolution layer in the at least one convolution layer and the pre-trained parameters of each convolution layer in the at least one convolution layer.

[0015] Optionally, the sensitivity determination unit may be configured to: for each convolution layer in the at least one convolution layer, calculate the Hessian matrix of the pre-trained parameters of each convolution layer; calculate the quantization sensitivity of each convolution layer based on the Hessian matrix of the pre-trained parameters of each convolution layer, the number of input channels, the number of output channels, and the sampling matrix.

[0016] Optionally, the probability information determination unit may be configured to: for each convolution layer in the at least one convolution layer, determine the quantization sensitivity data distribution of each convolution layer; determine the quantization sensitivity envelope points from the quantization sensitivity data distribution of each convolution layer; determine the quantization bit-width probability information of each convolution layer based on the quantization sensitivity envelope points of each convolution layer.

[0017] Optionally, the model quantization unit may be configured to: determine the quantized candidate models of the model to be quantized by quantizing each convolution layer in the at least one convolution layer respectively based on the quantization bit-width probability information of each convolution layer in the at least one convolution layer; determine the preset number of quantized candidate models from the quantized candidate models of the model to be quantized based on a preset quantization memory occupancy limit, wherein, between any two of the preset number of quantized candidate models, one or more convolution layers in the at least one convolution layer have different candidate quantization bit-widths.

[0018] Optionally, the model quantization unit may be configured to: for each convolution layer in the at least one convolution layer, sequentially determine the quantization bit-width of each convolution layer as the candidate quantization bit-widths in the order of decreasing quantization bit-width probability in the quantization bit-width probability information of each convolution layer, so as to quantize each convolution layer to obtain the quantized candidate models.

[0019] Optionally, the model quantization unit may be configured to: determine whether a quantization candidate model of the model to be quantized meets the preset quantization memory occupancy limit; based on determining that the quantization candidate model of the model to be quantized meets the preset quantization memory occupancy limit, use the quantization candidate model that meets the preset quantization memory occupancy limit as a first quantization candidate model until the number of first quantization candidate models reaches a first preset number; perform a repeatability determination on the first preset number of first quantization candidate models, and select non-repeating first quantization candidate models among the first preset number of first quantization candidate models as a second preset number of second quantization candidate models, wherein, between every two of the second preset number of second quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit widths; select the preset number of second quantization candidate models that meet the preset quantization memory occupancy limit from the second preset number of second quantization candidate models as the preset number of quantization candidate models.

[0020] Optionally, the accuracy selection unit may be configured to: perform accuracy verification on the preset number of quantization candidate models, and use the quantization candidate model with the highest accuracy among the preset number of quantization candidate models as the quantization model that meets the preset conditions.

[0021] Optionally, the at least one convolutional layer may not include the first convolutional layer in the model to be quantized.

[0022] According to an exemplary embodiment of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, which when executed by a processor, implements the model quantization method according to the exemplary embodiment of the present disclosure.

[0023] According to an exemplary embodiment of the present disclosure, there is provided a computing device, including: at least one processor; at least one memory storing a computer program, which when executed by the at least one processor, implements the model quantization method according to the exemplary embodiment of the present disclosure.

[0024] According to an exemplary embodiment of the present disclosure, there is provided a computer program product, and instructions in the computer program product can be executed by a processor of a computer device to complete the model quantization method according to the exemplary embodiment of the present disclosure.

[0025] A model quantization method and apparatus according to an exemplary embodiment of the present disclosure determine quantization sensitivities of at least one convolutional layer of a model to be quantized, and respectively determine quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, where the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width from a plurality of candidate quantization bit-widths, quantize the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantized candidate models, and select a quantized model that meets preset conditions from the preset number of quantized candidate models, thereby improving the effect and efficiency of model quantization.

[0026] Additional aspects and / or advantages of the inventive concept will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the drawings of the exemplary embodiments, in which: Figure 1 FIG. shows a flowchart of a model quantization method according to an exemplary embodiment of the present disclosure; Figure 2 FIG. shows a flowchart of calculating quantization sensitivity according to an exemplary embodiment of the present disclosure; Figure 3 FIG. shows a flowchart of a quantization sensitivity processor according to an exemplary embodiment of the present disclosure; Figure 4 FIG. shows a schematic diagram of the operation of a quantization model searcher according to an exemplary embodiment of the present disclosure; Figure 5 FIG. shows a block diagram of a model quantization apparatus according to an exemplary embodiment of the present disclosure; and Figure 6 FIG. shows a schematic diagram of a computing device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Reference will now be made in detail to the exemplary embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals refer to like elements throughout. The following embodiments will be described with reference to the accompanying drawings to explain the present disclosure.

[0029] Currently, there are still various limitations in the mixed-precision quantization methods implemented for models: The random method based on reinforcement learning has a large blindness. After determining the mixed bit-width of the convolutional layer, a complete training step needs to be performed on each generated candidate model. Therefore, it will consume a large amount of time to evaluate the candidate models, and the time to obtain an excellent model is too slow. The method based on quantization sensitivity analysis does not fully consider the data distribution between the convolutional layers of the model. The selection of the mixed-precision quantization bit-width is easily affected by the distribution of model parameters. Simply setting thresholds or extreme values to judge sensitivity will ignore quite a lot of useful information, making the distribution of the low-precision quantization layer relatively concentrated, resulting in serious accuracy loss of the final quantization model. Therefore, the present disclosure proposes a mixed-precision quantization search method based on the quantitative sensitivity estimates (QSE) of the convolutional layers of the model (hereinafter simply referred to as the QSE method), and designs an evaluation index for the memory occupancy of the mixed quantization model (Quantify Memory Access Cost Limit, hereinafter referred to as B-MAC), which can specify the range of model memory occupancy and search for different mixed-precision quantization models. This method can be executed by an electronic device including the following three modules: a preprocessing module (Preprocessing Module, hereinafter referred to as PM), a search module (Search Module, hereinafter referred to as SM), and a quantization module (Quantization Module, hereinafter referred to as QM). The preprocessing module parses the model to be quantized and calculates the data of the quantitative sensitivity estimates of the convolutional layers of the model; subsequently, the search module modifies the data size of the quantization bit-width probability of the quantization search according to the data of the quantitative sensitivity estimates, expands the number of candidate models and screens out a sufficient number of excellent models; finally, all the excellent candidate models are verified by the quantization module for the final accuracy.

[0030] Figure 1 FIG. shows a flowchart of a model quantization method according to an exemplary embodiment of the present disclosure. Figure 2 FIG. shows a flowchart of calculating quantization sensitivity according to an exemplary embodiment of the present disclosure. Figure 3 FIG. shows a flowchart of a quantization sensitivity processor according to an exemplary embodiment of the present disclosure. Figure 4 FIG. shows a working schematic diagram of a quantization model searcher according to an exemplary embodiment of the present disclosure.

[0031] Referring to Figure 1 , in step S101, determine the quantization sensitivity of at least one convolutional layer of the model to be quantized.

[0032] In an exemplary embodiment of the present disclosure, the at least one convolutional layer may not include the first convolutional layer in the model to be quantized. For example, the model to be quantized may include a first convolutional layer, a fully connected layer, and the at least one convolutional layer.

[0033] In an exemplary embodiment of the present disclosure, the step of determining the quantization sensitivity of at least one convolutional layer of the model to be quantized may include: parsing the model to be quantized to obtain convolutional layer information of each convolutional layer in the at least one convolutional layer, where the convolutional layer information includes the number of input channels, the number of output channels, and the sampling matrix; calculating the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized respectively based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters (e.g., pre-trained weights) of each convolutional layer in the at least one convolutional layer, thereby improving the speed of determining the quantization sensitivity of the convolutional layer.

[0034] In an exemplary embodiment of the present disclosure, the step of calculating the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized respectively based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, calculating the Hessian matrix of the pre-trained parameters of each convolutional layer; calculating the quantization sensitivity of each convolutional layer based on the Hessian matrix, the number of input channels, the number of output channels, and the sampling matrix, thereby improving the speed of determining the quantization sensitivity of the convolutional layer.

[0035] As an example, as Figure 2 shown, the model to be quantized may be analyzed first to obtain convolutional layers 1-n of the model to be quantized (e.g., but not limited to, Figure 2 the 7×7 convolutional layer, 3×3 convolutional layer, 1×1 convolutional layer, linear layer, etc. in Figure 2 ). Then, based on the convolutional layers and pre-trained parameters of the model to be quantized obtained after the analysis, the quantization sensitivity of each convolutional layer in convolutional layers 1-n is calculated. For example, as

[0036] shown, the analyzed convolutional layers and pre-trained parameters are input into a quantization sensitivity calculator to extract all convolutional layer information, and quantization sensitivity data is calculated in combination with the pre-trained parameters corresponding to the convolutional layers.

[0036] For example, for a neural network with m parameters, its gradient vector can be expressed as: . Here, represents an m-dimensional vector, represents the weight the first-order derivative of the variable .

[0037] And the derivative of the gradient, that is, the second-order derivative of the weight, is a matrix: . Here, represents the Hessian matrix, represents a matrix of, i.e., a matrix with m rows and m columns, represents the weight for the variable second derivative, represents the gradient vector for the variable first derivative.

[0038] Generally, the matrix is the Hessian matrix. According to the obtained Hessian matrix of the convolutional layer, the quantization sensitivity (QSE) of the convolutional layer is calculated, and the Hessian matrices of the weights on all channels of the convolutional layer are used for the calculation of the quantization sensitivity. Assume that the Hessian matrix of the weights of the th convolutional layer is , then the corresponding quantization sensitivity calculation formula is as follows: .

[0039] Here, and represent the number of input channels and the number of output channels of the current convolutional layer respectively, represents a sampling matrix that conforms to a Gaussian distribution with a mean of 0 and a variance of 1, is a vector with all elements being 1, the Hessian matrix of the weights of the th convolutional layer, represents the input channel,

[0040] To eliminate the calculation error of the QSE value caused by the randomness of the sampling matrix, multiple quantization sensitivities QSE of each convolutional layer can be calculated in multiple rounds, and their average value is taken. Subsequently, the average value calculated layer by layer is normalized to obtain the final quantization sensitivity QSE of the convolutional layer of the target network. For example, it can be as shown in the following formula: . Here, represents the final quantization sensitivity, represents the average value of multiple rounds of calculations of the quantization sensitivity of each convolutional layer, represents layer by layer (from the first layer to all layers) for to be normalized.

[0041] In step S102, quantization bit-width probability information (also referred to as quantization bit-width weight) of each convolutional layer in the at least one convolutional layer is determined respectively according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer. Here, the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width from a plurality of candidate quantization bit-widths.

[0042] In an exemplary embodiment of the present disclosure, the step of determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, determining the quantization sensitivity data distribution of each convolutional layer based on the quantization sensitivity of each convolutional layer; determining the quantization sensitivity envelope points of each convolutional layer from the quantization sensitivity data distribution of each convolutional layer; and determining the quantization bit-width probability information of each convolutional layer based on the quantization sensitivity envelope points of each convolutional layer.

[0043] In an embodiment of the present disclosure, quantization sensitivity QSE calculations are respectively performed on a MobileNet-V2 network (with a total of 52 convolutional layers), where this network includes 52 convolutional layers except for the first convolutional layer. Since the number of channels included in the deep convolutional layers of the MobileNet-V2 network is extremely large, the predicted value trend of the quantization sensitivity QSE is gradually decreasing, that is, the value of the quantization sensitivity QSE will gradually decrease in distribution magnitude as the number of channels in the convolutional layer increases.

[0044] In an embodiment of the present disclosure, quantization sensitivity QSE calculations are respectively performed on a ResNet-50 network (with a total of 53 convolutional layers), where this network includes 53 convolutional layers except for the first convolutional layer. The value of the quantization sensitivity QSE of the ResNet-50 network starts to sharply decrease to a lower order of magnitude at the 28th layer and has almost no fluctuations in the normalized graph. If the threshold method is adopted, the QSE information of the ResNet-50 network from the 28th layer to the 53rd layer will be meaningless. In addition, due to the modular convolutional structure design of the ResNet-50 network, repeated structures will appear in groups, and the convolutional layers within the same group have similar quantization sensitivity QSE data distributions, while the quantization sensitivity QSE distributions of the convolutional layers between different groups may vary greatly.

[0045] Therefore, in order to eliminate the influence of the QSE data magnitude difference between convolutional layer groups without ignoring the important information of the similarity of the quantization sensitivity QSE data distribution within the convolutional layer group, in the embodiments of the present disclosure, the quantization sensitivity QSE data of the convolutional layer can be processed in the way of envelope points, that is, for the result distribution of the quantization sensitivity QSE data of the network, the upper envelope and the lower envelope are used to obtain the envelope points of the quantization sensitivity QSE data distribution, and the mixed-precision quantization method is used for these envelope points.

[0046] In the embodiments of the present disclosure, the envelope analysis is used to model the quantization sensitivity QSE data distribution, and the mixed-precision quantization method is used for the envelope points to adjust the quantization bit width of this layer. Since the envelope points are representative compared with their surrounding adjacent points and reflect the data distribution information of the group to a certain extent. Compared with the adjacent points of the envelope points, the envelope points have the smallest / largest quantization sensitivity QSE value, and the difference from their true sensitivity is relatively small, making the mixed quantization points evenly distributed and preventing the quantization points from concentrating in the deep convolutional layer.

[0047] Since some convolutional layers will greatly lose the verification accuracy of the quantization model when using low-bit-width quantization (2bit), and too many 2bit quantization layers will severely slow down the convergence speed of the quantization model. Therefore, in the embodiments of the present disclosure, the 2bit quantization bit width probability (that is, quantization weight) can be set to zero, that is, each convolutional layer only makes a choice within 4bit, 8bit, and 16bit. Figure 3 A flowchart of a quantization sensitivity processor according to an exemplary embodiment of the present disclosure is shown. As Figure 3 shown, after the quantization sensitivity (QSE) data of all convolutional layers are input into the quantization sensitivity (QSE) processor, the points passed by the upper envelope in the quantization sensitivity QSE data result distribution will be listed in the 16bit candidate layer list by the quantization sensitivity processor (for example Figure 3 the 14th convolutional layer and the 17th convolutional layer in Figure 3In the 1st and 4th convolutional layers (in [ID]), the 4-bit candidate layer will most likely be quantized with a 4-bit width, and the quantization width probabilities (i.e., quantization weights) of 8-bit and 16-bit will also decrease. These layers have the lowest QSE values compared to neighboring points and can be relatively safely quantized with low precision. The remaining points will be included in the candidate list of 8-bit quantization width probabilities (i.e., quantization weights) (e.g., Figure 3 in the 6th and 9th convolutional layers in [ID]) and participate in the subsequent search process.

[0048] In step S103, the model to be quantized is quantized based on the quantization width probability information of each convolutional layer in the at least one convolutional layer, to obtain a preset number of quantized candidate models.

[0049] In an exemplary embodiment of the present disclosure, the step of quantizing the model to be quantized based on the quantization width probability information of each convolutional layer in the at least one convolutional layer may include: by respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization width probability information of each convolutional layer in the at least one convolutional layer, determining the quantized candidate models of the model to be quantized; based on a preset quantization memory occupancy limit, determining the preset number of quantized candidate models from the quantized candidate models of the model to be quantized, wherein, between every two of the preset number of quantized candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization widths.

[0050] In an exemplary embodiment of the present disclosure, the step of determining the quantized candidate models of the model to be quantized by respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization width probability information of each convolutional layer in the at least one convolutional layer may include: for each convolutional layer in the at least one convolutional layer, sequentially determining the quantization width of each convolutional layer to be the candidate quantization widths in the order from high to low of the quantization width probabilities in the quantization width probability information of each convolutional layer, to quantize each convolutional layer and obtain the quantized candidate models. For example, the quantized candidate models may be determined in the order from high to low of the sum of the quantization width probabilities of the at least one convolutional layer (e.g., the sum of the quantization width probabilities of 3 convolutional layers).

[0051] In an exemplary embodiment of the present disclosure, the step of determining the preset number of quantization candidate models from the quantization candidate models of the to-be-quantized model based on a preset quantization memory occupancy limit may include: determining whether the quantization candidate models of the to-be-quantized model satisfy the preset quantization memory occupancy limit; based on determining that the quantization candidate models of the to-be-quantized model satisfy the preset quantization memory occupancy limit, taking the quantization candidate models that satisfy the preset quantization memory occupancy limit as the first quantization candidate models until the number of the first quantization candidate models reaches a first preset number; performing a repeatability determination on the first preset number of first quantization candidate models, and selecting non-repeating first quantization candidate models from the first preset number of first quantization candidate models as the second preset number of second quantization candidate models, where, between any two of the second preset number of second quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit widths; selecting the preset number of second quantization candidate models that satisfy the preset quantization memory occupancy limit from the second preset number of second quantization candidate models as the preset number of quantization candidate models, thereby not only greatly reducing the memory occupancy of the quantization candidate models, but also maintaining the final verification accuracy of the quantization candidate models.

[0052] As an example, different quantization bit width selection probabilities are assigned to these convolutional layers according to the quantization sensitivity QSE value of the convolutional layer, and model search is performed by a Quantized Model Searcher according to the different quantization bit width probability data of each convolutional layer.

[0053] In order to distinguish the structural differences of these search models, in an exemplary embodiment of the present disclosure, a quantization memory occupancy limit (Quantify Memory Access Cost Limit, abbreviated as B-MAC) is defined according to the memory occupancy of the quantization model. As an example, the formula is as follows: . Here, B-MAC represents the quantization memory occupancy limit, represents the memory occupancy, represents, represents, represents the first convolutional layer, represents all convolutional layers, represents from the first convolutional layer to the end of all convolutional layers.

[0054] Such as Figure 4As shown, the candidate results searched by the quantization model searcher will be judged by the quantization memory occupancy limit, and all unsatisfactory results (B-MAC overflow) will be marked and discarded. The quantization model searcher will operate in a loop strictly following the limit conditions until a sufficient number (Set Amount) of candidate results are searched. All candidate results will be sent to the Candidate Model Pool for duplicate judgment to prevent the same model from being selected repeatedly. Finally, for example, three results with the smallest B-MAC will be selected from these qualified candidate results and sent to the quantization module QM to verify the accuracy of the quantization model.

[0055] In step S104, a quantization model meeting the preset conditions is selected from the preset number of quantization candidate models. Here, the preset conditions can be that the accuracy meets the conditions. For example, the preset conditions can be that the accuracy is the highest among the accuracies of all quantization candidate models, or the accuracy exceeds the accuracy threshold. In addition, the preset conditions can also be that the memory occupancy meets the conditions. For example, the preset conditions can be that the memory occupancy is the smallest among the memory occupancies of all quantization candidate models, or the memory occupancy is lower than the occupancy threshold. In addition, the preset conditions can be that the accuracy meets the conditions and the memory occupancy also meets the conditions. For example, the preset conditions can be that the accuracy of the quantization candidate model exceeds the accuracy threshold and the memory occupancy is lower than the occupancy threshold.

[0056] In an exemplary embodiment of the present disclosure, the step of selecting a quantization model meeting the preset conditions from the preset number of quantization candidate models may include: performing accuracy verification on the preset number of quantization candidate models, and using the quantization candidate model with the highest accuracy among the preset number of quantization candidate models as the quantization model meeting the preset conditions, thereby improving the model quantization effect.

[0057] As an example, after obtaining the candidate results searched by the quantization model searcher, the QIL method will be used for model quantization in sequence. All convolutional layers (nn.Conv2d) in the input model to be quantized will be replaced by the quantization convolutional layer (QConv2d) by the Model Quantizer, and the input information (nBitAct) and weight information (nBitWei) of each convolutional layer will be quantized to the same bit width (i.e., nBitAct is equal to nBitWei). To avoid accuracy loss caused by quantization, the first convolutional layer (1st Conv2d Layer) will remain in full precision without any quantization, and all fully connected layers (Linear Layer) will also remain in full precision without any quantization. Finally, the quantization model will be sent to the Quantized Model Trainer for accuracy verification.

[0058] Benefiting from the effective guidance of the quantization sensitivity QSE information of the convolutional layer, the quantization model search based on the quantization sensitivity QSE of the convolutional layer greatly reduces the time-consuming of the quantization model search, and the adaptive quantization sensitivity QSE data processing method has good generalization. By using the model quantization method according to the exemplary embodiments of the present disclosure, the effect and efficiency of model quantization can be improved.

[0059] The above has been combined with Figures 1 to 4 to describe the model quantization method according to the exemplary embodiments of the present disclosure. In the following, reference will be made to Figure 5 to describe the model quantization device and its units according to the exemplary embodiments of the present disclosure.

[0060] Figure 5 The block diagram showing the model quantization device according to the exemplary embodiments of the present disclosure.

[0061] Referring to Figure 5 , the model quantization device includes a sensitivity determination unit 51, a probability information determination unit 52, a model quantization unit 53, and a precision selection unit 54. The sensitivity determination unit 51 is configured to determine the quantization sensitivity of at least one convolutional layer of the model to be quantized.

[0062] In the exemplary embodiments of the present disclosure, the sensitivity determination unit 51 may be configured to: parse the model to be quantized to obtain the convolutional layer information of each convolutional layer in the at least one convolutional layer, where the convolutional layer information includes the number of input channels, the number of output channels, and the sampling matrix; calculate the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized respectively based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters of each convolutional layer in the at least one convolutional layer.

[0063] In the exemplary embodiments of the present disclosure, the sensitivity determination unit 51 may be configured to: for each convolutional layer in the at least one convolutional layer, calculate the Hessian matrix of the pre-trained parameters of each convolutional layer; calculate the quantization sensitivity of each convolutional layer based on the Hessian matrix, the number of input channels, the number of output channels, and the sampling matrix.

[0064] In the exemplary embodiments of the present disclosure, the at least one convolutional layer does not include the first convolutional layer in the model to be quantized.

[0065] The probability information determination unit 52 is configured to respectively determine the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, where the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width among a plurality of candidate quantization bit-widths.

[0066] In an exemplary embodiment of the present disclosure, the probability information determination unit 52 may be configured to: for each convolutional layer in the at least one convolutional layer, determine the quantization sensitivity data distribution of each convolutional layer based on the quantization sensitivity of each convolutional layer; determine the quantization sensitivity envelope points from the quantization sensitivity data distribution of each convolutional layer; and determine the quantization bit-width probability information of each convolutional layer based on the quantization sensitivity envelope points of each convolutional layer.

[0067] The model quantization unit 53 is configured to quantize the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer, to obtain a preset number of quantized candidate models.

[0068] In an exemplary embodiment of the present disclosure, the model quantization unit 53 may be configured to: determine the quantized candidate models of the model to be quantized by respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer; and determine the preset number of quantized candidate models from the quantized candidate models of the model to be quantized based on a preset quantization memory occupancy limit, where between any two of the preset number of quantized candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit-widths.

[0069] In an exemplary embodiment of the present disclosure, the model quantization unit 53 may be configured to: for each convolutional layer in the at least one convolutional layer, sequentially determine the quantization bit-width of each convolutional layer as the candidate quantization bit-widths in the order from high to low of the quantization bit-width probabilities in the quantization bit-width probability information of each convolutional layer, to quantize each convolutional layer and obtain the quantized candidate models.

[0070] In an exemplary embodiment of the present disclosure, the model quantization unit 53 may be configured to: determine whether the quantization candidate model of the model to be quantized meets the preset quantization memory occupancy limit; based on determining that the quantization candidate model of the model to be quantized meets the preset quantization memory occupancy limit, use the quantization candidate model that meets the preset quantization memory occupancy limit as the first quantization candidate model until the number of first quantization candidate models reaches a first preset number; perform a repeatability determination on the first preset number of first quantization candidate models, and select non-repeating first quantization candidate models among the first preset number of first quantization candidate models as a second preset number of second quantization candidate models, wherein, between every two of the second preset number of second quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit widths; select the preset number of second quantization candidate models that meet the preset quantization memory occupancy limit from the second preset number of second quantization candidate models as the preset number of quantization candidate models.

[0071] The accuracy selection unit 54 is configured to select a quantization model that meets the preset conditions from the preset number of quantization candidate models.

[0072] In an exemplary embodiment of the present disclosure, the accuracy selection unit 54 may be configured to: perform accuracy verification on the preset number of quantization candidate models, and use the quantization candidate model with the highest accuracy among the preset number of quantization candidate models as the quantization model that meets the preset conditions.

[0073] In addition, according to an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the model quantization method according to the exemplary embodiment of the present disclosure is implemented.

[0074] In an exemplary embodiment of the present disclosure, the computer-readable storage medium may carry one or more programs, and when the computer program is executed, the following steps may be implemented: determine the quantization sensitivity of at least one convolutional layer of the model to be quantized; respectively determine the quantization bit width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit width among a plurality of candidate quantization bit widths; quantize the model to be quantized based on the quantization bit width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantization candidate models; select a quantization model that meets the preset conditions from the preset number of quantization candidate models.

[0075] A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a computer program, and the computer program may be used by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium may be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above. The computer-readable storage medium may be contained in any device; it may also exist alone without being assembled into the device.

[0076] In addition, according to an exemplary embodiment of the present disclosure, there is also provided a computer program product, and the instructions in the computer program product can be executed by a processor of a computer device to complete the method of model quantization according to the exemplary embodiment of the present disclosure.

[0077] The above has been combined with Figure 5 to describe the model quantization device according to the exemplary embodiment of the present disclosure. Next, in combination with Figure 6 to describe the computing device according to the exemplary embodiment of the present disclosure.

[0078] Figure 6 A schematic diagram showing a computing device according to an exemplary embodiment of the present disclosure.

[0079] Referring to Figure 6 , a computing device 6 according to an exemplary embodiment of the present disclosure includes a memory 61 and a processor 62. A computer program is stored on the memory 61. When the computer program is executed by the processor 62, the method of model quantization according to the exemplary embodiment of the present disclosure is implemented.

[0080] In an exemplary embodiment of the present disclosure, when the computer program is executed by the processor 62, the following steps may be implemented: determining the quantization sensitivity of at least one convolutional layer of the model to be quantized; respectively determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width from a plurality of candidate quantization bit-widths; quantizing the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantized candidate models; and selecting a quantized model that meets the preset conditions from the preset number of quantized candidate models.

[0081] The computing device in the embodiments of the present disclosure may include, but is not limited to, devices such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 6 The illustrated computing device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0082] The above has been described with reference to Figures 1 to 6 the model quantization method and device according to the exemplary embodiments of the present disclosure. However, it should be understood that: Figure 5 the model quantization device and its units shown in Figure 6 may be respectively configured as software, hardware, firmware, or any combination of the above items that perform specific functions,

[0083] The model quantization method and device according to the exemplary embodiments of the present disclosure, by determining the quantization sensitivity of at least one convolutional layer of the model to be quantized, respectively determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width from a plurality of candidate quantization bit-widths, quantizing the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantized candidate models, and selecting a quantized model that meets the preset conditions from the preset number of quantized candidate models, thereby improving the effect and efficiency of model quantization.

[0084] Although the present disclosure has been specifically shown and described with reference to its exemplary embodiments, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of the present disclosure as defined by the claims.

Claims

1. A model quantization method, comprising: Determining the quantization sensitivity of at least one convolutional layer of the model to be quantized; Respectively determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability that each convolutional layer in the at least one convolutional layer selects each candidate quantization bit-width from multiple candidate quantization bit-widths; Quantizing the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer to obtain a preset number of quantized candidate models; Selecting a quantized model that meets the preset conditions from the preset number of quantized candidate models.

2. The method according to claim 1, wherein, The step of determining the quantization sensitivity of at least one convolutional layer of the model to be quantized comprises: Parsing the model to be quantized to obtain the convolutional layer information of each convolutional layer in the at least one convolutional layer, wherein the convolutional layer information includes the number of input channels, the number of output channels, and the sampling matrix; Based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters of each convolutional layer in the at least one convolutional layer, respectively calculating the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized.

3. The method according to claim 2, wherein, The step of respectively calculating the quantization sensitivity of each convolutional layer in the at least one convolutional layer of the model to be quantized based on the convolutional layer information of each convolutional layer in the at least one convolutional layer and the pre-trained parameters of each convolutional layer in the at least one convolutional layer comprises: For each convolutional layer in the at least one convolutional layer, calculating the Hessian matrix of the pre-trained parameters of each convolutional layer; Based on the Hessian matrix, the number of input channels, the number of output channels, and the sampling matrix, calculating the quantization sensitivity of each convolutional layer.

4. The method according to claim 1, wherein, The step of respectively determining the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer comprises: For each convolutional layer in the at least one convolutional layer, determining the quantization sensitivity data distribution of each convolutional layer based on the quantization sensitivity of each convolutional layer; Determining the quantization sensitivity envelope points of each convolutional layer from the quantization sensitivity data distribution of each convolutional layer; Based on the quantization sensitivity envelope points of each convolutional layer, determining the quantization bit-width probability information of each convolutional layer.

5. The method according to claim 1, wherein, The step of quantizing the model to be quantized based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer comprises: Determining the quantized candidate models of the model to be quantized by respectively quantizing each convolutional layer in the at least one convolutional layer based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer; Determine the preset number of quantization candidate models from the quantization candidate models of the model to be quantized based on a preset quantization memory occupancy limit, wherein, between every two quantization candidate models among the preset number of quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit widths.

6. The method according to claim 5, wherein, The step of determining the quantization candidate models of the model to be quantized by quantizing each convolutional layer in the at least one convolutional layer respectively based on the quantization bit width probability information of each convolutional layer in the at least one convolutional layer includes: For each convolutional layer in the at least one convolutional layer, sequentially determine the quantization bit width of each convolutional layer as the candidate quantization bit widths in the order from high to low of the quantization bit width probabilities in the quantization bit width probability information of each convolutional layer, to quantize each convolutional layer, and obtain the quantization candidate models.

7. The method according to claim 5, wherein, The step of determining the preset number of quantization candidate models from the quantization candidate models of the model to be quantized based on a preset quantization memory occupancy limit includes: Determine whether the quantization candidate models of the model to be quantized meet the preset quantization memory occupancy limit; Based on determining that the quantization candidate models of the model to be quantized meet the preset quantization memory occupancy limit, use the quantization candidate models that meet the preset quantization memory occupancy limit as the first quantization candidate models until the number of the first quantization candidate models reaches a first preset number; Perform a repeatability determination on the first preset number of first quantization candidate models, and select the non-repeating first quantization candidate models among the first preset number of first quantization candidate models as the second preset number of second quantization candidate models, wherein, between every two second quantization candidate models among the second preset number of second quantization candidate models, one or more convolutional layers in the at least one convolutional layer have different candidate quantization bit widths; Select the preset number of second quantization candidate models that meet the preset quantization memory occupancy limit from the second preset number of second quantization candidate models as the preset number of quantization candidate models.

8. The method according to claim 1, wherein, The step of selecting a quantization model that meets a preset condition from the preset number of quantization candidate models includes: Perform accuracy verification on the preset number of quantization candidate models, and use the quantization candidate model with the highest accuracy among the preset number of quantization candidate models as the quantization model that meets the preset condition.

9. The method according to claim 1, wherein, The at least one convolutional layer does not include the first convolutional layer in the model to be quantized.

10. A model quantization device, comprising: A sensitivity determination unit configured to determine the quantization sensitivity of at least one convolutional layer of a model to be quantized; A probability information determination unit configured to respectively determine the quantization bit width probability information of each convolutional layer in the at least one convolutional layer according to the quantization sensitivity of each convolutional layer in the at least one convolutional layer, wherein the quantization bit width probability information of each convolutional layer in the at least one convolutional layer respectively represents the probability of each convolutional layer in the at least one convolutional layer selecting each candidate quantization bit width among a plurality of candidate quantization bit widths; A model quantization unit, configured to quantize the to-be-quantized model based on the quantization bit-width probability information of each convolutional layer in the at least one convolutional layer, to obtain a preset number of quantized candidate models; and An accuracy selection unit, configured to select a quantized model meeting a preset condition from the preset number of quantized candidate models.

11. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, the model quantization method according to any one of claims 1 to 9 is implemented.

12. A computing device, comprising: At least one processor; At least one memory storing a computer program, which when executed by the at least one processor, implements the model quantization method according to any one of claims 1 to 9.