Model Quantization Method, Apparatus, Storage Medium, and Electronic Device

By quantizing the deep learning model and selecting a reference quantized model with excellent performance, the storage and processor load problems caused by the huge number of model parameters are solved, and the efficient computing and real-time performance of the model are achieved.

CN116468082BActive Publication Date: 2025-07-08伟光有限公司(CN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310445134.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2025-07-08
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

The neural network structure of the deep learning model is complex, resulting in a huge number of parameters, a large memory space and a high processor requirement, which affects the real-time and performance of model operations.

Method used

By obtaining multiple candidate parameter values of the model to be quantized, using the preset parameter determination strategy for quantization, selecting a reference quantization model with excellent performance until the preset stop condition is met, and the final quantized model is determined.

Benefits of technology

It reduces the storage requirements and processor load of the model, improves the computing efficiency and performance of the model, and realizes the real-time and efficient operation of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468082B_ABST
    Figure CN116468082B_ABST
Patent Text Reader

Abstract

The present application discloses a model quantization method, apparatus, storage medium and electronic device. The method includes: performing quantization processing on the model to be quantized according to the first candidate parameter value of the quantization parameter of the model to be quantized, to obtain a first candidate quantized model; determining a second candidate parameter value of the quantization parameter of the model to be quantized according to the first candidate parameter value through a preset parameter determination strategy, performing quantization processing on the model to be quantized according to the second candidate parameter value, to obtain a second candidate quantized model; selecting a reference quantized model from the first candidate quantized model and the second candidate quantized model according to the performance values of the first candidate quantized model and the second candidate quantized model, determining the parameter value corresponding to the reference quantized model as the first candidate parameter value, taking the reference quantized model as the first candidate quantized model, and returning to execute the corresponding steps until a preset stop condition is satisfied, to obtain the quantized model corresponding to the model to be quantized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, and particularly relates to a model quantization method, apparatus, computer-readable storage medium, and electronic device. Background Art

[0002] With the development of deep learning technology, deep learning technology has achieved great success in many application fields. In deep learning technology, the quality of the neural network structure has a very important impact on the effect of the model. In practice, in order to obtain high performance, the structure complexity of the neural network is relatively high, and correspondingly, the number of network parameters is huge. Storing the parameters of the neural network model requires a large amount of memory space, and when running the neural network model, due to the large number of parameters and high precision, the requirements for the processor are relatively high.

[0003] In order to ensure the real-time performance of model operations, reduce the computing pressure on the processor, and at the same time ensure the performance of the model, it is necessary to quantize the model. Summary of the Invention

[0004] Embodiments of this application provide a model quantization method, apparatus, storage medium, and electronic device, which can obtain a quantized model corresponding to the model to be quantized.

[0005] In a first aspect, embodiments of this application provide a model quantization method, including:

[0006] Obtain multiple first candidate parameter values of the quantization parameters of the model to be quantized, and perform quantization processing on the model to be quantized according to each of the first candidate parameter values to obtain multiple first candidate quantized models;

[0007] According to the multiple first candidate parameter values, determine multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy, and perform quantization processing on the model to be quantized according to each of the second candidate parameter values to obtain multiple second candidate quantized models;

[0008] According to the performance value of each of the first candidate quantized models and the performance value of each of the second candidate quantized models, select multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models, determine the parameter value corresponding to each reference quantized model as the first candidate parameter value, use each reference quantized model as the first candidate quantized model, and return to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy according to the multiple first candidate parameter values until a preset stop condition is met;

[0009] Determine the quantized model corresponding to the model to be quantized from multiple reference quantized models that meet the preset stop condition.

[0010] In a second aspect, an embodiment of the present application provides a model quantization device, including:

[0011] A first processing module, configured to obtain multiple first candidate parameter values of quantization parameters of a model to be quantized, and perform quantization processing on the model to be quantized according to each of the first candidate parameter values to obtain multiple first candidate quantized models;

[0012] A second processing module, configured to determine multiple second candidate parameter values of the quantization parameters of the model to be quantized according to the multiple first candidate parameter values through a preset parameter determination strategy, and perform quantization processing on the model to be quantized according to each of the second candidate parameter values to obtain multiple second candidate quantized models;

[0013] A first determination module, configured to select multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models according to the performance value of each of the first candidate quantized models and the performance value of each of the second candidate quantized models, determine the parameter value corresponding to each reference quantized model as the first candidate parameter value, use each reference quantized model as the first candidate quantized model, and return to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized according to the multiple first candidate parameter values through a preset parameter determination strategy until the preset stop condition is met;

[0014] A second determination module, configured to determine the quantized model corresponding to the model to be quantized from multiple reference quantized models that meet the preset stop condition.

[0015] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is caused to execute the model quantization method provided by the embodiment of the present application.

[0016] In a fourth aspect, an embodiment of the present application further provides an electronic device, including a memory and a processor. The processor is configured to execute the model quantization method provided by the embodiment of the present application by calling the computer program stored in the memory.

[0017] In the embodiments of the present application, by obtaining multiple first candidate parameter values of the quantization parameters of the model to be quantized, and performing quantization processing on the model to be quantized according to each of the first candidate parameter values, multiple first candidate quantized models are obtained; according to the multiple first candidate parameter values, multiple second candidate parameter values of the quantization parameters of the model to be quantized are determined through a preset parameter determination strategy, and quantization processing is performed on the model to be quantized according to each of the second candidate parameter values, obtaining multiple second candidate quantized models; according to the performance value of each of the first candidate quantized models and the performance value of each of the second candidate quantized models, multiple reference quantized models are selected from the multiple first candidate quantized models and the multiple second candidate quantized models, the parameter value corresponding to each reference quantized model is determined as the first candidate parameter value, each reference quantized model is used as the first candidate quantized model, and the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy according to the multiple first candidate parameter values is returned and executed until a preset stop condition is met; determining the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition can achieve obtaining the quantized model corresponding to the model to be quantized. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of the present application, will make the technical solutions and their beneficial effects of the present application obvious.

[0019] Figure 1 is the first flowchart of the model quantization method provided by the embodiments of the present application.

[0020] Figure 2 is the second flowchart of the model quantization method provided by the embodiments of the present application.

[0021] Figure 3 is the structural schematic diagram of the model quantization device provided by the embodiments of the present application.

[0022] Figure 4 is the structural schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] It should be noted that the terms "first", "second", "third", etc. in this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but some embodiments also include steps or modules that are not listed, or some embodiments also include other steps or modules inherent to these processes, methods, products, or devices.

[0024] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0025] An embodiment of this application provides a model quantization method, a model quantization device, a storage medium, and an electronic device. Among them, the model quantization method can be applied to the quantization of neural network models of any field and any structure, such as images, sounds, natural languages, controls, etc. The execution subject of the model quantization method can be the model quantization device provided by the embodiment of this application, or an electronic device integrated with the model quantization device, where the model quantization device can be implemented in a hardware or software manner. Among them, the electronic device can be a device configured with a processor and having model quantization capabilities, such as a smart phone, a tablet computer, a palm computer, a notebook computer, etc.

[0026] Please refer to Figure 1 , Figure 1 which is the first flowchart of the model quantization method provided by the embodiment of this application. The process may include:

[0027] In 101, obtain multiple first candidate parameter values of the quantization parameters of the model to be quantized, and according to each first candidate parameter value, perform quantization processing on the model to be quantized to obtain multiple first candidate quantized models.

[0028] It can be understood that the main targets of quantization in the model are the weight tensors (weights) of the weight layer and the activation tensors (activations) in the inference process. Then, the quantization parameters in the model include quantization parameters for quantizing the weight tensors and quantization parameters for quantizing the activation tensors. In current deep learning-based neural networks, the weight layer mainly includes convolutional layers (Convolution Layer, Conv) and fully connected layers (Linear Layer) of various dimensions.

[0029] The quantization of a model is divided into two types: symmetric quantization and asymmetric quantization. Quantization parameters are the parameters used to quantize the model. In symmetric quantization, the quantization parameters only include mapping parameters. In asymmetric quantization, the quantization parameters include mapping parameters and offset zero-point parameters. Taking 8-bit symmetric quantization as an example, that is, the quantization mapping range is symmetric based on the value of 0. Therefore, the quantization parameter only has the mapping parameter S, and its offset zero point Z = 0, that is, fixed at the zero point. The quantization bandwidth is 8 bits, that is, the quantization mapping range is [-2 7 +1, 2 7 -1], that is, [-127, 127].

[0030] Whether it is symmetric quantization or asymmetric quantization, the quantization of the weight tensor includes hierarchical quantization or per-channel quantization. And whether it is hierarchical quantization or per-channel quantization, their processing processes are the same. Only the weight tensor of hierarchical quantization is the weight tensor of each weight layer, that is, the weight tensors of different weight layers are quantized according to different quantization parameters. And the weight tensor of per-channel quantization is the weight tensor of each channel in each weight layer, that is, the weight tensors of different channels are quantized according to different quantization parameters.

[0031] The following will take the symmetric quantization of the i-th weight layer in a model to be quantized as an example to illustrate the model quantization method provided by the embodiments of the present application. Wherein, i = 1,..., N, and N is the number of quantization layers.

[0032] Suppose X i represents the activation tensor of the i-th weight layer in the model to be quantized, and W i represents the weight tensor of the i-th weight layer in the model to be quantized. The reference parameter values for the quantization parameters used to quantize the activation tensor and the reference parameter values for the quantization parameters used to quantize the weight tensor can be determined respectively first.

[0033] Subsequently, according to the reference parameter values of the quantization parameters used to quantize the activation tensor, multiple first candidate parameter values for the quantization parameters used to quantize the activation tensor can be obtained. Wherein, the difference between each first candidate parameter value and the reference parameter value is less than a preset difference.

[0034] According to the reference parameter values of the weight parameters used to quantize the activation tensor, multiple first candidate parameter values for the quantization parameters used to quantize the weight tensor can be obtained. Wherein, the difference between each first candidate parameter value and the reference parameter value is less than a preset difference.

[0035] Among them, the preset difference can be set by those skilled in the art or set by a computer device based on certain rules.

[0036] It can be understood that through the same process as the above process, a plurality of first candidate parameter values for quantifying the quantization parameters of the activation tensors of other weight layers in the to-be-quantized model except the i-th layer weight layer and a plurality of first candidate parameter values for quantifying the quantization parameters of the weight tensors of other weight layers in the to-be-quantized model except the i-th layer weight layer can also be obtained.

[0037] For example, assume that the to-be-quantized model includes two layers of weight layers. Through the above process, a plurality of first candidate parameter values S111, S112, and S113 for quantifying the quantization parameters of the activation tensor of the first layer weight layer in the to-be-quantized model and a plurality of first candidate parameter values S121, S122, and S123 for quantifying the quantization parameters of the weight tensor of the first layer weight layer in the to-be-quantized model are obtained, as well as a plurality of first candidate parameter values S211, S212, and S213 for quantifying the quantization parameters of the activation tensor of the second layer weight layer in the to-be-quantized model and a plurality of first candidate parameter values S221, S222, and S223 for quantifying the quantization parameters of the weight tensor of the second layer weight layer in the to-be-quantized model. The first candidate parameter values S111, S121, S211, and S221 can be used as the first candidate parameter value set G11 for quantifying the to-be-quantized model, the first candidate parameter values S112, S122, S212, and S222 can be used as the first candidate parameter value set G12 for quantifying the to-be-quantized model, and the first candidate parameter values S113, S123, S213, and S223 can be used as the first candidate parameter value set G13 for quantifying the to-be-quantized model.

[0038] For example, the activation tensor of the first layer weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S111, the weight tensor of the first layer weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S121, the activation tensor of the second layer weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S211, and the weight tensor of the second layer weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S221 to obtain a second candidate quantized model.

[0039] The activation tensor of the first layer weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S112, the weight tensor of the first layer weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S122, the activation tensor of the second layer weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S212, and the weight tensor of the second layer weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S222 to obtain a first candidate quantized model.

[0040] By analogy, the quantization of the model to be quantized can also be performed according to the same process as the above process based on other obtained first candidate parameter values, and other first candidate quantized models can be obtained, so as to obtain multiple first candidate quantized models.

[0041] It should be noted that the model to be quantized in the embodiments of the present application can be a neural network model in any field such as images, sounds, natural languages, controls, etc., such as an image classification model, a speech recognition model, an object detection model, and so on.

[0042] In 102, according to multiple first candidate parameter values, multiple second candidate parameter values of the quantization parameters of the model to be quantized are determined through a preset parameter determination strategy, and the model to be quantized is quantized according to each second candidate parameter value to obtain multiple second candidate quantized models.

[0043] Suppose the preset parameter determination strategy includes randomly selecting a set from the first candidate parameter value sets G11, G12, and G13, and mutating each parameter value in the set with a preset probability to obtain multiple second candidate parameter value sets. Among them, each second candidate parameter value set includes second candidate parameter values of the quantization parameters of the activation tensors for quantizing the first layer weight layer and the second layer weight layer of the model to be quantized, respectively, and second candidate parameter values of the quantization parameters of the weight tensors for quantizing the first layer weight layer and the second layer weight layer of the model to be quantized, respectively.

[0044] Suppose two second candidate parameter value sets G21 and G22 are obtained. The second candidate parameter value set G21 includes the second candidate parameter value S114 of the quantization parameter of the activation tensor for quantizing the first layer weight layer of the model to be quantized, the second candidate parameter value S124 of the quantization parameter of the weight tensor for quantizing the first layer weight layer of the model to be quantized, the second candidate parameter value S214 of the quantization parameter of the activation tensor for quantizing the second layer weight layer of the model to be quantized, and the second candidate parameter value S224 of the quantization parameter of the weight tensor for quantizing the second layer weight layer of the model to be quantized. The second candidate parameter value set G22 includes the second candidate parameter value S115 of the quantization parameter of the activation tensor for quantizing the first layer weight layer of the model to be quantized, the second candidate parameter value S125 of the quantization parameter of the weight tensor for quantizing the first layer weight layer of the model to be quantized, the second candidate parameter value S215 of the quantization parameter of the activation tensor for quantizing the second layer weight layer of the model to be quantized, and the second candidate parameter value S225 of the quantization parameter of the weight tensor for quantizing the second layer weight layer of the model to be quantized.

[0045] It should be noted that the above, such as S111, S112, and S113, are just examples of the number of the first candidate parameter values and do not limit this application. In actual applications, the number of the first candidate parameter values can be far more than this, and it is specifically set according to actual needs.

[0046] It should also be noted that the number of the second candidate parameter value sets can be less than, greater than, or equal to the number of the first parameter value set. Optionally, the number of the second candidate parameter value sets can be half of the number of the first parameter value set.

[0047] For example, the excitation tensor of the first weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S114, the weight tensor of the first weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S124, the excitation tensor of the second weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S214, and the weight tensor of the second weight layer in the to-be-quantized model can be quantized according to the second candidate parameter value S224, to obtain a second candidate quantized model.

[0048] The excitation tensor of the first weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S115, the weight tensor of the first weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S125, the excitation tensor of the second weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S215, and the weight tensor of the second weight layer in the to-be-quantized model can be quantized according to the first candidate parameter value S225, to obtain a second candidate quantized model.

[0049] And so on. Similarly, according to the other obtained second candidate parameter values, the to-be-quantized model can be quantized in the same process as above to obtain other second candidate quantized models, so as to obtain multiple second candidate quantized models.

[0050] In an optional embodiment, the excitation tensor in a to-be-quantized model can be quantized through formula (1) to obtain the excitation tensor in the corresponding candidate quantized model, and the weight tensor in a to-be-quantized model can also be quantized through formula (2) to obtain the weight tensor in the corresponding candidate quantized model, so as to obtain a candidate quantized model, such as the first candidate quantized model or the second candidate quantized model.

[0051]

[0052]

[0053] Q A (·),QW (·) respectively represent X i and W i 's quantization and dequantization processes. Clip(·) represents a truncation operation, represents rounding down, represents the activation tensor of the i-th weight layer of the candidate quantized model, represents the weight tensor of the i-th weight layer of the candidate quantized model, X i represents the activation tensor of the i-th weight layer of the model to be quantized, W i represents the weight tensor of the i-th weight layer of the model to be quantized, represents the candidate parameter values for quantizing the activation tensor of the i-th weight layer of the model to be quantized, including the first candidate parameter value or the second candidate parameter value, represents the candidate parameter values for quantizing the weight tensor of the i-th weight layer of the model to be quantized, including the first candidate parameter value or the second candidate parameter value. n represents the quantization bandwidth corresponding to the model to be quantized.

[0054] In an optional embodiment, if asymmetric quantization is adopted, the offset zero point parameter can be obtained through the same process as obtaining the quantization parameter, i.e., the mapping parameter, for quantizing the weight tensor or the activation tensor. Then, the activation tensor in a certain model to be quantized can be quantized through formula (3) to obtain the corresponding activation tensor in the candidate quantized model. The weight tensor in a certain model to be quantized can also be quantized through formula (4) to obtain the corresponding weight tensor in the candidate quantized model, thereby obtaining the candidate quantized model, such as the first candidate quantized model or the second candidate quantized model.

[0055]

[0056]

[0057] Q A (·), Q W (·) respectively represent X i and W i 's quantization and dequantization processes. Clip(·) represents a truncation operation, represents rounding down, represents the activation tensor of the i-th weight layer of the candidate quantized model, represents the weight tensor of the i-th weight layer of the candidate quantized model, X i represents the activation tensor of the i-th weight layer of the model to be quantized, W i represents the weight tensor of the i-th weight layer of the model to be quantized, Represents candidate parameter values for quantifying the activation tensor of the i-th weight layer of the model to be quantified, including a first candidate parameter value or a second candidate parameter value, Represents candidate parameter values for quantifying the weight tensor of the i-th weight layer of the model to be quantified, including a first candidate parameter value or a second candidate parameter value, where n represents the quantization bandwidth corresponding to the model to be quantified, Represents the offset zero point characterizing the offset of the activation tensor of the i-th weight layer of the model to be quantified, Represents the offset zero point characterizing the offset of the weight tensor of the i-th weight layer of the model to be quantified.

[0058] It should be noted that after the model to be quantified is trained, the activation tensor and weight tensor of each weight layer of the model to be quantified are known, and the value ranges of the activation tensor and weight tensor of the model to be quantified are also known.

[0059] In 103, according to the performance values of each first candidate quantized model and the performance values of each second candidate quantized model, multiple reference quantized models are selected from multiple first candidate quantized models and multiple second candidate quantized models, the parameter values corresponding to each reference quantized model are determined as the first candidate parameter values, each reference quantized model is used as the first candidate quantized model, and the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantified according to multiple first candidate parameter values through a preset parameter determination strategy is returned until a preset stop condition is met.

[0060] After obtaining multiple candidate quantized models, the performance value of each candidate quantized model can be determined respectively.

[0061] Among them, the performance value of the model is used to characterize the performance of the model. The better the performance of the model, the higher the performance value, and the performance of the model is characterized by the accuracy of the model prediction. The higher the accuracy, the better the performance.

[0062] For example, as shown in formula (5), let perfs = {} represent the set of performance records of candidate quantized models, such as the first candidate quantized model or the second candidate quantized model, infer(·) represents the test process, and based on the test set D val Measure the performance of each candidate quantized model p i Performance performance.

[0063] perfs[i] = infer(p i , D val ) (5)

[0065] For example, taking the candidate quantized model as an image classification model as an example, multiple test images can be obtained and input into a certain candidate quantized model to obtain the classification results of each test image; determine the proportion of the correct classification results among all the classification results to obtain the accuracy of the candidate quantized model; the accuracy can be used as the performance value of the candidate quantized model.

[0066] For example, candidate quantized models with performance values greater than the first performance value among multiple first candidate quantized models and multiple second candidate quantized models can be determined as reference quantized models. It is also possible to sort the multiple first candidate quantized models and multiple second candidate quantized models in descending order of performance values, and determine the candidate quantized models with the top preset number as reference quantized models.

[0067] Among them, the first preset performance value and the preset number can be set by those skilled in the art, or can be set by the electronic device based on certain rules.

[0068] For example, assuming that the number of the first candidate parameter value sets is twice the number of the second candidate parameter value sets, the preset number can be the number of the first candidate parameter value sets. Assuming that the number of the first candidate parameter value sets is 20, then the preset number is 20. Then, after sorting the multiple first candidate quantized models and multiple second candidate quantized models in descending order of performance values, the top 20 candidate quantized models can be determined as reference quantized models.

[0069] In an optional embodiment, for example, as shown in formulas (6) and (7), taking the accuracy (Accuracy, Acc) of the classification task as an example, sort in descending order according to the Acc value, eliminate the inferior candidate quantized models, such as the first candidate quantized model or the second candidate quantized model, and maintain the number of candidate quantized models as K. The calculation is as follows.

[0070]

[0071]

[0072] argsort(·) represents sorting in descending order and obtaining its index, topk(·) represents taking the top K coefficients, and it is i remain , that is, the K candidate quantized models to be retained. K can be the number of the first candidate parameter value sets.

[0073] After selecting the reference quantized model, the parameter values corresponding to each weight layer of each reference quantized model can be determined as the first candidate parameter values, each reference quantized model can be used as the first candidate quantized model, and the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized according to multiple first candidate parameter values through a preset parameter determination strategy can be returned until a preset stop condition is met.

[0074] For example, assume that a certain reference quantized model quantizes the activation tensor of the first weight layer in the model to be quantized according to the first candidate parameter value S111, quantizes the weight tensor of the first weight layer in the model to be quantized according to the first candidate parameter value S121, quantizes the activation tensor of the second weight layer in the model to be quantized according to the first candidate parameter value S211, and quantizes the weight tensor of the second weight layer in the model to be quantized according to the first candidate parameter value S221. Then the parameter values corresponding to this reference quantized model are the first candidate parameter values S111, S121, S211, and S221.

[0075] Another example, assume that a certain reference quantized model quantizes the activation tensor of the first weight layer in the model to be quantized according to the second candidate parameter value S114, quantizes the weight tensor of the first weight layer in the model to be quantized according to the second candidate parameter value S124, quantizes the activation tensor of the second weight layer in the model to be quantized according to the second candidate parameter value S214, and quantizes the weight tensor of the second weight layer in the model to be quantized according to the second candidate parameter value S224. Then the parameter values corresponding to this reference quantized model are the second candidate parameter values S114, S124, S214, and S224.

[0076] Among them, the preset stop condition can include that the iteration period reaches the preset period, or there is a reference quantized model with a performance value greater than the second preset performance value among the obtained reference quantized models. The second preset performance value can be set by those skilled in the art or set by the electronic device based on certain rules. The second preset performance value is greater than the first preset performance value.

[0077] In 104, determine the quantized model corresponding to the model to be quantized from multiple reference quantized models that meet the preset stop condition.

[0078] For example, when the preset stop condition is that the iteration period reaches the preset period, any reference quantized model with a performance value greater than the third preset performance value among the multiple reference quantized models that meet the preset stop condition can be determined as the quantized model corresponding to the model to be quantized, or the reference quantized model with the largest performance value among the reference quantized models can be determined as the quantized model corresponding to the model to be quantized.

[0079] For another example, when the preset stop condition is that there is a reference quantized model with a performance value greater than the second preset performance value in the obtained reference quantized models, any reference quantized model with a performance value greater than the second preset performance value in the reference quantized models can be determined as the quantized model corresponding to the model to be quantized, or the reference quantized model with the largest performance value in the reference quantized models can be determined as the quantized model corresponding to the model to be quantized.

[0080] Among them, the third preset performance value can be set by those skilled in the art or set by the electronic device based on certain rules. The third preset performance value is greater than the first preset performance value.

[0081] In this embodiment, by obtaining multiple first candidate parameter values of the quantization parameters of the model to be quantized, and according to each first candidate parameter value, performing quantization processing on the model to be quantized to obtain multiple first candidate quantized models; according to the multiple first candidate parameter values, through a preset parameter determination strategy, determining multiple second candidate parameter values of the quantization parameters of the model to be quantized, and according to each second candidate parameter value, performing quantization processing on the model to be quantized to obtain multiple second candidate quantized models; according to the performance value of each first candidate quantized model and the performance value of each second candidate quantized model, selecting multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models, determining the parameter value corresponding to each reference quantized model as the first candidate parameter value, taking each reference quantized model as the first candidate quantized model, and returning to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through the preset parameter determination strategy according to the multiple first candidate parameter values until the preset stop condition is met; determining the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition can implement obtaining the quantized model corresponding to the model to be quantized.

[0082] It can be understood that the model quantization method provided in this embodiment can also implement global tuning of the global quantization parameters of the model to be quantized.

[0083] In an optional embodiment, determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy according to the multiple first candidate parameter values includes:

[0084] Selecting multiple third candidate parameter values from the multiple first candidate parameter values;

[0085] Performing mutation processing on each third candidate parameter value according to a preset mutation probability to obtain multiple second candidate parameter values of the quantization parameters of the model to be quantized.

[0086] Among them, the preset mutation probability can be set by those skilled in the art or by the electronic device based on certain rules. For example, the preset mutation probability can be 0.01.

[0087] For example, assume that the model to be quantized includes two weight layers. Multiple first candidate parameter values S111, S112, and S113 for the quantization parameters of the activation tensor for quantizing the first weight layer in the model to be quantized and multiple first candidate parameter values S121, S122, and S123 for the quantization parameters of the weight tensor for quantizing the first weight layer in the model to be quantized, as well as multiple first candidate parameter values S211, S212, and S213 for the quantization parameters of the activation tensor for quantizing the second weight layer in the model to be quantized and multiple first candidate parameter values S221, S222, and S223 for the quantization parameters of the weight tensor for quantizing the second weight layer in the model to be quantized are obtained. The first candidate parameter values S111, S121, S211, and S221 can be used as the first candidate parameter value set G11 for quantizing the model to be quantized, the first candidate parameter values S112, S122, S212, and S222 can be used as the first candidate parameter value set G12 for quantizing the model to be quantized, and the first candidate parameter values S113, S123, S213, and S223 can be used as the first candidate parameter value set G13 for quantizing the model to be quantized. One first parameter value set can be randomly selected from the first candidate parameter value sets G11, G12, and G13. Assume it is the first candidate parameter value set G11. Thus, the third candidate parameter values S111 and S211 can be obtained. For the third candidate parameter values S111 and S211, they can be mutated respectively according to the probability of 0.01 to obtain the second candidate parameter value S123 for the quantization parameters of the weight tensor and the second candidate parameter value S223 for the quantization parameters of the weight tensor. One first parameter value set can also be randomly selected from the first parameter value sets G1, G2, and G3. Assume it is the first parameter value set G2. Thus, the third candidate parameter values S111, S121, S211, and S221 can be obtained. For the third candidate parameter values S111, S121, S211, and S221, they can be mutated respectively according to the probability of 0.01 to obtain the second candidate parameter value S114 for the quantization parameters of the activation tensor for the first weight layer in the model to be quantized, the second candidate parameter value S124 for the quantization parameters of the weight tensor for the first weight layer in the model to be quantized, the second candidate parameter value S214 for the quantization parameters of the activation tensor for the second weight layer in the model to be quantized, and the second candidate parameter value S224 for the quantization parameters of the weight tensor for the second weight layer in the model to be quantized.

[0088] And so on. Other second candidate parameter values can be obtained according to the same process as the above process, which will not be elaborated here.

[0089] In an alternative embodiment, each second candidate parameter value follows a Gaussian distribution with each third candidate parameter value as the mean and a value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum and minimum values of the parameter to be quantized required by the quantization parameter as the variance, that is, S’←N(S,σ 2 ). For example, the second candidate parameter value S114 follows a Gaussian distribution with the third candidate parameter value 111 as the mean and a value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum and minimum values of the parameter to be quantized required by the quantization parameter as the variance, and the second candidate parameter value S124 follows a Gaussian distribution with the third candidate parameter value 121 as the mean and a value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum and minimum values of the parameter to be quantized required by the quantization parameter as the variance. Wherein, S’ represents the second candidate parameter value, S represents the third candidate parameter value, and σ 2 represents the variance.

[0090] Wherein, the variance can be obtained through formula (8).

[0091]

[0092] Wherein, σ 2 represents the variance, max(F) represents the maximum value of the weight tensor or activation tensor, min(F) represents the minimum value of the weight tensor or activation tensor, and bw represents the quantization bandwidth corresponding to the model to be quantized.

[0093] It should be noted that the model to be quantized is a trained model, and after training, the value ranges of the weight tensor and activation tensor of the model to be quantized are both determined. Therefore, the maximum and minimum values of the weight tensor can be determined based on the value range of the weight tensor, and the maximum and minimum values of the activation tensor can be determined based on the value range of the activation tensor.

[0094] In an alternative embodiment, obtaining a plurality of first candidate parameter values of the quantization parameter of the model to be quantized includes:

[0095] Obtaining the parameter value range of the quantization parameter of the model to be quantized;

[0096] Obtaining a plurality of first candidate parameter values of the quantization parameter from the parameter value range.

[0097] For example, taking the quantization parameter including the weight parameter for quantizing the weight tensor of a certain weight layer of the model to be quantized as an example, the parameter value range of the weight parameter can be obtained, and then new values are randomly generated according to a uniform distribution, so as to obtain a plurality of first candidate parameter values for quantizing the weight tensor of a certain weight layer of the model to be quantized.

[0098] In an optionally implemented embodiment, obtaining the value range of the quantization parameters of the model to be quantized includes:

[0099] Determining the reference parameter value of the quantization parameter of the model to be quantized according to the maximum value of the parameter to be quantized that needs to be quantized by the quantization parameter and the quantization bandwidth corresponding to the model to be quantized;

[0100] Determining the value range of the quantization parameter of the model to be quantized according to the reference parameter value.

[0101] Among them, the reference parameter value of the quantization parameter for quantizing the weight tensor of the i-th weight layer of the model to be quantized can be obtained according to formula (9).

[0102]

[0103] Among them, represents the reference parameter value of the quantization parameter for quantizing the weight tensor of the i-th weight layer of the model to be quantized, max(W i ) represents the maximum value of the weight tensor of the i-th weight layer of the model to be quantized, and bw represents the quantization bandwidth corresponding to the model to be quantized.

[0104] Among them, the reference parameter value of the quantization parameter for quantizing the activation tensor of the i-th weight layer of the model to be quantized can be obtained according to formula (10).

[0105]

[0106] Among them, represents the reference parameter value of the quantization parameter for quantizing the activation tensor of the i-th weight layer of the model to be quantized, max(X i ) represents the maximum value of the activation tensor of the i-th weight layer of the model to be quantized, and bw represents the quantization bandwidth corresponding to the model to be quantized.

[0107] After obtaining the reference parameter value of the quantization parameter for quantizing the weight tensor of the i-th weight layer of the model to be quantized, the value range can be obtained based on this reference parameter value

[0108] After obtaining the reference parameter value of the quantization parameter for quantizing the activation tensor of the i-th weight layer of the model to be quantized, the value range can be obtained based on this reference parameter value

[0109] Among them, the values of α and β can be set by those skilled in the art, or can be set by the electronic device based on certain rules. For example, α = 0.1 and β = 10.

[0110] In an alternative embodiment, to improve the quantization accuracy, we can perform per-channel quantization on the weight tensor, that is, there is a quantization scale for each channel (usually the last dimension) of the weight tensor. That is, the reference parameter value for the quantization parameter of the weight tensor of the i-th weight layer of the model to be quantized can be obtained according to formula (11).

[0111]

[0112] Wherein, represents the reference parameter value for the quantization parameter of the weight tensor of the i-th weight layer of the model to be quantized, which represents a vector (i.e., ), c represents the value of the last dimension of W i , that is, the number of channels, max(W i ) represents the maximum value of the weight tensor of the i-th weight layer of the model to be quantized, bw represents the quantization bandwidth corresponding to the model to be quantized, and axis=-1 indicates the dimension of the channel.

[0113] It should be noted that the model to be quantized is a trained model. After training, the value ranges of the weight tensor and the activation tensor of each layer of the model to be quantized are both determined. Therefore, the maximum and minimum values of the weight tensor of each layer of the model to be quantized can be determined based on the value range of the weight tensor of each layer of the model to be quantized, and the maximum and minimum values of the activation tensor of each layer of the model to be quantized can be determined based on the value range of the activation tensor of each layer of the model to be quantized.

[0114] In an alternative embodiment, determining the quantized model corresponding to the model to be quantized from multiple reference quantized models that meet the preset stop condition includes:

[0115] Determining the reference quantized model with the maximum performance value from multiple reference quantized models that meet the preset stop condition to obtain the target quantized model;

[0116] Determining the quantized model corresponding to the model to be quantized according to the target quantized model.

[0117] For example, the reference quantized model with the maximum performance value can be determined from multiple reference quantized models that meet the preset stop condition to obtain the target quantized model, and the target quantized model can be determined as the quantized model corresponding to the model to be quantized.

[0118] In an alternative embodiment, determining the quantized model corresponding to the model to be quantized according to the target quantized model includes:

[0119] Determining the target parameter value range of the target quantization parameter according to the parameter value of the target quantization parameter corresponding to the target quantized model;

[0120] Retrieve the local optimal solution of the target quantization parameter from the target parameter value interval to obtain the target parameter value of the target quantization parameter;

[0121] Quantize the model to be quantized according to the target parameter value to obtain the target quantized model corresponding to the model to be quantized.

[0122] For example, for the parameter value of the target quantization parameter of the i-th weight layer of the target quantized model (i = 1, …, N, N is the number of quantization layers), perform local fine-tuning using the Alternating Direction Method of Multipliers (ADMM) algorithm, where represents the parameter value of the target quantization parameter corresponding to the activation tensor of the i-th weight layer of the target quantized model, represents the parameter value of the target quantization parameter corresponding to the weight tensor of the i-th weight layer of the target quantized model. The basic idea is to fix one of the parameter values and find the optimal solution for the other parameter, and iterate sequentially until both parameters converge to the optimal solution; that is, fix Take as the variable, and retrieve the local optimal solution around it (such as α = 0.5, β = 1.5) and update Subsequently, fix Search within for the local optimal solution of and update Repeat this process until convergence. Specifically, it is expressed as follows:

[0123] Use k = 0, …, K - 1 to represent the search step, and the array D represents the metric recording each step

[0124] Search for Fix D = []

[0125]

[0126] (Quantization: Quantized activation tensor and weight tensor)

[0127] (Dequantization and convolution process: Performance value of the quantized model)

[0128] (Calculate Under the condition of, the quantization error, such as the output result of the quantized model and the output result O of the model to be quantized i of the error)

[0129] k best = argmin(D), (Select the k corresponding to the minimum error and update it.)

[0130] Search Fix Perform similar operations as above. In the above calculation process, the metric of L(·) can be mean squared error, cosine similarity, etc.

[0131] Through the above process, the local optimal solution of the target quantization parameter can be obtained, that is, the target parameter value of the target quantization parameter; the quantization process is performed on the model to be quantized according to the target parameter value, and the target quantized model corresponding to the model to be quantized can further slightly improve the performance.

[0132] It should be noted that for the specific method of quantizing the model to be quantized according to the target parameter value and obtaining the target quantized model corresponding to the model to be quantized, reference can be made to the previous embodiments, which will not be elaborated here.

[0133] It should also be noted that considering time, the original ADMM iteration requires multiple cycles to converge. To save time, only one cycle of iteration can be performed on it (that is, update only once respectively).

[0134] It should also be noted that the above local fine-tuning process can also be applied to the quantization process of the model to be quantized according to the second candidate parameter value. After obtaining the candidate quantized model, local optimization is performed on the parameter values of the quantization parameters in the candidate quantized model, and then the optimized candidate quantized model is obtained, which is then applied to the subsequent process of selecting multiple reference quantized models.

[0135] In an optional embodiment, for the per-channel quantization type of weight quantization, Mutation processing can be performed on each element in each first candidate parameter vector, and finally multiple second candidate parameter vectors of the quantization parameters of the model to be quantized are obtained.

[0136] Please refer to Figure 2 , Figure 2 which is the second flow diagram of the model quantization method provided by the embodiments of the present application. The flow may include:

[0137] In 201, according to the maximum value of the quantization parameters to be quantized by the quantization parameters of the model to be quantized and the quantization bandwidth corresponding to the model to be quantized, the reference parameter value of the quantization parameters of the model to be quantized is determined.

[0138] In 202, according to the reference parameter value, the parameter value range of the quantization parameters of the model to be quantized is determined.

[0139] In 203, a plurality of first candidate parameter values of the quantization parameter are obtained from the parameter value range.

[0140] In 204, according to each first candidate parameter value, the quantization processing is performed on the model to be quantized, and a plurality of first candidate quantized models are obtained.

[0141] In 205, a plurality of third candidate parameter values are selected from the plurality of first candidate parameter values.

[0142] In 206, each third candidate parameter value is mutated according to a preset mutation probability, and a plurality of second candidate parameter values of the quantization parameter of the model to be quantized are obtained. Each second candidate parameter value follows a Gaussian distribution with each third candidate parameter value as the mean and the value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum and minimum values of the quantization parameter to be quantized as the variance.

[0143] In 207, according to each second candidate parameter value, the quantization processing is performed on the model to be quantized, and a plurality of second candidate quantized models are obtained.

[0144] In 208, according to the performance value of each first candidate quantized model and the performance value of each second candidate quantized model, a plurality of reference quantized models are selected from the plurality of first candidate quantized models and the plurality of second candidate quantized models. The parameter value corresponding to each reference quantized model is determined as the first candidate parameter value, and each reference quantized model is used as the first candidate quantized model, and the step of determining a plurality of second candidate parameter values of the quantization parameter of the model to be quantized according to a plurality of first candidate parameter values through a preset parameter determination strategy is returned until a preset stop condition is satisfied.

[0145] In 209, the reference quantized model with the largest performance value is determined from the plurality of reference quantized models that satisfy the preset stop condition, and the target quantized model is obtained.

[0146] In 210, according to the target quantized model, the quantized model corresponding to the model to be quantized is determined.

[0147] It should be noted that for the specific implementation of steps 201 to 210, reference may be made to the previous embodiments, which will not be elaborated here.

[0148] Please refer to Figure 3 , Figure 3 which is the structural schematic diagram of the model quantization device provided by the embodiment of the present application. The model quantization device 300 includes: a first processing module 301, a second processing module 302, a first determination module 303, and a second determination module 304.

[0149] The first processing module 301 is configured to obtain multiple first candidate parameter values of the quantization parameters of the model to be quantized, and perform quantization processing on the model to be quantized according to each of the first candidate parameter values, so as to obtain multiple first candidate quantized models.

[0150] The second processing module 302 is configured to determine multiple second candidate parameter values of the quantization parameters of the model to be quantized according to the multiple first candidate parameter values through a preset parameter determination strategy, and perform quantization processing on the model to be quantized according to each of the second candidate parameter values, so as to obtain multiple second candidate quantized models.

[0151] The first determination module 303 is configured to select multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models according to the performance values of each of the first candidate quantized models and the performance values of each of the second candidate quantized models, determine the parameter value corresponding to each reference quantized model as the first candidate parameter value, use each reference quantized model as the first candidate quantized model, and return to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy according to the multiple first candidate parameter values until a preset stop condition is met.

[0152] The second determination module 304 is configured to determine the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition.

[0153] In an optional embodiment, the second processing module 302 may be configured to: select multiple third candidate parameter values from the multiple first candidate parameter values; perform mutation processing on each of the third candidate parameter values according to a preset mutation probability to obtain multiple second candidate parameter values of the quantization parameters of the model to be quantized.

[0154] In an optional embodiment, each of the second candidate parameter values follows a Gaussian distribution with each of the third candidate parameter values as the mean and a value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum value and the minimum value of the quantization parameters to be quantized as the variance.

[0155] In an optional embodiment, the first processing module 301 may be configured to: obtain the parameter value range of the quantization parameters of the model to be quantized; obtain multiple first candidate parameter values of the quantization parameters from the parameter value range.

[0156] In an optional embodiment, the first processing module 301 may be configured to: determine a reference parameter value of the quantization parameter of the model to be quantized according to the maximum value of the parameter to be quantized required by the quantization parameter and the quantization bandwidth corresponding to the model to be quantized; and determine a parameter value range of the quantization parameter of the model to be quantized according to the reference parameter value.

[0157] In an optional embodiment, the second determination module 304 may be configured to: determine a reference quantized model with the maximum performance value from multiple reference quantized models that meet a preset stop condition, to obtain a target quantized model; and determine a quantized model corresponding to the model to be quantized according to the target quantized model.

[0158] In an optional embodiment, the second determination module 304 may be configured to: determine a target parameter value range of the target quantization parameter according to the parameter value of the target quantization parameter corresponding to the target quantized model; retrieve a local optimal solution of the target quantization parameter from the target parameter value range, to obtain a target parameter value of the target quantization parameter; and perform quantization processing on the model to be quantized according to the target parameter value, to obtain a target quantized model corresponding to the model to be quantized.

[0159] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is enabled to execute the model quantization method provided in this embodiment.

[0160] The embodiment of the present application further provides an electronic device, including a memory and a processor. The processor is configured to execute the model quantization method provided in this embodiment by calling the computer program stored in the memory.

[0161] For example, the above-mentioned electronic device may be a mobile terminal such as a tablet computer or a smart phone. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the electronic device provided in the embodiment of the present application.

[0162] The electronic device 400 may include components such as a processor 401 and a memory 402. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in

[0163] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing the application programs stored in the memory 402 and calling the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0164] The memory 402 can be used to store application programs and data. The application programs stored in the memory 402 contain executable codes. The application programs can form various functional modules. The processor 401 executes various functional applications and data processing by running the application programs stored in the memory 402.

[0165] In this embodiment, the processor 401 in the electronic device will, according to the following instructions, load the executable codes corresponding to the processes of one or more application programs into the memory 402, and the processor 401 will run the application programs stored in the memory 402, so as to achieve:

[0166] Obtain multiple first candidate parameter values of the quantization parameters of the model to be quantized, and according to each of the first candidate parameter values, perform quantization processing on the model to be quantized to obtain multiple first candidate quantized models;

[0167] According to the multiple first candidate parameter values, through a preset parameter determination strategy, determine multiple second candidate parameter values of the quantization parameters of the model to be quantized, and according to each of the second candidate parameter values, perform quantization processing on the model to be quantized to obtain multiple second candidate quantized models;

[0168] According to the performance value of each of the first candidate quantized models and the performance value of each of the second candidate quantized models, select multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models, determine the parameter value corresponding to each reference quantized model as the first candidate parameter value, use each reference quantized model as the first candidate quantized model, and return to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized according to the multiple first candidate parameter values until a preset stop condition is met;

[0169] Determine the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition.

[0170] In an optional embodiment, when the processor 401 determines multiple second candidate parameter values of the quantization parameters of the to-be-quantized model according to the multiple first candidate parameter values through a preset parameter determination strategy, it may execute: selecting multiple third candidate parameter values from the multiple first candidate parameter values; performing mutation processing on each of the third candidate parameter values according to a preset mutation probability to obtain multiple second candidate parameter values of the quantization parameters of the to-be-quantized model.

[0171] In an optional embodiment, each of the second candidate parameter values follows a Gaussian distribution with each of the third candidate parameter values as the mean and a value determined according to the quantization bandwidth corresponding to the to-be-quantized model, the maximum and minimum values of the to-be-quantized parameters to be quantized by the quantization parameters as the variance.

[0172] In an optional embodiment, when the processor 401 obtains multiple first candidate parameter values of the quantization parameters of the to-be-quantized model, it may execute: obtaining a parameter value range of the quantization parameters of the to-be-quantized model; obtaining multiple first candidate parameter values of the quantization parameters from the parameter value range.

[0173] In an optional embodiment, when the processor 401 obtains the parameter value range of the quantization parameters of the to-be-quantized model, it may execute: determining a reference parameter value of the quantization parameters of the to-be-quantized model according to the maximum value of the to-be-quantized parameters to be quantized by the quantization parameters and the quantization bandwidth corresponding to the to-be-quantized model; determining a parameter value range of the quantization parameters of the to-be-quantized model according to the reference parameter value.

[0174] In an optional embodiment, when the processor 401 determines the quantized model corresponding to the to-be-quantized model from multiple reference quantized models that meet a preset stop condition, it may execute: determining the reference quantized model with the maximum performance value from multiple reference quantized models that meet a preset stop condition to obtain a target quantized model; determining the quantized model corresponding to the to-be-quantized model according to the target quantized model.

[0175] In an optional embodiment, when the processor 401 determines the quantized model corresponding to the to-be-quantized model according to the target quantized model, it may execute: determining a target parameter value range of the target quantization parameters according to the parameter values of the target quantization parameters corresponding to the target quantized model; retrieving a local optimal solution of the target quantization parameters from the target parameter value range to obtain a target parameter value of the target quantization parameters; performing quantization processing on the to-be-quantized model according to the target parameter value to obtain the target quantized model corresponding to the to-be-quantized model.

[0176] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not elaborated in a certain embodiment, reference may be made to the detailed description of the model quantization method above, and details will not be repeated here.

[0177] The model quantization device provided in the embodiments of the present application and the model quantization method in the above embodiments belong to the same concept. Any method provided in the embodiments of the model quantization method can be run on the model quantization device. The specific implementation process can be seen in the embodiments of the model quantization method, and details will not be repeated here.

[0178] It should be noted that for the model quantization method of the embodiments of the present application, those of ordinary skill in the art can understand that all or part of the process of implementing the model quantization method of the embodiments of the present application can be completed by controlling relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, such as stored in a memory and executed by at least one processor. During the execution process, it can include the process of the embodiments of the model quantization method. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), etc.

[0179] It can be understood that in the specific implementation of the present application, user information is involved, such as data related to application usage behavior data, logs, etc. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0180] For the model quantization device of the embodiments of the present application, its various functional modules can be integrated in a processing chip, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. When the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk.

[0181] The above has introduced in detail a model quantization method, device, storage medium, and electronic device provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A model quantization method, characterized in that, Including: Obtaining multiple first candidate parameter values of quantization parameters of a model to be quantized, and performing quantization processing on the model to be quantized according to each of the first candidate parameter values to obtain multiple first candidate quantized models; According to the multiple first candidate parameter values, determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy, and performing quantization processing on the model to be quantized according to each of the second candidate parameter values to obtain multiple second candidate quantized models; According to the performance value of each of the first candidate quantized models and the performance value of each of the second candidate quantized models, selecting multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models, determining the parameter value corresponding to each reference quantized model as the first candidate parameter value, taking each reference quantized model as the first candidate quantized model, and returning to execute the step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized through a preset parameter determination strategy until a preset stop condition is satisfied; Determining the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition; When the candidate quantized model is an image classification model, obtaining multiple images to be tested and inputting them into the candidate quantized model to obtain the classification result of each image to be tested; Determining the proportion of the correct classification result in all classification results to obtain the accuracy of the candidate quantized model; taking the accuracy as the performance value of the candidate quantized model.

2. The model quantization method according to claim 1, characterized in that The step of determining multiple second candidate parameter values of the quantization parameters of the model to be quantized according to the multiple first candidate parameter values through a preset parameter determination strategy includes: Selecting multiple third candidate parameter values from the multiple first candidate parameter values; Performing mutation processing on each of the third candidate parameter values according to a preset mutation probability to obtain multiple second candidate parameter values of the quantization parameters of the model to be quantized.

3. The model quantization method according to claim 2, wherein Each of the second candidate parameter values follows a Gaussian distribution with each of the third candidate parameter values as the mean and a value determined according to the quantization bandwidth corresponding to the model to be quantized, the maximum value and the minimum value of the quantization parameters to be quantized as the variance.

4. The model quantization method according to claim 1, wherein, The step of obtaining multiple first candidate parameter values of the quantization parameters of the model to be quantized includes: Obtaining the parameter value range of the quantization parameters of the model to be quantized; Obtaining multiple first candidate parameter values of the quantization parameters from the parameter value range.

5. The model quantization method according to claim 4, wherein The step of obtaining the parameter value range of the quantization parameters of the model to be quantized includes: Determining a reference parameter value of the quantization parameters of the model to be quantized according to the maximum value of the quantization parameters to be quantized and the quantization bandwidth corresponding to the model to be quantized; Determining the parameter value range of the quantization parameters of the model to be quantized according to the reference parameter value.

6. The model quantization method according to any one of claims 1 to 5, characterized in that The step of determining the quantized model corresponding to the model to be quantized from the multiple reference quantized models that meet the preset stop condition includes: Determine the reference quantized model with the largest performance value from multiple reference quantized models that meet the preset stop condition to obtain the target quantized model; Determine the quantized model corresponding to the model to be quantized according to the target quantized model.

7. The model quantization method according to claim 6, wherein The determining the quantized model corresponding to the model to be quantized according to the target quantized model includes: Determine the target parameter value range of the target quantization parameter according to the parameter value of the target quantization parameter corresponding to the target quantized model; Retrieve the local optimal solution of the target quantization parameter from the target parameter value range to obtain the target parameter value of the target quantization parameter; Quantize the model to be quantized according to the target parameter value to obtain the target quantized model corresponding to the model to be quantized.

8. A model quantization device, characterized in that, including: The first processing module is used to obtain multiple first candidate parameter values of the quantization parameter of the model to be quantized, and quantize the model to be quantized according to each first candidate parameter value to obtain multiple first candidate quantized models; The second processing module is used to determine multiple second candidate parameter values of the quantization parameter of the model to be quantized according to the multiple first candidate parameter values through a preset parameter determination strategy, and quantize the model to be quantized according to each second candidate parameter value to obtain multiple second candidate quantized models; The first determination module is used to select multiple reference quantized models from the multiple first candidate quantized models and the multiple second candidate quantized models according to the performance value of each first candidate quantized model and the performance value of each second candidate quantized model, determine the parameter value corresponding to each reference quantized model as the first candidate parameter value, use each reference quantized model as the first candidate quantized model, and return to execute the step of determining multiple second candidate parameter values of the quantization parameter of the model to be quantized through a preset parameter determination strategy according to the multiple first candidate parameter values until the preset stop condition is met; The second determination module is used to determine the quantized model corresponding to the model to be quantized from multiple reference quantized models that meet the preset stop condition; When the candidate quantized model is an image classification model, obtain multiple test images and input them into the candidate quantized model to obtain the classification result of each test image; Determine the proportion of the correct classification result in all classification results to obtain the accuracy of the candidate quantized model; use the accuracy as the performance value of the candidate quantized model.

9. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium. When the computer program runs on a computer, the computer is caused to execute the model quantization method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory. A computer program is stored in the memory. The processor is used to execute the model quantization method according to any one of claims 1 to 7 by calling the computer program stored in the memory.

Citation Information

Patent Citations

  • Model quantification method and device, vehicle and storage medium

    CN114648116A

  • Model training method and device, storage medium and electronic equipment

    CN115034396A