Neural network model quantization method and apparatus, storage medium, and electronic device

By quantizing the pre-trained model before quantization perception training and optimizing the weight distribution, the problem of insufficient quantization accuracy in existing technologies is solved, enabling efficient application of deep learning models on mobile terminals.

CN115879525BActive Publication Date: 2026-07-21OPPO CHONGQING INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OPPO CHONGQING INTELLIGENT TECH CO LTD
Filing Date
2022-12-01
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing model quantization methods have poor quantization accuracy, resulting in significant accuracy loss and high computational cost when deep learning models are applied on mobile terminals.

Method used

Before quantization perception training, the pre-trained neural network model is quantized to determine the preset quantization accuracy. Then, the intermediate neural network model is quantized perception training using training data to optimize the weight distribution and improve the quantization accuracy.

Benefits of technology

By improving the weight distribution of the pre-trained model, the accuracy loss of the target neural network model is reduced, the quantization accuracy is improved, and it is suitable for mobile terminal devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879525B_ABST
    Figure CN115879525B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of model quantization, and particularly relates to a neural network model quantization method and device, a computer readable storage medium and an electronic device. The method comprises the following steps: obtaining a floating-point pre-training neural network model; determining a preset quantization precision, and performing quantization on the pre-training neural network model according to the preset quantization precision to obtain an intermediate neural network model; obtaining training data, and performing quantization perception training on the intermediate neural network model according to the preset quantization precision by using the training data to obtain a target neural network model. The technical scheme of the embodiment of the present disclosure improves the precision of the model quantization method, and overcomes the problem of large loss of model precision in the quantization process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of model quantization technology, and more specifically, to a method and apparatus for quantizing neural network models, a computer-readable storage medium, and an electronic device. Background Technology

[0002] With the rapid development of deep learning, the accuracy of deep learning models has been continuously improved. However, these deep learning models require huge hardware resources to apply and are not suitable for mobile devices. To solve the problem of applying high-precision deep learning models on mobile devices, quantization is usually used to obtain models that can be used on mobile devices.

[0003] However, the model quantization methods in related technologies have poor quantization accuracy, which will cause a loss of model accuracy, and the amount of computation in the quantization process is large.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this disclosure is to provide a neural network model quantization method, a neural network model quantization device, a computer-readable medium, and an electronic device, thereby improving the accuracy of the model quantization method to at least a certain extent and overcoming the problem of significant model accuracy loss during the quantization process.

[0006] According to a first aspect of this disclosure, a method for quantizing a neural network model is provided, comprising: acquiring a floating-point pre-trained neural network model; determining a preset quantization precision, and quantizing the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model; acquiring training data, and using the training data to perform quantization-aware training of the intermediate neural network model with a preset quantization precision to obtain a target neural network model.

[0007] According to a second aspect of this disclosure, a neural network model quantization apparatus is provided, comprising: an acquisition module for acquiring a floating-point pre-trained neural network model; a quantization module for determining a preset quantization precision and quantizing the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model; and a training module for acquiring training data and using the training data to perform quantization-aware training of the intermediate neural network model at a preset quantization precision to obtain a target neural network model.

[0008] According to a third aspect of this disclosure, a computer-readable medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method described above.

[0009] According to a fourth aspect of this disclosure, an electronic device is provided, characterized in that it includes: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described above.

[0010] One embodiment of this disclosure provides a method for quantizing a neural network model, which specifically includes acquiring a floating-point pre-trained neural network model; determining a preset quantization precision, and quantizing the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model; acquiring training data, and using the training data to perform quantization-aware training on the intermediate neural network model with a preset quantization precision to obtain a target neural network model. Compared to existing technologies, quantizing the pre-trained model before performing quantization-aware training improves the weight distribution of the pre-trained model, achieves high quantization precision, and reduces the precision loss of the obtained target neural network model.

[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0013] Figure 1 A schematic diagram of an exemplary system architecture to which embodiments of the present disclosure may be applied is shown;

[0014] Figure 2 This schematically illustrates a flowchart of a neural network model quantization method according to an exemplary embodiment of the present disclosure;

[0015] Figure 3 This schematically illustrates a flowchart of quantizing a pre-trained neural network model in an exemplary embodiment of the present disclosure;

[0016] Figure 4 This schematically illustrates a flowchart of determining a quantization range in an exemplary embodiment of the present disclosure;

[0017] Figure 5 This schematically illustrates another flowchart for determining a quantization range in an exemplary embodiment of the present disclosure;

[0018] Figure 6This diagram schematically illustrates a data structure for determining a quantization range in an exemplary embodiment of the present disclosure.

[0019] Figure 7 This schematically illustrates another data structure diagram for determining the quantization range in an exemplary embodiment of the present disclosure;

[0020] Figure 8 This schematically illustrates a flowchart of training an intermediate neural network model in an exemplary embodiment of the present disclosure;

[0021] Figure 9 This schematically illustrates a flowchart of updating quantization parameters in an exemplary embodiment of this disclosure;

[0022] Figure 10 This schematic diagram illustrates the composition of a neural network model quantization apparatus in an exemplary embodiment of the present disclosure.

[0023] Figure 11 A schematic diagram of an electronic device to which embodiments of the present disclosure may be applied is shown. Detailed Implementation

[0024] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0025] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0026] In related technologies, with the continuous development of machine learning and deep learning, deep neural networks are widely used in various fields such as autonomous driving, computer vision, natural language processing, speech recognition, and data processing. Taking material inspection in the field of autonomous driving as an example, a pre-trained neural network model is usually used to predict and process the input data to obtain the corresponding material inspection results. However, the data in the pre-trained neural network model is usually 32-bit floating-point numbers. In order to enable the pre-trained neural network model to be deployed on different hardware devices (such as mobile devices), the neural network model is usually quantized.

[0027] Neural network model quantization is a popular deep learning optimization method. For example, converting model data from 32-bit floating-point numbers (FP32) to fixed-point numbers (such as INT8) reduces the size of the quantized model, resulting in less memory usage. Furthermore, computation uses fixed-point multipliers instead of floating-point multipliers, leading to faster speeds and lower bandwidth consumption for memory access. Based on these advantages, model quantization has become an important means of accelerating network models on embedded devices. Therefore, how to make trained neural network models suitable for embedded devices has become a pressing problem to be solved.

[0028] When quantizing a model, related technologies typically first process the pre-trained model, then perform quantization-aware training on the pre-trained model to obtain the quantized model. However, the weights of the model trained on floating-point scales may not be suitable for quantization, ultimately leading to the quantization-aware training model's accuracy falling short of expectations. For example, the weight data distribution may be too scattered, or the extreme points may be too large.

[0029] To address the aforementioned drawbacks, this disclosure provides a method for quantizing neural network models. Figure 1 A schematic diagram of a system architecture for implementing the aforementioned neural network model quantization method is shown. This system architecture 100 may include a terminal 110 and a server 120. The terminal 110 may be a smartphone, tablet, desktop computer, laptop, or other terminal device. The server 120 generally refers to a backend system providing services related to the neural network model quantization method in this exemplary embodiment, and may be a single server or a cluster of multiple servers. The terminal 110 and the server 120 can be connected via a wired or wireless communication link for data interaction.

[0030] In one implementation, the aforementioned neural network model quantization method can be executed by terminal 110. For example, a user loads a pre-trained neural network model using terminal 110, and terminal 110 performs quantization-based perceptual training on the pre-trained neural network model to obtain the target neural network model.

[0031] In one implementation, the above-described neural network model quantization method can be executed by server 120. For example, a user loads a pre-trained neural network model using terminal 110, which uploads the pre-trained neural network model to server 120. Server 120 then performs quantization-based perceptual training to obtain the target neural network model and returns the target neural network model to terminal 110.

[0032] As can be seen from the above, the execution subject of the neural network model quantization method in this exemplary embodiment can be the aforementioned terminal 110 or server 120, and this disclosure does not limit it in this regard.

[0033] The following is combined with Figure 2 The method for quantizing neural network models in this exemplary embodiment will be described. It should be noted that the neural network model in this disclosure can be a model used in a target domain. This target domain can be image processing, autonomous driving, computer vision, natural language processing, speech recognition, and data processing, among other fields. In other words, the algorithm for improving the accuracy of neural network model training through quantization provided in this application can be applied to image target detection and segmentation, LiDAR point cloud target detection, speech recognition, and natural language processing, among other fields.

[0034] Figure 2 An exemplary flow of the neural network model quantization method is shown, which may include:

[0035] Step S210: Obtain the floating-point pre-trained neural network model;

[0036] Step S220: Determine the preset quantization precision, and quantize the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model;

[0037] Step S230: Obtain training data and use the training data to perform quantization perception training on the intermediate neural network model with a preset quantization precision to obtain the target neural network model.

[0038] Based on the above method, the pre-trained model was quantized before quantization perception training, which improved the weight distribution of the pre-trained model, achieved high quantization accuracy, and reduced the accuracy loss of the obtained target neural network model.

[0039] The following is about Figure 2 Each step in the process will be explained in detail.

[0040] refer to Figure 2 In step S210, a floating-point pre-trained neural network model is obtained.

[0041] In one exemplary embodiment, the model data (such as model weights and activation layer data) of the above-mentioned floating-point pre-trained neural network model can be 32-bit floating-point numbers, or 64-bit floating-point numbers, or can be customized according to user needs. No specific limitation is made in this example embodiment.

[0042] In step S220, a preset quantization precision is determined, and the pre-trained neural network model is quantized according to the preset quantization precision to obtain an intermediate neural network model.

[0043] In one example embodiment of this disclosure, after obtaining the above-mentioned floating-point pre-trained neural network model, a preset quantization precision can be determined. The preset quantization precision can be 8 bits, or it can be customized according to user needs. In this example embodiment, no specific limitation is made.

[0044] In this example implementation, refer to Figure 3 As shown, the above-mentioned quantization of the pre-trained neural network model according to the preset quantization precision to obtain the intermediate neural network model may include steps S310 to S320.

[0045] In step S310, the quantization range of each layer of the pre-trained neural network model is determined.

[0046] In this example implementation, refer to Figure 4 As shown, steps S410 to S420 can be performed on any layer of the network in the pre-trained neural network model.

[0047] In step S410, the set of weight values ​​in each layer of the network is obtained.

[0048] In this example implementation, the weight values ​​of each layer of the pre-trained neural network model can be extracted first to obtain all the weight values ​​between each node in each layer, thus forming a set of weight values.

[0049] In step S420, the quantization range of each network layer is determined based on the set of weight values ​​and the preset scaling factor.

[0050] In this example, in real-time mode, refer to Figure 5 As shown, after obtaining the set of weight values ​​in each layer of the network, the quantization range of each layer of the network can be determined based on the set of weight values ​​and the preset scaling factor. Specifically, this can include steps S510 to S520.

[0051] In step S510, the maximum and minimum values ​​are determined in the set of weight values ​​according to the preset proportional coefficient.

[0052] In this example implementation, a preset ratio coefficient can be set first, for example, 99%. The preset ratio coefficient can also be 98%, 97%, etc., or it can be customized according to user needs. No specific limitation is made in this example implementation.

[0053] In this example implementation, after obtaining the preset ratio coefficient, the maximum and minimum values ​​can be determined in the weight set based on the preset ratio coefficient.

[0054] For example, quantizing the aforementioned pre-trained neural network model can include object quantization and asymmetric quantization. When the quantization process is symmetric quantization, refer to... Figure 6 As shown, the absolute values ​​of the maximum and minimum values ​​are equal, which allows us to determine the maximum value (max) and the minimum value (min), ensuring that the weight value of the preset proportional coefficient falls between the maximum and minimum values.

[0055] When the above quantization is asymmetric quantization, refer to Figure 7 As shown, the maximum and minimum values ​​can be determined separately. For example, if the preset ratio coefficient is 99%, the weight values ​​are sorted to determine a maximum value so that the weight values ​​of 99% greater than or equal to 0 are less than or equal to the maximum value, and a minimum value is determined so that the weight values ​​of 99% less than or equal to 0 are less than or equal to the minimum value.

[0056] In step S520, the quantization range is determined based on the maximum value and the minimum value.

[0057] In this example implementation, after obtaining the maximum and minimum values, the range between the maximum and minimum values ​​can be used as the quantization range. That is, when quantizing the pre-trained model, only the weight values ​​within the quantization range are quantized.

[0058] It should be noted that the above-mentioned methods for determining the quantization range can also include the KLD (Kullback-Leiblerdivergense) method, the percentile method, MSE (Mean Squared Error), L2-norm, etc. The percentile quantization method can directly select 99.9% or 99.99% of the total data elements after sorting, or it can be customized according to user needs; this example implementation does not impose specific limitations.

[0059] In step S320, at least one layer of the pre-trained neural network model is quantized according to the preset quantization precision and the quantization range to obtain an intermediate neural network model.

[0060] In this example implementation, after obtaining the above quantization range, the intermediate quantization bit number can be determined based on the above preset quantization precision. Specifically, the intermediate quantization bit number can be less than or equal to the above preset quantization precision. For example, if the above preset quantization precision is 8 bits, then the intermediate quantization bit number can be 6 bits, 7 bits, 4 bits, etc., or it can be customized according to user needs. In this example implementation, no specific limitation is made.

[0061] In this example implementation, after determining the intermediate quantization bit depth, the at least one layer of the network can be quantized based on the intermediate quantization precision and the quantization range to obtain an intermediate neural network model. Specifically, the quantization process involves converting the original 32-bit floating-point number (FP32) into a fixed-point number, such as INT6.

[0062] After obtaining the above intermediate neural network model, step S230 can be executed.

[0063] In step S230, training data is acquired, and the intermediate neural network model is subjected to quantization perception training with a preset quantization precision using the training data to obtain the target neural network model.

[0064] In this example implementation, training data can be acquired first. The training data can be determined according to different application fields. For example, if the above model is applied to the field of image processing technology, the training data can include multiple original images and calibration data of the original images.

[0065] In this example implementation, after obtaining the training data, the training data can be preprocessed. The preprocessing operation may include image scaling, feature extraction, etc., and may also be determined according to the input type of the pre-trained model. No specific limitation is made here.

[0066] After obtaining the training data, the target neural network model can be obtained by performing quantization perception training on the intermediate neural network model with a preset quantization precision using the training data.

[0067] Specifically, the above training data can be used to train the above intermediate neural network model for a preset number of iterations. The preset number of iterations can be 100 times, 200 times, or can be customized according to user needs. In this example implementation, no specific limitation is made.

[0068] During iterative training, refer to Figure 8 As shown, steps S10 to S840 may be included.

[0069] In step S810, the quantization parameters of the intermediate neural network model are determined during each iteration of training.

[0070] In this example embodiment, where the application technology is image processing, the aforementioned intermediate neural network model can be a convolutional neural network model. This intermediate neural network model includes multiple layers, each containing parameters (weights). The output value (activation value) of the previous layer is used as the input value of the next layer. When training a deep learning model based on image data, the output value of the previous layer is used as the input value of the next layer for convolution calculation to obtain a feature map composed of pixels (activation values). During quantization-aware training, model quantization can be completed through multiple iterations of model quantization training.

[0071] During the training of the aforementioned intermediate neural network using the training data, the model weights will change as the training progresses. In each iteration of model training, the quantization parameters need to be redefined based on the numerical distribution characteristics of the model's weights or activation values. Since the weights, input values, and output values ​​of each network layer in the model may be floating-point numbers or integers of different lengths, the quantization parameters are determined based on the numerical distribution characteristics of the weights or activation values ​​of each network layer in the model to make it easier to find the optimal model parameters when quantizing the weights and activation values ​​separately according to the quantization parameters.

[0072] In step S820, the weights and activation values ​​of each layer of the intermediate neural network are quantized according to the preset quantization precision and quantization parameters to obtain weight quantization values ​​and activation quantization values.

[0073] In this example implementation, after obtaining the above quantization parameters, during iteration, the weights and activation values ​​of each layer of the intermediate neural network can be quantized according to the above preset quantization precision and quantization parameters to obtain weight quantization values ​​and activation quantization values.

[0074] Specifically, the weight and activation values ​​of each layer can be quantized to obtain the quantized weight and activation values ​​of each layer of the network.

[0075] In step S830, a forward operation is performed based on the weighted quantization value and the activation quantization value to obtain the next activation value, and the quantization parameters are updated.

[0076] In this example implementation, the above model can be applied to the field of computer vision. After quantizing the weights and activation values ​​of each layer of the model according to quantization parameters, and obtaining the quantized weight and activation values, a forward pass is performed using the simulated quantized weight and activation values ​​to perform image analysis and output the next activation value. Since computer vision deep learning models can be intermediate neural network models for solving any problem such as image classification, object detection, object segmentation, image style transfer, image colorization, image reconstruction, image super-resolution, or image synthesis, when performing the forward pass using simulated quantized weight and activation values, image analysis is performed according to the model's function. For example, if the intermediate neural network model is an image classification model, then when performing the forward pass using simulated quantized weight and activation values, image classification is performed, and the classification result, i.e., the next activation value, is output. When training a deep learning model based on image data, a fixed-point network model can be trained by performing forward computation based on the quantized weight and activation values.

[0077] In this example implementation, the quantization parameters are updated after each iteration. Specifically, refer to... Figure 9 As shown, steps S910 to S920 may be included.

[0078] In step S910, a first update ratio and a second update ratio are set;

[0079] In this example implementation, the first update ratio can be 80%, 90%, or customized according to user needs. No specific limitation is made in this example implementation.

[0080] The second update ratio mentioned above can be 60%, 50%, etc., or it can be customized according to user needs. In this example implementation, no specific limitation is made.

[0081] In step S920, the range of activation values ​​is updated based on the activation value of each iteration, the number of iterations, the first update ratio, and the second update ratio.

[0082] In one exemplary embodiment of this disclosure, the quantization parameters can be updated using the quantization parameters obtained after the first to Nth iterations. Specifically, a first number of quantization parameters obtained from the first to Nth iterations that are greater than the initial quantization parameters is determined; a second number of quantization parameters obtained from the first to Nth iterations that are less than the initial quantization parameters is determined; in response to the ratio of the first number to the iteration number N being greater than a first update ratio, the initial quantization parameters are updated using the average value of the quantization parameters obtained from the first to Nth iterations; in response to the ratio of the second number to the iteration number N being greater than a second update ratio, the initial quantization parameters are updated using the average value of the quantization parameters obtained from the first to Nth iterations. Wherein, N is a positive integer greater than or equal to 2.

[0083] In this example implementation, the first quantity can represent the number of activation values ​​obtained from the first to the Nth iterations of training whose maximum value is greater than the maximum value of the activation values ​​in the initial quantization parameters. The second quantity can represent the number of activation values ​​obtained from the first to the Nth iterations of training whose maximum value is less than the maximum value of the activation values ​​in the initial quantization parameters.

[0084] When the ratio of the first quantity to the number of iterations N is greater than the first update ratio, max1 = 0.999 * max0 + 0.001 * max 平均 Where max1 is the maximum value of the updated quantization parameters, max0 is the maximum value of the initial quantization parameters, and max... 平均 This is the average of the maximum values ​​of the quantized parameters obtained from the first to the Nth iterations of training.

[0085] When the ratio of the second quantity to the number of iterations N is greater than the second update ratio, max1 = 0.999 * max0 - 0.001 * max 平均 Where max1 is the maximum value of the updated quantization parameters, max0 is the maximum value of the initial quantization parameters, and max... 平均 This is the average of the maximum values ​​of the quantization parameters obtained from the first to the Nth iterations of training.

[0086] In this example implementation, the first quantity can be the number of times the minimum value among the activation values ​​obtained from the first to the Nth iterations of training is greater than the minimum value among the initial quantization parameters. The second quantity can represent the number of times the minimum value among the activation values ​​obtained from the first to the Nth iterations of training is greater than the minimum value among the initial quantization parameters.

[0087] When the ratio of the first quantity to the number of iterations N is greater than the first update ratio, min1 = 0.999 * min0 + 0.001 * min 平均 Where, min1 is the minimum value of the updated quantization parameters, min0 is the minimum value of the initial quantization parameters, and min... 平均 This is the average of the minimum values ​​of the quantization parameters obtained from the first to the Nth iterations of training.

[0088] When the ratio of the second quantity to the number of iterations N is greater than the second update ratio, min1 = 0.999*min0 - 0.001*min 平均 Where, min1 is the minimum value of the updated quantization parameters, min0 is the minimum value of the initial quantization parameters, and min... 平均 This is the average of the minimum values ​​of the quantization parameters obtained from the first to the Nth iterations of training.

[0089] In step S840, the weights of the intermediate neural network are updated according to the intermediate weight values ​​and the intermediate activation values ​​to obtain the target neural network.

[0090] In this example implementation, the error can also be calculated based on the model's weight quantization and activation quantization values, and the gradient can be backpropagated based on this error so that the intermediate neural network model adjusts the weights of each network layer according to the gradient. After updating the original floating-point weights of the model by performing forward and backward calculations using the weight quantization and activation quantization values, one round of model quantization training is completed. After updating the weights of the intermediate neural network model, the target neural network model is obtained.

[0091] In summary, in this exemplary embodiment, the pre-trained model is quantized before quantization-based training, which improves the weight distribution of the pre-trained model, achieves high quantization accuracy, and reduces the accuracy loss of the obtained target neural network model. Furthermore, when quantizing the pre-trained model, the quantization range is considered, and the weight distribution is optimized, thereby improving the quantization accuracy of the pre-trained model.

[0092] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0093] Further reference Figure 10 As shown, this example embodiment also provides a neural network model quantization device 1000, including an acquisition module 1010, a quantization module 1020, and a training module 30. Wherein:

[0094] The acquisition module 1010 can be used to acquire floating-point pre-trained neural network models.

[0095] The quantization module 1020 can be used to determine a preset quantization precision and quantize the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model.

[0096] In one example implementation, the quantization module 1020 can be configured to determine an intermediate quantization bit depth based on the preset quantization precision; and to quantize the pre-trained neural network model based on the quantization bit depth to obtain an intermediate neural network model.

[0097] In one example implementation, the quantization module 1020 can be configured to determine the quantization range of each layer in the pre-trained neural network model; and to quantize at least one layer in the pre-trained neural network model according to the preset quantization precision and the quantization range to obtain an intermediate neural network model. Specifically, it obtains the set of weight values ​​in each layer; and determines the quantization range of each layer according to the set of weight values ​​and a preset scaling factor.

[0098] In this real-time example, the maximum and minimum values ​​can first be determined in the set of weight values ​​according to the preset proportional coefficient; then the quantization range can be determined according to the maximum and minimum values.

[0099] The training module 1030 can be used to acquire training data and use the training data to perform quantization perception training on the intermediate neural network model with a preset quantization precision to obtain the target neural network model.

[0100] In one example implementation, the training module 1030 can be configured to perform a preset number of iterative training cycles on the intermediate neural network model using the training data; during each iteration, determine the quantization parameters of the intermediate neural network model; quantize the weights and activation values ​​of each layer of the intermediate neural network according to the preset quantization precision and quantization parameters to obtain weight quantization values ​​and activation quantization values; perform forward operations based on the weight quantization values ​​and activation quantization values ​​to obtain the next activation value and update the quantization parameters; and update the weights of the intermediate neural network according to the quantization parameters to obtain the target neural network.

[0101] In one example implementation, the training module 1030 can be configured such that the quantization parameters include a range of activation values, and updating the quantization parameters includes: setting a first update ratio and a second update ratio; and updating the range of activation values ​​based on the activation values ​​of each iteration, the number of iterations, the first update ratio, and the second update ratio.

[0102] The specific details of each module in the above-mentioned device have been described in detail in the method section of the implementation. For any undisclosed details, please refer to the implementation content of the method section, and therefore will not be repeated here.

[0103] Exemplary embodiments of this disclosure also provide an electronic device for performing the above-described neural network model quantization method. This electronic device may be the terminal 110 or the server 120 described above. Generally, the electronic device may include a processor and a memory, the memory storing executable instructions of the processor, and the processor configured to perform the above-described neural network model quantization method by executing the executable instructions.

[0104] The following is based on Figure 11Taking the mobile terminal 1100 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 11 The structure can also be applied to fixed types of equipment.

[0105] like Figure 11 As shown, the mobile terminal 1100 may specifically include: a processor 1101, a memory 1102, a bus 1103, a mobile communication module 1104, an antenna 1, a wireless communication module 1105, an antenna 2, a display screen 1106, a camera module 1107, an audio module 1108, a power module 1109, and a sensor module 1110.

[0106] Processor 1101 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The neural network model quantization method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU.

[0107] The processor 1101 can be connected to the memory 1102 or other components via the bus 1103.

[0108] The memory 1102 can be used to store computer executable program code, which includes instructions. The processor 1101 executes various functional applications and data processing of the mobile terminal 1100 by running the instructions stored in the memory 1102. The memory 1102 can also store application data, such as images, videos, and other files.

[0109] The communication function of mobile terminal 1100 can be implemented through mobile communication module 1104, antenna 1, wireless communication module 1105, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1104 can provide 2G, 3G, 4G, and 5G mobile communication solutions for mobile terminal 1100. Wireless communication module 1105 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1100.

[0110] The display screen 1106 is used to implement display functions, such as displaying the user interface, images, and videos. The camera module 1107 is used to implement shooting functions, such as capturing images and videos. The audio module 208 is used to implement audio functions, such as playing audio and capturing voice. The power module 209 is used to implement power management functions, such as charging the battery, powering the device, and monitoring the battery status. The sensor module 1110 may include a depth sensor 11101, a pressure sensor 11102, a gyroscope sensor 11103, a barometric pressure sensor 11104, etc., to implement corresponding sensing and detection functions.

[0111] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0112] Exemplary embodiments of this disclosure also provide a computer-readable storage medium having a program product stored thereon capable of implementing the methods described above in this specification. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.

[0113] It should be noted that the computer-readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0114] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0115] Furthermore, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0116] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0117] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A neural network model quantization method, applied in the field of image processing technology, characterized in that, include: Obtain a floating-point pre-trained neural network model, wherein the pre-trained neural network model is used for deployment and operation on mobile terminals and embedded devices; A preset quantization precision is determined, and the pre-trained neural network model is quantized according to the preset quantization precision to obtain an intermediate neural network model, wherein the quantization includes converting the model data from floating-point numbers to fixed-point numbers; Acquire training data and use the training data to perform quantization perception training on the intermediate neural network model with a preset quantization precision to obtain the target neural network model; the training data includes multiple original images and calibration data for the original images; The obtained target neural network model includes: The intermediate neural network model is trained iteratively a predetermined number of times using the training data; During each iteration of training, the quantization parameters of the intermediate neural network model are determined; The weights and activation values ​​of each layer of the intermediate neural network are quantized according to the preset quantization precision and quantization parameters to obtain the weight quantization value and activation quantization value. A forward operation is performed based on the weighted quantization value and the activation quantization value to obtain the next activation value, and the quantization parameters are updated. The target neural network is obtained by updating the weights of the intermediate neural network according to the quantization parameters; The quantization parameters include a range of activation values, and updating the quantization parameters includes: Set the first update ratio and the second update ratio; The activation value range is updated based on the activation value, the number of iterations, the first update ratio, and the second update ratio for each iteration.

2. The method according to claim 1, characterized in that, The step of quantizing the pre-trained neural network model according to the preset quantization precision to obtain the intermediate neural network model includes: The intermediate quantization bit depth is determined according to the preset quantization precision; The pre-trained neural network model is quantized according to the quantization bit depth to obtain an intermediate neural network model.

3. The method according to claim 1, characterized in that, The step of quantizing the pre-trained neural network model according to the preset quantization precision to obtain the intermediate neural network model includes: Determine the quantization range of each layer in the pre-trained neural network model; The intermediate neural network model is obtained by quantizing at least one layer of the pre-trained neural network model according to the preset quantization precision and the quantization range.

4. The method according to claim 3, characterized in that, Determining the quantization range of each layer in the pre-trained neural network model includes: Obtain the set of weight values ​​for each layer of the network; The quantization range of each network layer is determined based on the set of weight values ​​and the preset scaling factor.

5. The method according to claim 4, characterized in that, The quantization range of each network layer is determined based on the set of weight values ​​and a preset scaling factor, including: The maximum and minimum values ​​are determined in the set of weight values ​​according to the preset proportional coefficient; The quantization range is determined based on the maximum value and the minimum value.

6. A neural network model quantization device, applied in the field of image processing technology, characterized in that, include: An acquisition module is used to acquire a floating-point pre-trained neural network model, wherein the pre-trained neural network model is used for deployment and operation on mobile terminals and embedded devices; A quantization module is used to determine a preset quantization precision and quantize the pre-trained neural network model according to the preset quantization precision to obtain an intermediate neural network model, wherein the quantization includes converting the model data from floating-point numbers to fixed-point numbers; The training module is used to acquire training data and use the training data to perform quantization perception training on the intermediate neural network model with a preset quantization precision to obtain the target neural network model; the training data includes multiple original images and calibration data of the original images; The process of obtaining the target neural network model includes: performing a preset number of iterative training iterations on the intermediate neural network model using the training data; determining the quantization parameters of the intermediate neural network model during each iteration; quantizing the weights and activation values ​​of each layer of the intermediate neural network according to the preset quantization precision and quantization parameters to obtain weight quantization values ​​and activation quantization values; performing forward operations based on the weight quantization values ​​and activation quantization values ​​to obtain the next activation value and updating the quantization parameters; and updating the weights of the intermediate neural network according to the quantization parameters to obtain the target neural network. The quantization parameters include the range of activation values. Updating the quantization parameters includes: setting a first update ratio and a second update ratio; and updating the range of activation values ​​based on the activation value, the number of iterations, the first update ratio, and the second update ratio for each iteration.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the neural network model quantization method as described in any one of claims 1 to 4.

8. An electronic device, characterized in that, include: One or more processors; as well as A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the neural network model quantization method as described in any one of claims 1 to 4.