Mixing precision quantification method and device of neural network, equipment and storage medium

By calculating the quantization error of the network to be quantized and determining the target network layer, the problem of inefficient hybrid accuracy quantization of neural networks is solved, a more efficient quantization strategy is achieved, and the balance between calculation quantity and accuracy is improved.

CN120218134APending Publication Date: 2025-06-27HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311813299.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The mixed accuracy quantization efficiency of neural networks in the prior art makes it difficult to achieve a good trade-off between network computing volume and accuracy.

Method used

By acquiring the network to be quantized and determining the evaluation layer, the quantization errors of low-precision quantization and high-precision quantization are calculated, and the target high-precision and low-precision network layers are determined based on these errors to form a target hybrid precision network.

Benefits of technology

The hybrid accuracy quantization efficiency of neural networks is improved, and the need to search for each layer of quantization strategy for the quantization network is reduced, thereby achieving a better balance between computational volume and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218134A_ABST
    Figure CN120218134A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid precision quantification method and device for a neural network, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-quantized network, and determining an evaluation layer under the to-be-quantized network; determining a first quantization error of the evaluation layer in the to-be-quantized network with low-precision quantization, and determining a second quantization error of the evaluation layer in the to-be-quantized network with high-precision quantization; and determining a target high-precision network layer and a target low-precision network layer in the to-be-quantized network based on the first quantization error and the second quantization error, and forming a target mixed precision network of the to-be-quantized network based on the target high-precision network layer and the target low-precision network layer. According to the method, the high-precision network layer and the low-precision network layer are determined by respectively calculating the first quantization error of low-precision quantization and the second quantization error of high-precision quantization, so that the target hybrid precision network of the to-be-quantized network is formed, and the hybrid precision quantization efficiency of the neural network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and particularly to a method, device, equipment and storage medium for mixed-precision quantization of a neural network. Background Art

[0002] Currently, learning models based on neural networks are widely used in various fields. With the development of deep learning, the performance of the models has been continuously improved, but at the same time, it has also caused a large number of parameters and calculations in the neural network models. Therefore, the field has proposed mixed-precision quantization of neural networks, where mixed-precision quantization is to quantize the weights of each layer of the neural network into different precisions, such as converting the floating-point calculations of some network layers into low-bit fixed-point calculations, and other network layers still collect floating-point calculations, so as to obtain a model with a smaller volume.

[0003] However, although the mixed-precision quantization of neural networks can effectively reduce the model calculation intensity, parameter size and memory consumption, it often brings a large accuracy loss. Therefore, how to search for a mixed-precision quantization strategy to achieve a better compromise between network calculation volume and accuracy has been widely studied.

[0004] In related technologies, usually, the quantization strategy of each layer of the network is searched, and after each search, the quantized network is fine-tuned to more accurately evaluate the quantization strategy. However, this method requires inference on a large number of validation sets provided by users, resulting in low efficiency of the mixed-precision quantization of neural networks. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, equipment and storage medium for mixed-precision quantization of a neural network, aiming to solve the technical problem of low efficiency of the mixed-precision quantization of neural networks in the prior art.

[0006] To achieve the above purpose, this application provides a method for mixed-precision quantization of a neural network, and the method for mixed-precision quantization of the neural network includes:

[0007] Obtain a network to be quantized, and determine an evaluation layer under the network to be quantized;

[0008] Determine a first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine a second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized;

[0009] Based on the first quantization error and the second quantization error, determine the target high-precision network layers and target low-precision network layers in the network to be quantized, and form a target mixed-precision network of the network to be quantized based on the target high-precision network layers and target low-precision network layers.

[0010] Optionally, the step of determining the target high-precision network layers and target low-precision network layers in the network to be quantized based on the first quantization error and the second quantization error includes:

[0011] Calculate the target influence factors of each network layer of the evaluation layer based on the first quantization error and the second quantization error;

[0012] Determine an influence factor threshold based on the target influence factors of each network layer of the evaluation layer;

[0013] Screen each network layer based on the target influence factors of each network layer of the evaluation layer and the influence factor threshold to obtain the target high-precision network layers and target low-precision network layers in the network to be quantized.

[0014] Optionally, the step of screening each network layer based on the target influence factors of each network layer of the evaluation layer and the influence factor threshold to obtain the target high-precision network layers and target low-precision network layers in the network to be quantized includes:

[0015] Compare the target influence factors of each network layer of the evaluation layer with the influence factor threshold to obtain a comparison result;

[0016] Use the network layers in the comparison result whose target influence factors are greater than or equal to the influence factor threshold as the target high-precision network layers, and use the network layers in the comparison result whose target influence factors are less than the influence factor threshold as the target low-precision network layers.

[0017] Optionally, the step of calculating the target influence factors of each network layer of the evaluation layer based on the first quantization error and the second quantization error includes:

[0018] Obtain a mode strategy, where the mode strategy includes a performance mode strategy and a precision mode strategy, and the performance mode strategy includes the following formula:

[0019]

[0020] The precision mode strategy includes the following formula:

[0021] impactFactor = errorln8Bit - errorlnMPLayer

[0022] Among them, impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient;

[0023] Based on the pattern strategy, determine the impact factor calculation algorithm;

[0024] Based on the impact factor calculation algorithm, calculate the impact factor for the first quantization error and the second quantization error to obtain the target impact factor of the evaluation layer for each network layer.

[0025] Optionally, the step of determining the impact factor threshold based on the target impact factor of the evaluation layer for each network layer includes:

[0026] Sort the target impact factors of the evaluation layer for each network layer in descending order, and extract a preset number of consecutive target impact factors from the sorted target impact factors to obtain multiple groups of consecutive target impact factors;

[0027] Calculate the mean of each group of consecutive target impact factors to obtain the mean of each group of target impact factors;

[0028] Determine the mean inflection point among the means of each group of target impact factors, and use the target impact factor corresponding to the mean inflection point as the impact factor threshold.

[0029] Optionally, the step of determining the evaluation layer under the network to be quantized includes:

[0030] Determine each network layer under the network to be quantized;

[0031] Based on the attribute information of each network layer, determine the evaluation layer under the network to be quantized, where the evaluation layer includes an output layer, a quantization-sensitive layer, and a private layer;

[0032] The steps of determining the first quantization error output by the evaluation layer in the network to be quantized with low precision and determining the second quantization error output by the evaluation layer in the network to be quantized with high precision include:

[0033] Based on the evaluation layer, perform multi-layer joint evaluation on the network to be quantized with low-precision quantization and the network to be quantized with high-precision quantization respectively, to obtain the first quantization error of each evaluation layer in the network to be quantized with low-precision quantization, and to obtain the second quantization error of each evaluation layer in the network to be quantized with high-precision quantization.

[0034] Optionally, the step of determining the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error includes:

[0035] Based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer, calculate the initial influence factor of each evaluation layer with respect to each network layer;

[0036] Perform joint calculation on the initial influence factors of each evaluation layer with respect to each network layer to obtain the target influence factor of each evaluation layer with respect to each network layer;

[0037] Based on the target influence factor of each evaluation layer with respect to each network layer, determine the target high-precision network layer and the target low-precision network layer in the network to be quantized.

[0038] Optionally, the step of determining the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization and determining the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization includes:

[0039] Based on a preset input sample and the evaluation layer, perform floating-point inference on the network to be quantized to obtain a floating-point result;

[0040] Set each network layer in the network to be quantized as a low-precision network layer, and based on the input sample, perform low-precision quantization inference on the low-precision network layer to obtain the low-precision quantization result of each low-precision network layer output by the evaluation layer, and based on the floating-point result and the low-precision quantization result, calculate the first quantization error of the evaluation layer;

[0041] Set each network layer in the network to be quantized as a high-precision network layer layer by layer, and based on the input sample, perform high-precision quantization inference on the high-precision network layer layer by layer to obtain the high-precision quantization result of each high-precision network layer output by the evaluation layer, and based on the floating-point result and the high-precision quantization result, obtain the second quantization error of the evaluation layer.

[0042] The present application also provides a neural network hybrid precision quantization device, and the neural network hybrid precision quantization device includes:

[0043] An acquisition module, configured to acquire a network to be quantized and determine an evaluation layer under the network to be quantized;

[0044] An error calculation module, configured to determine a first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine a second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized;

[0045] A determination module, configured to determine a target high-precision network layer and a target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error, and form a target mixed-precision network of the network to be quantized based on the target high-precision network layer and the target low-precision network layer.

[0046] And / or, the determination module includes: an influence factor calculation module, configured to calculate a target influence factor of the evaluation layer with respect to each network layer based on the first quantization error and the second quantization error; a threshold determination module, configured to determine an influence factor threshold based on the target influence factor of the evaluation layer with respect to each network layer; a screening module, configured to screen each network layer based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold, to obtain a target high-precision network layer and a target low-precision network layer in the network to be quantized;

[0047] And / or, the screening module includes: a comparison module, configured to compare the target influence factor of the evaluation layer with respect to each network layer with the influence factor threshold to obtain a comparison result; a network layer quantization module, configured to use the network layer with the target influence factor greater than or equal to the influence factor threshold in the comparison result as the target high-precision network layer, and use the network layer with the target influence factor less than the influence factor threshold in the comparison result as the target low-precision network layer;

[0048] And / or, the influence factor calculation module includes: a strategy acquisition module, configured to acquire a mode strategy, where the mode strategy includes a performance mode strategy and a precision mode strategy, and the performance mode strategy includes the following formula:

[0049]

[0050] The precision mode strategy includes the following formula:

[0051] impactFactor = errorln8Bit - errorlnMPLayer

[0052] Among them, impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient; the algorithm determination module is used to determine the impact factor calculation algorithm based on the mode strategy; the target impact factor calculation module is used to calculate the impact factor of the first quantization error and the second quantization error based on the impact factor calculation algorithm to obtain the target impact factor of each network layer of the evaluation layer;

[0053] And / or, the threshold determination module includes: an extraction module, which is used to sort the target impact factors of each network layer of the evaluation layer in descending order, and extract a preset number of consecutive target impact factors from the sorted target impact factors to obtain multiple groups of consecutive target impact factors; a mean calculation module, which is used to calculate the mean of each group of consecutive target impact factors to obtain the mean of each group of target impact factors; an inflection point determination module, which is used to determine the mean inflection point in the mean of each group of target impact factors, and use the target impact factor corresponding to the mean inflection point as the impact factor threshold;

[0054] And / or, the acquisition module includes: a network layer determination module, which is used to determine each network layer under the network to be quantized; an evaluation layer determination module, which is used to determine the evaluation layer under the network to be quantized based on the attribute information of each network layer, where the evaluation layer includes an output layer, a quantization-sensitive layer, and a private layer; a multi-layer joint evaluation module, which is used to perform multi-layer joint evaluation on the network to be quantized with low-precision quantization and the network to be quantized with high-precision quantization based on the evaluation layer to obtain the first quantization error of each evaluation layer in the network to be quantized with low-precision quantization, and obtain the second quantization error of each evaluation layer in the network to be quantized with high-precision quantization;

[0055] And / or, the determination module further includes: an initial impact factor calculation module, which is used to calculate the initial impact factor of each evaluation layer with respect to each network layer based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer; a joint calculation module, which is used to perform joint calculation on the initial impact factors of each evaluation layer with respect to each network layer to obtain the target impact factor of each evaluation layer with respect to each network layer; a quantization determination module, which is used to determine the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the target impact factor of each evaluation layer with respect to each network layer;

[0056] And / or, the error calculation module includes: a floating-point inference module, configured to perform floating-point inference on the network to be quantized based on a preset input sample and the evaluation layer, and obtain a floating-point result; a first error determination module, configured to set each network layer in the network to be quantized as a low-precision network layer, perform low-precision quantization inference on the low-precision network layer based on the input sample, obtain a low-precision quantization result of the evaluation layer for each low-precision network layer, and calculate a first quantization error of the evaluation layer based on the floating-point result and the low-precision quantization result; a second error determination module, configured to set each network layer in the network to be quantized as a high-precision network layer layer by layer, perform high-precision quantization inference on the high-precision network layer based on the input sample, obtain a high-precision quantization result of the evaluation layer for each high-precision network layer, and calculate a second quantization error of the evaluation layer based on the floating-point result and the high-precision quantization result.

[0057] The present application further provides a neural network hybrid precision quantization device, where the neural network hybrid precision quantization device includes: a memory, a processor, and a program stored on the memory for implementing the neural network hybrid precision quantization method.

[0058] The memory is used to store a program for implementing the neural network hybrid precision quantization method.

[0059] The processor is configured to execute the program for implementing the neural network hybrid precision quantization method to implement the steps of the neural network hybrid precision quantization method.

[0060] The present application further provides a storage medium, where a program for implementing the neural network hybrid precision quantization method is stored on the storage medium, and the program for implementing the neural network hybrid precision quantization method is executed by a processor to implement the steps of the neural network hybrid precision quantization method.

[0061] The present application calculates a first quantization error of the network to be quantized with low-precision quantization and a second quantization error of the network to be quantized with high-precision quantization respectively through the evaluation layer under the network to be quantized, and determines high-precision network layers that meet the conditions under the network to be quantized based on the first quantization error and the second quantization error, so as to form a target hybrid precision network of the network to be quantized, without performing a quantization strategy search for each layer of the network to be quantized, thereby improving the efficiency of hybrid precision quantization of the neural network. Description of the Drawings

[0062] The accompanying drawings herein are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. To more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic flowchart of the first embodiment of the mixed-precision quantization method for the neural network of the present application;

[0064] Figure 2 It is a schematic flowchart of the second embodiment of the mixed-precision quantization method for the neural network of the present application;

[0065] Figure 3 It is a schematic flowchart of the third embodiment of the mixed-precision quantization method for the neural network of the present application;

[0066] Figure 4 It is a schematic diagram of the modules of the mixed-precision quantization device for the neural network of the present application;

[0067] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the present application.

[0068] The realization, functional features, and advantages of the purpose of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific Embodiments

[0069] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] Referring to Figure 1 Figure 1 It is a schematic flowchart of the first embodiment of the mixed-precision quantization method for the neural network of the present application.

[0071] In the first embodiment, the mixed-precision quantization method for the neural network includes the following steps:

[0072] Step S100, obtain the network to be quantized and determine the evaluation layer under the network to be quantized;

[0073] It should be noted that the execution subject of the mixed-precision quantization method for the neural network is the mixed-precision quantization device for the neural network. Preferably, the mixed-precision quantization device for the neural network is a software client, or it can also be other terminals with data transmission and data processing functions, and specific limitations are not made here.

[0074] ​It is understandable that the network to be quantized refers to a machine learning model based on a neural network, including but not limited to object detection models, action recognition models, speech recognition models, etc. The number of parameters of the network to be quantized is usually large, and the parameters of each layer of the network to be quantized can be calculated using floating-point arithmetic (such as FP32 or FP16). Due to the huge number of parameters, the computational cost of performing operations using the network to be quantized is also huge. To obtain a model with a smaller volume, the method provided in the embodiments of the present application can be used to quantize the parameters (weights) of each layer of the neural network into different precisions. For example, convert the floating-point arithmetic of some network layers into low-bit fixed-point arithmetic, and convert the floating-point arithmetic of some network layers into high-bit fixed-point arithmetic, so that the neural network model can exchange a small precision loss for a large reduction in computational cost.

[0075] In a specific implementation, the network to be quantized includes multiple network layers, including but not limited to convolutional layers, pooling layers, normalization layers, activation functions, Flatten layers, fully connected layers, and output layers, etc. The evaluation layer is a network layer used to evaluate the precision error of the quantized model, usually the output layer. Specifically, the model outputs corresponding prediction results according to the training samples input by the user, and the evaluation layer judges the precision error of the prediction results. Among them, in the present application, a quantization-sensitive layer and a private layer are also proposed as the evaluation layer. The quantization-sensitive layer refers to a network layer with a large change in the quantized value, such as softmax, concat, and eltwise, etc.; the private layer refers to a network layer other than the public layer, and the public layer is a network layer with public convolution. When the private layer is used as the output, the device needs to select its input layer. Therefore, the determination of the network layer can be distinguished by the network layer attributes or specified according to the user's evaluation layer designation command.

[0076] Step S200: Determine the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the smallest precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized;

[0077] It is understandable that the low-precision quantization refers to the smallest precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized, where the low-precision quantization and the high-precision quantization are determined according to the chip configuration of the neural network model.

[0078] For example, the network to be quantized supports int8 and int16, where low-precision quantization is int8 and high-precision quantization is int16. It should be noted that int8 refers to quantization and calculation supporting 8 bits, and int16 refers to quantization and calculation supporting 16 bits.

[0079] In a specific implementation, the method for the device to determine the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization and the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization further includes the following steps:

[0080] Based on a preset input sample and the evaluation layer, perform floating-point inference on the network to be quantized to obtain a floating-point result; set each network layer in the network to be quantized as a low-precision network layer, and based on the input sample, perform low-precision quantization inference on the low-precision network layer to obtain low-precision quantization results of the evaluation layer for each low-precision network layer, and based on the floating-point result and the low-precision quantization results, calculate the first quantization error of the evaluation layer; set each network layer in the network to be quantized as a high-precision network layer layer by layer, and based on the input sample, perform high-precision quantization inference on the high-precision network layer layer by layer to obtain high-precision quantization results of the evaluation layer for each high-precision network layer, and based on the floating-point result and the high-precision quantization results, calculate the second quantization error of the evaluation layer.

[0081] It should be noted that the first quantization error refers to the calculation error caused before and after low-precision quantization of the model after the network to be quantized is converted from floating-point calculation to low-bit fixed-point calculation, where this calculation error is obtained based on the prediction result output by the evaluation layer. Similarly, the second quantization error refers to the calculation error caused before and after high-precision quantization of the model after the network to be quantized is converted from floating-point calculation to high-bit fixed-point calculation, where this calculation error is obtained based on the prediction result output by the evaluation layer.

[0082] For example, the user inputs sample X into the network to be quantized. The network to be quantized performs floating-point fp32 calculation, and the evaluation layer outputs prediction result Y1. The network to be quantized performs low-bit int8 fixed-point calculation, and the evaluation layer outputs prediction result Y2. The network to be quantized performs high-bit int16 fixed-point calculation, and the evaluation layer outputs prediction result Y3. The first quantization error is Y2 - Y1, and the second quantization error is Y3 - Y1.

[0083] In a specific implementation, the calculation error analysis is obtained through any of the following error analysis functions:

[0084]

[0085] CosineSimilarity

[0086] Among them, the first function is the mean square error calculation function, the second function is the relative error calculation function, and the third function is the cosine similarity calculation function.

[0087] It can be understood that in this application, all the networks to be quantized are set as low-precision quantization layers, while the high-precision quantization layers are set layer by layer. The concept of this design is that this application is based on the low-precision quantization network, and according to the data prediction performance of the neural network after being set as high-precision quantization, the network layers that meet the conditions are selected as high-precision quantization layers. The network layers in the network to be quantized except the high-precision quantization layers are all set as low-precision quantization layers, and the target mixed-precision network of the network to be quantized is composed of these.

[0088] Step S300: Based on the first quantization error and the second quantization error, determine the target high-precision network layer and the target low-precision network layer in the network to be quantized, and based on the target high-precision network layer and the target low-precision network layer, compose the target mixed-precision network of the network to be quantized.

[0089] It should be noted that the first quantization error represents the precision performance of the network to be quantized after being converted to low-bit fixed-point calculation, and the second quantization error represents the precision performance of the network to be quantized after being converted to high-bit fixed-point calculation. Under the strategy of combining precision and model volume, the device determines that some network layers in the network to be quantized are set as the target high-precision network layer and another part of the network layers are set as the target low-precision network layer according to the precision performance of low-precision quantization and the precision performance of high-precision quantization, and finally composes the target mixed-precision network of the network to be quantized.

[0090] In specific implementation, the method for the device to determine the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error further includes the following steps:

[0091] Based on the first quantization error and the second quantization error, calculate the target influence factor of the evaluation layer for each network layer; based on the target influence factor of the evaluation layer for each network layer, determine the influence factor threshold; based on the target influence factor of the evaluation layer for each network layer and the influence factor threshold, screen each network layer to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized.

[0092] It should be noted that the target influence factor refers to the quantization value of the influence on the prediction precision after the network to be quantized is converted from floating-point calculation to low-bit fixed-point calculation or high-bit fixed-point calculation. The larger the target influence factor, the higher the precision after setting this network layer as high-bit fixed-point calculation. On the contrary, the smaller the target influence factor, the lower the precision after setting this network layer as high-bit fixed-point calculation.

[0093] It is understandable that the impact factor threshold refers to the screening condition of the impact factor. The network layers that meet the screening conditions are set as high-precision quantization layers, and the remaining network layers are default set as low-precision quantization layers.

[0094] In a specific implementation, the method for the device to calculate the target impact factor of each network layer of the evaluation layer based on the first quantization error and the second quantization error further includes the following steps:

[0095] Obtain a mode strategy, where the mode strategy includes a performance mode strategy and a precision mode strategy. The performance mode strategy includes the following formula:

[0096]

[0097] The precision mode strategy includes the following formula:

[0098] impactFactor = errorln8Bit - errorlnMPLayer

[0099] where impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient; based on the mode strategy, determine the impact factor calculation algorithm; based on the impact factor calculation algorithm, calculate the impact factor for the first quantization error and the second quantization error to obtain the target impact factor of each network layer of the evaluation layer.

[0100] It should be noted that the mode strategy is a selection strategy specified according to user needs. User needs include quantization requirements with performance priority and quantization requirements with precision priority. The device formulates and implements different impact factor calculation algorithms according to different user needs. Among them, the impact factor calculation algorithm corresponding to the performance mode strategy includes the following algorithm formula:

[0101]

[0102] The precision mode strategy includes the following formula:

[0103] impactFactor = errorln8Bit - errorlnMPLayer

[0104] Among them, impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient.

[0105] It can be understood that the optimization scheme is to determine a suitable impact factor threshold, and screen the eligible high-precision layers according to the impact factor threshold. The selection of the impact factor threshold can be determined according to the mean value of the target impact factors of each network layer; it can also be determined according to the impact factor corresponding to the inflection point of the interval mean value of each network layer. By continuously calculating the mean value of N inflection points of the interval mean value of each network layer, the impact factor corresponding to the inflection point of the mean value is used as the threshold. Specifically, the determination scheme of the inflection point of the interval mean value includes the following steps:

[0106] Sort the target impact factors of each network layer of the evaluation layer in descending order, and extract a preset number of consecutive target impact factors from the sorted target impact factors to obtain multiple groups of consecutive target impact factors; calculate the mean value of each group of consecutive target impact factors to obtain the mean value of each group of target impact factors; determine the inflection point of the mean value of each group of target impact factors, and use the target impact factor corresponding to the inflection point of the mean value as the impact factor threshold.

[0107] It should be noted that the concept of selecting the mean value between target impact factors in this application is based on the fact that the value after calculating the mean value of target impact factors is more stable, so the obtained impact factor threshold is more accurate, making the accuracy of the mixed-precision quantization of the neural network higher.

[0108] In specific implementation, the method for the device to screen each network layer based on the target impact factors of each network layer of the evaluation layer and the impact factor threshold to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized further includes the following steps:

[0109] Compare the target impact factors of each network layer of the evaluation layer with the impact factor threshold to obtain a comparison result; use the network layer whose target impact factor is greater than or equal to the impact factor threshold in the comparison result as the target high-precision network layer, and use the network layer whose target impact factor is less than the impact factor threshold in the comparison result as the target low-precision network layer.

[0110] In this application, through the evaluation layer under the network to be quantized, the first quantization error of the network to be quantized with low-precision quantization and the second quantization error of the network to be quantized with high-precision quantization are calculated respectively. Based on the first quantization error and the second quantization error, with the low-precision quantization network as the basis, the high-precision network layers that meet the conditions under the network to be quantized are determined, so as to form the target mixed-precision network of the network to be quantized, without performing quantization strategy search for each layer of the network to be quantized, thereby improving the efficiency of mixed-precision quantization of the neural network.

[0111] Based on the above first embodiment, this application also provides another embodiment. Referring to Figure 2 , the method for mixed-precision quantization of the neural network includes:

[0112] Step A100: Obtain the network to be quantized and determine each network layer under the network to be quantized;

[0113] Step A200: Based on the attribute information of each network layer, determine the evaluation layer under the network to be quantized, where the evaluation layer is at least two network layers;

[0114] In a specific implementation, since the output layer is the network layer that outputs the results of the neural network model, the evaluation layer is usually the output layer. This application also proposes a method of jointly evaluating multiple network layers as the evaluation layer on the basis of the output layer, and uses the output layer, quantization-sensitive layer, and private layer under the network to be quantized as the evaluation layer. Specifically, the device distinguishes the evaluation layer through the network layer attributes.

[0115] Step A300: Based on the evaluation layer, perform multi-layer joint evaluation on the network to be quantized with low-precision quantization and the network to be quantized with high-precision quantization respectively, to obtain the first quantization error of each evaluation layer in the network to be quantized with low-precision quantization, and to obtain the second quantization error of each evaluation layer in the network to be quantized with high-precision quantization.

[0116] Step A400: Based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer, calculate the initial influence factor of each evaluation layer with respect to each network layer; perform joint calculation on the initial influence factor of each evaluation layer with respect to each network layer to obtain the target influence factor of the evaluation layer with respect to each network layer; based on the target influence factor of the evaluation layer with respect to each network layer, determine the target high-precision network layer and target low-precision network layer in the network to be quantized, and based on the target high-precision network layer and target low-precision network layer, form the target mixed-precision network of the network to be quantized.

[0117] In a specific implementation, since multiple network layers serve as evaluation layers, when calculating the first quantization error of each network layer in the network to be quantized with low-precision quantization or the second quantization error of each network layer in the network to be quantized with high-precision quantization, each evaluation layer will output the corresponding first quantization error and second quantization error, and after calculating based on the first quantization error and the second quantization error, each evaluation layer will output the corresponding initial influence factor. The present application proposes a multi-layer joint influence factor calculation formula. Specifically, by taking the mean value of the initial influence factors as the target influence factor, or taking the extreme value (maximum value or minimum value) of the initial influence factors as the target influence factor; or determining the target influence factor in a weighted manner. Specifically, different weights exist for the output layer and the sensitive layer, and the sensitive layer determines the weight according to the network depth.

[0118] Based on the above first embodiment and second embodiment, the present application also provides another embodiment. Refer to Figure 3 , the method for mixed-precision quantization of the neural network includes:

[0119] The device first determines appropriate evaluation layers (including the output layer, quantization sensitive layer, etc.) in the network to be quantized and a mixed-precision influence factor calculation scheme. Among them, the mixed-precision influence factor calculation scheme includes a performance mode and a precision mode. The performance mode focuses on the balance between performance and precision, and the precision mode only cares about the optimal precision.

[0120] The device executes a multi-layer joint influence factor calculation scheme and a high-precision layer selection scheme. Based on a preset input sample, it performs network floating-point fp32 inference and low-precision quantization (int8) inference on the network to be quantized, and calculates the low-precision (int8) relative to fp32 error errorint8bit and the time-consuming networktimefor8bit (MSA\MAE\cosine similarity, etc.).

[0121] Then the device sets high-precision (int16) layer by layer and sequentially obtains the error errorInMPlayer and the time-consuming networktimeformplayer at high precision for each layer.

[0122] The device further calculates and obtains the joint influence factor (maximum, average value, weighted average value) of each layer and multiple evaluation layers, sorts the joint influence factors in descending order, determines the influence factor threshold (average value or interval inflection point), screens the layers that meet the threshold to construct a high-precision layer candidate, and finally determines the mixed-precision network.

[0123] The present application also provides a device for mixed-precision quantization of a neural network. Refer to Figure 4 , the device for mixed-precision quantization of the neural network includes:

[0124] An acquisition module 10, configured to acquire a network to be quantized and determine an evaluation layer under the network to be quantized;

[0125] An error calculation module 20, configured to determine a first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine a second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized;

[0126] A determination module 30, configured to determine a target high-precision network layer and a target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error, and form a target mixed-precision network of the network to be quantized based on the target high-precision network layer and the target low-precision network layer.

[0127] And / or, the determination module 30 includes: an influence factor calculation module, configured to calculate a target influence factor of the evaluation layer with respect to each network layer based on the first quantization error and the second quantization error; a threshold determination module, configured to determine an influence factor threshold based on the target influence factor of the evaluation layer with respect to each network layer; a screening module, configured to screen each network layer based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold, and obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized;

[0128] And / or, the screening module includes: a comparison module, configured to compare the target influence factor of the evaluation layer with respect to each network layer with the influence factor threshold to obtain a comparison result; a network layer quantization module, configured to use the network layer with the target influence factor greater than or equal to the influence factor threshold in the comparison result as the target high-precision network layer, and use the network layer with the target influence factor less than the influence factor threshold in the comparison result as the target low-precision network layer;

[0129] And / or, the influence factor calculation module includes: a strategy acquisition module, configured to acquire a mode strategy, where the mode strategy includes a performance mode strategy and a precision mode strategy, and the performance mode strategy includes the following formula:

[0130]

[0131] The precision mode strategy includes the following formula:

[0132] impactFactor = errorln8Bit - errorlnMPLayer

[0133] Among them, impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient; the algorithm determination module is used to determine the impact factor calculation algorithm based on the mode strategy; the target impact factor calculation module is used to calculate the impact factor of the first quantization error and the second quantization error based on the impact factor calculation algorithm, and obtain the target impact factor of each network layer of the evaluation layer;

[0134] And / or, the threshold determination module includes: an extraction module, which is used to sort the target impact factors of each network layer of the evaluation layer in descending order, and extract a preset number of consecutive target impact factors from the sorted target impact factors to obtain multiple groups of consecutive target impact factors; a mean calculation module, which is used to calculate the mean of each group of consecutive target impact factors to obtain the mean of each group of target impact factors; an inflection point determination module, which is used to determine the mean inflection point in the mean of each group of target impact factors, and use the target impact factor corresponding to the mean inflection point as the impact factor threshold;

[0135] And / or, the acquisition module 10 further includes: a network layer determination module, which is used to determine each network layer under the network to be quantized; an evaluation layer determination module, which is used to determine the evaluation layer under the network to be quantized based on the attribute information of each network layer, where the evaluation layer includes an output layer, a quantization-sensitive layer, and a private layer; a multi-layer joint evaluation module, which is used to perform multi-layer joint evaluation on the network to be quantized with low-precision quantization and the network to be quantized with high-precision quantization based on the evaluation layer, and obtain the first quantization error of each evaluation layer in the network to be quantized with low-precision quantization, and obtain the second quantization error of each evaluation layer in the network to be quantized with high-precision quantization;

[0136] And / or, the determination module 30 further includes: an initial impact factor calculation module, which is used to calculate the initial impact factor of each network layer of each evaluation layer based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer; a joint calculation module, which is used to perform joint calculation on the initial impact factors of each network layer of each evaluation layer to obtain the target impact factor of each network layer of the evaluation layer; a quantization determination module, which is used to determine the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the target impact factor of each network layer of the evaluation layer;

[0137] And / or, the error calculation module 20 includes: a floating-point inference module, configured to perform floating-point inference on the network to be quantized based on a preset input sample and the evaluation layer, so as to obtain a floating-point result; a first error determination module, configured to set each network layer in the network to be quantized as a low-precision network layer, perform low-precision quantization inference on the low-precision network layer based on the input sample, so as to obtain a low-precision quantization result of the evaluation layer for each low-precision network layer, and calculate a first quantization error of the evaluation layer based on the floating-point result and the low-precision quantization result; a second error determination module, configured to set each network layer in the network to be quantized as a high-precision network layer layer by layer, perform high-precision quantization inference on the high-precision network layer based on the input sample, so as to obtain a high-precision quantization result of the evaluation layer for each high-precision network layer, and calculate a second quantization error of the evaluation layer based on the floating-point result and the high-precision quantization result.

[0138] The specific implementation manner of the hybrid precision quantization device of the neural network in this application is basically the same as each embodiment of the above-mentioned hybrid precision quantization method of the neural network, and will not be elaborated here.

[0139] Refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a terminal of the hardware operating environment involved in the solution of the embodiment of this application.

[0140] As Figure 5 shown, the terminal may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.

[0141] Optionally, the mixed-precision quantization device of the neural network may further include a rectangular user interface, a network interface, a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, and so on. The rectangular user interface may include a display screen and an input sub-module such as a keyboard. Optionally, the rectangular user interface may further include a standard wired interface and a wireless interface. The network interface may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0142] Those skilled in the art can understand that Figure 4 the structure of the mixed-precision quantization device of the neural network shown in does not constitute a limitation on the mixed-precision quantization device of the neural network, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0143] As Figure 5 shown, in the memory 1005 as a storage medium, there may be included an operating system, a network communication module, and a mixed-precision quantization program of the neural network. The operating system is a program for managing and controlling the hardware and software resources of the mixed-precision quantization device of the neural network, and supports the operation of the mixed-precision quantization program of the neural network and other software and / or programs. The network communication module is used to implement communication between various components inside the memory 1005, as well as communication between other hardware and software in the mixed-precision quantization system of the neural network.

[0144] In Figure 5 the mixed-precision quantization device of the neural network shown, the processor 1001 is used to execute the mixed-precision quantization program stored in the memory 1005 to implement the steps of the mixed-precision quantization method of the neural network described in any one of the above.

[0145] The specific implementation manner of the mixed-precision quantization device of the neural network in this application is basically the same as each embodiment of the above-mentioned mixed-precision quantization method of the neural network, and will not be elaborated here.

[0146] This application also provides a storage medium, on which a program for implementing the mixed-precision quantization method of the neural network is stored. The program for implementing the mixed-precision quantization method of the neural network is executed by a processor to implement the following mixed-precision quantization method of the neural network:

[0147] Obtain the network to be quantized, and determine the evaluation layer under the network to be quantized;

[0148] Determine the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized;

[0149] Based on the first quantization error and the second quantization error, determine the target high-precision network layer and the target low-precision network layer in the network to be quantized, and based on the target high-precision network layer and the target low-precision network layer, form the target mixed-precision network of the network to be quantized.

[0150] Optionally, the step of determining the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error includes:

[0151] Based on the first quantization error and the second quantization error, calculate the target influence factor of the evaluation layer with respect to each network layer;

[0152] Based on the target influence factor of the evaluation layer with respect to each network layer, determine the influence factor threshold;

[0153] Based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold, screen each network layer to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized.

[0154] Optionally, the step of screening each network layer based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized includes:

[0155] Compare the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold to obtain a comparison result;

[0156] Use the network layer with the target influence factor greater than or equal to the influence factor threshold in the comparison result as the target high-precision network layer, and use the network layer with the target influence factor less than the influence factor threshold in the comparison result as the target low-precision network layer.

[0157] Optionally, the step of calculating the target influence factor of the evaluation layer with respect to each network layer based on the first quantization error and the second quantization error includes:

[0158] Obtaining mode strategies, where the mode strategies include a performance mode strategy and a precision mode strategy, and the performance mode strategy includes the following formula:

[0159]

[0160] The precision mode strategy includes the following formula:

[0161] impactFactor = errorln8Bit - errorlnMPLayer

[0162] where impactFactor represents the impact factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient;

[0163] Based on the mode strategy, determine the impact factor calculation algorithm;

[0164] Based on the impact factor calculation algorithm, calculate the impact factor for the first quantization error and the second quantization error to obtain the target impact factor of the evaluation layer for each network layer.

[0165] Optionally, the step of determining the impact factor threshold based on the target impact factor of the evaluation layer for each network layer includes:

[0166] Sort the target impact factors of the evaluation layer for each network layer in descending order, and extract a preset number of consecutive target impact factors from the sorted target impact factors to obtain multiple groups of consecutive target impact factors;

[0167] Calculate the mean of each group of consecutive target impact factors to obtain the mean of each group of target impact factors;

[0168] Determine the mean inflection point among the means of each group of target impact factors, and use the target impact factor corresponding to the mean inflection point as the impact factor threshold.

[0169] Optionally, the step of determining the evaluation layer under the network to be quantized includes:

[0170] Determine each network layer under the network to be quantized;

[0171] Based on the attribute information of each network layer, determine the evaluation layer under the network to be quantized, where the evaluation layer includes an output layer, a quantization-sensitive layer, and a private layer;

[0172] The steps of determining the first quantization error of the output of the evaluation layer in the quantization network to be quantized with low precision and determining the second quantization error of the output of the evaluation layer in the quantization network to be quantized with high precision include:

[0173] Based on the evaluation layer, perform multi-layer joint evaluation on the quantization network to be quantized with low precision and the quantization network to be quantized with high precision respectively, to obtain the first quantization error of each evaluation layer in the quantization network to be quantized with low precision, and to obtain the second quantization error of each evaluation layer in the quantization network to be quantized with high precision.

[0174] Optionally, the steps of determining the target high-precision network layer and the target low-precision network layer in the quantization network based on the first quantization error and the second quantization error include:

[0175] Based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer, calculate the initial influence factor of each evaluation layer with respect to each network layer;

[0176] Perform joint calculation on the initial influence factor of each evaluation layer with respect to each network layer to obtain the target influence factor of each evaluation layer with respect to each network layer;

[0177] Based on the target influence factor of each evaluation layer with respect to each network layer, determine the target high-precision network layer and the target low-precision network layer in the quantization network.

[0178] Optionally, the steps of determining the first quantization error of the output of the evaluation layer in the quantization network to be quantized with low precision and determining the second quantization error of the output of the evaluation layer in the quantization network to be quantized with high precision include:

[0179] Based on a preset input sample and the evaluation layer, perform floating-point inference on the quantization network to be quantized to obtain a floating-point result;

[0180] Set each network layer in the quantization network to a low-precision network layer, and based on the input sample, perform low-precision quantization inference on the low-precision network layer to obtain the low-precision quantization result of the evaluation layer with respect to each low-precision network layer, and based on the floating-point result and the low-precision quantization result, calculate the first quantization error of the evaluation layer;

[0181] Set each network layer in the quantization network to a high-precision network layer layer by layer, and based on the input sample, perform high-precision quantization inference on the high-precision network layer layer by layer to obtain the high-precision quantization result of the evaluation layer with respect to each high-precision network layer, and based on the floating-point result and the high-precision quantization result, calculate the second quantization error of the evaluation layer.

[0182] The specific implementation of the storage medium of this application is basically the same as that of each embodiment of the above-mentioned mixed-precision quantization method for neural networks, and will not be elaborated here.

[0183] This application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned mixed-precision quantization method for neural networks when executed by a processor.

[0184] The specific implementation of the computer program product of this application is basically the same as that of each embodiment of the above-mentioned mixed-precision quantization method for neural networks, and will not be elaborated here.

[0185] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0186] The serial numbers of the above-mentioned embodiments of this application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0187] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.

[0188] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.

Claims

1. A mixed-precision quantization method for a neural network, characterized in that, The method for mixed-precision quantization of the neural network includes: Obtain the network to be quantized and determine the evaluation layer under the network to be quantized; Determine the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization, and determine the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized; Based on the first quantization error and the second quantization error, determine the target high-precision network layer and the target low-precision network layer in the network to be quantized, and based on the target high-precision network layer and the target low-precision network layer, form the target mixed-precision network of the network to be quantized.

2. The mixed-precision quantization method for a neural network according to claim 1, wherein, The step of determining the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error includes: Based on the first quantization error and the second quantization error, calculate the target influence factor of the evaluation layer with respect to each network layer; Based on the target influence factor of the evaluation layer with respect to each network layer, determine the influence factor threshold; Based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold, screen each network layer to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized.

3. The mixed-precision quantization method of the neural network according to claim 2, wherein The step of screening each network layer based on the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold to obtain the target high-precision network layer and the target low-precision network layer in the network to be quantized includes: Compare the target influence factor of the evaluation layer with respect to each network layer and the influence factor threshold to obtain a comparison result; Use the network layer with the target influence factor greater than or equal to the influence factor threshold in the comparison result as the target high-precision network layer, and use the network layer with the target influence factor less than the influence factor threshold in the comparison result as the target low-precision network layer.

4. The mixed-precision quantization method for a neural network according to claim 2, wherein The step of calculating the target influence factor of the evaluation layer with respect to each network layer based on the first quantization error and the second quantization error includes: Obtain the mode strategy, where the mode strategy includes a performance mode strategy and a precision mode strategy, and the performance mode strategy includes the following formula: The precision mode strategy includes the following formula: impactFactor = errorln8Bit - errorlnMPLayer where impactFactor represents the influence factor, errorln8Bit represents the first quantization error, errorlnMPLayer represents the second quantization error, networkTimeForMPLayer represents the running time of the high-precision quantization layer, networkTimeFor8Bit represents the running time of the low-precision quantization layer, and eps represents the small value coefficient; Based on the mode strategy, determine the influence factor calculation algorithm; Based on the influence factor calculation algorithm, calculate the influence factors of the first quantization error and the second quantization error to obtain the target influence factors of each network layer in the evaluation layer.

5. The mixed-precision quantization method for a neural network according to claim 2, wherein The step of determining the influence factor threshold based on the target influence factors of each network layer in the evaluation layer includes: Sort the target influence factors of each network layer in the evaluation layer in descending order, and extract a preset number of consecutive target influence factors from the sorted target influence factors to obtain multiple groups of consecutive target influence factors. Calculate the mean value of each group of consecutive target influence factors to obtain the mean value of each group of target influence factors. Determine the mean inflection point among the mean values of each group of target influence factors, and use the target influence factor corresponding to the mean inflection point as the influence factor threshold.

6. The mixed-precision quantization method of the neural network according to claim 1, wherein, The step of determining the evaluation layer under the network to be quantized includes: Determine each network layer under the network to be quantized. Based on the attribute information of each network layer, determine the evaluation layer under the network to be quantized, where the evaluation layer includes an output layer, a quantization-sensitive layer, and a private layer. The steps of determining the first quantization error output by the evaluation layer in the network to be quantized with low-precision quantization and determining the second quantization error output by the evaluation layer in the network to be quantized with high-precision quantization include: Based on the evaluation layer, perform multi-layer joint evaluation on the network to be quantized with low-precision quantization and the network to be quantized with high-precision quantization respectively, to obtain the first quantization error of each evaluation layer in the network to be quantized with low-precision quantization and the second quantization error of each evaluation layer in the network to be quantized with high-precision quantization.

7. The mixed-precision quantization method of the neural network according to claim 6, characterized in that, The steps of determining the target high-precision network layer and the target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error include: Based on the first quantization error of each evaluation layer and the second quantization error of each evaluation layer, calculate the initial influence factor of each evaluation layer for each network layer. Perform joint calculation on the initial influence factors of each evaluation layer for each network layer to obtain the target influence factors of each network layer in the evaluation layer. Based on the target influence factors of each network layer in the evaluation layer, determine the target high-precision network layer and the target low-precision network layer in the network to be quantized.

8. A hybrid-precision quantization device for a neural network, characterized in that, The hybrid precision quantization device of the neural network includes: An acquisition module, configured to acquire a network to be quantized and determine the evaluation layer under the network to be quantized. An error calculation module, configured to determine the first quantization error of the evaluation layer in the network to be quantized with low-precision quantization and determine the second quantization error of the evaluation layer in the network to be quantized with high-precision quantization, where the low-precision quantization refers to the minimum precision quantization value among the precision quantization values supported by the network to be quantized, and the high-precision quantization refers to the precision quantization value higher than the low-precision quantization among the precision quantization values supported by the network to be quantized. A determination module, configured to determine a target high-precision network layer and a target low-precision network layer in the network to be quantized based on the first quantization error and the second quantization error, and form a target mixed-precision network of the network to be quantized based on the target high-precision network layer and the target low-precision network layer.

9. A hybrid-precision quantization device for a neural network, characterized in that, The mixed-precision quantization device of the neural network includes: a memory, a processor, and a program stored on the memory for implementing the mixed-precision quantization method of the neural network. The memory is used to store a program for implementing the mixed-precision quantization method of the neural network. The processor is configured to execute the program for implementing the mixed-precision quantization method of the neural network to implement the steps of the mixed-precision quantization method of the neural network according to any one of claims 1 to 7.

10. A storage medium, characterized in that, A program for implementing the mixed-precision quantization method of the neural network is stored on the storage medium, and the program for implementing the mixed-precision quantization method of the neural network is executed by the processor to implement the steps of the mixed-precision quantization method of the neural network according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method for determining mixed quantization configuration of neural network model executed on heterogeneous computing platform

    RU2870036C1