Neural network algorithm processor design method for mixing precision quantization

By designing a neural network algorithm processor that supports mixed precision calculation, using a hybrid precision calculation module and a quantized scheduling module, the accuracy scheduling of neural networks at different levels in the inference process is realized, and the model accuracy reduction caused by fixed precision calculation in the existing technology is solved, and the calculation efficiency and accuracy are improved.

CN119990207APending Publication Date: 2025-05-13西安翔腾微电子科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411749973.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing neural network processors usually only support fixed accuracy calculations, and cannot balance model calculation accuracy with hardware calculation efficiency in neural network inference tasks, resulting in a significant decrease in model accuracy.

Method used

Design a neural network algorithm processor for hybrid precision quantization. Through the hybrid precision calculation module and the quantization scheduling module, the neural network is allowed to use different calculation accuracy during the inference process, the critical layer uses high-precision calculation, and the non-critical layer uses low-precision calculation.

Benefits of technology

By reducing unnecessary high-precision calculations, the computational complexity in the neural network inference process is reduced, the computing efficiency is improved, and the overall computational accuracy of the neural network is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990207A_ABST
    Figure CN119990207A_ABST
Patent Text Reader

Abstract

The invention relates to a neural network algorithm processor design method for mixing precision quantization. The method comprises the following steps: 1) performing mixing precision quantification on a neural network model according to a calibration data set; 2) a quantitative scheduling module maps each layer of the neural network with the corresponding calculation precision according to the quantitative configuration table; and 3) the mixing precision calculation module comprises a precision conversion unit and a calculation unit, the precision conversion unit in the mixing precision calculation module performs precision conversion on the input data to ensure the precision consistency of the input data, and after the precision conversion is completed, the corresponding MAC calculation unit is called to perform multiplication and addition calculation. According to the method, unnecessary high-precision calculation is reduced, the calculation complexity in the neural network reasoning process is reduced, the calculation efficiency is improved, meanwhile, a hierarchical precision distribution method is adopted, high-precision calculation is used for a key layer, the overall calculation precision of the neural network is kept, and the calculation efficiency is improved. The method can be applied to scenes with strict requirements on the reasoning precision and performance of the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a neural network algorithm processor design method that supports mixed-precision quantized inference calculations of neural networks. Background Art

[0002] With the rapid development of deep learning, neural networks are widely used in computer vision, natural language processing, speech recognition and other fields. Neural networks usually use 32-bit floating-point numbers to represent their weights and activation values. While ensuring high precision, they bring huge computing and storage overheads, and cannot be directly deployed on resource-constrained edge embedded devices. Model quantization methods are usually required to convert 32-bit floating-point calculations into 16-bit floating-point calculations or INT8 integer calculations, reducing the amount of calculation while reducing its storage requirements. Converting 32-bit floating-point data to 16-bit floating-point data generally has less loss and can be directly converted. Therefore, the present invention mainly considers quantizing 32-bit floating-point numbers into INT8 integers.

[0003] However, quantizing all layers of the neural network from 32-bit floating point calculations to 8-bit integer calculations will significantly reduce the model accuracy. Mixed precision allows different layers of the neural network to use different precision calculations, which can not only retain the high-precision calculations of key layers, but also reduce the computational overhead of non-key layers through low-precision calculations. However, existing neural network processors usually only support fixed-precision calculations, that is, all layers of the neural network use the same calculation precision (FP16 or INT8 calculation precision), which cannot balance the model calculation accuracy and hardware calculation efficiency when performing neural network inference tasks. Summary of the invention

[0004] In order to solve the above-mentioned technical problems existing in the background technology, the present invention provides a neural network algorithm processor design method for mixed precision quantization, which supports the use of different calculation precisions by the neural network during the reasoning process. The key layers of the model use high-precision calculations, and the non-key layers use low-precision calculations. While improving the model reasoning performance, its calculation accuracy is greatly improved.

[0005] The technical solution of the present invention is: the present invention is a design method of a neural network algorithm processor for mixed precision quantization, and its special feature is that the method comprises the following steps:

[0006] 1) Perform mixed precision quantization on the neural network model based on the calibration dataset;

[0007] 2) The quantization scheduling module maps each layer of the neural network to its corresponding calculation accuracy according to the quantization configuration table;

[0008] 3) The mixed precision computing module includes a precision conversion unit and a computing unit. The precision conversion unit in the mixed precision computing module performs precision conversion on the input data to ensure the precision consistency of the input data. After the precision conversion is completed, the corresponding computing unit is called to perform multiplication and addition calculations.

[0009] Further, the specific steps of step 1) are as follows:

[0010] 1.1) Quantize 32-bit floating-point data to INT8 integer data. The data conversion between FP32 floating-point data and INT8 integer data is achieved by the following formula:

[0011] Value fp32 =Scale fp32 *Value int8

[0012] Scale fp32 is a quantization parameter, in order to determine the effective Scale fp32 The relative entropy is usually used to describe the difference between the probability distribution of FP32 data and INT8 data. The smaller the relative entropy, the smaller the difference between the two probability distributions, and the closer the shape and value of the probability density function are. The relative entropy search method is used to determine the appropriate quantization parameter Scale fp32 , to minimize the difference between FP32 data distribution and INT8 data distribution, thereby reducing the accuracy error of INT8 quantization calculation;

[0013] 1.2) During the quantization process, the neural network model is analyzed layer by layer, and the activation distribution, weight distribution, gradient information of each layer and its impact on the final result are comprehensively considered, so as to select different quantization precisions for each layer and generate a quantization configuration table.

[0014] Furthermore, the specific steps of step 2) are: the quantization scheduling module maps each layer of the neural network to its corresponding computing accuracy according to the quantization configuration table, and schedules the corresponding computing unit according to the accuracy requirement of each layer.

[0015] Further, the specific steps of step 3) are as follows:

[0016] 3.1) The mixed precision computing module is specially designed for the mixed precision quantized inference computing task of neural network. Its two inputs are the feature data and parameter data in the neural network calculation. The feature data refers to the data generated in real time in the neural network calculation, and the parameter data refers to the weight data generated offline during the neural network training process.

[0017] 3.2) The precision conversion between FP32 data format and INT8 data format can be completed offline for parameter data, that is, it does not need to be completed by the precision conversion unit in the mixed precision computing module, while the precision conversion between FP32 data format and INT8 data format needs to be completed online for feature data according to actual needs.

[0018] Furthermore, in step 3.2), the precision conversion is completed by the precision conversion unit in the mixed precision computing module, and the type of precision conversion and the scale quantization parameter are determined by the quantization configuration table.

[0019] Furthermore, in step 3.2), when both input data of the computing unit of the mixed precision computing module adopt FP16 precision, we say that the mixed precision computing module is in FP16 working mode, and the output data of the computing unit is also expressed in FP16 floating point; when both input data of the computing unit adopt INT8 integer expression, we say that the mixed precision operator is in INT8 working mode, and the output data of the computing unit adopts INT8 integer expression. The precision conversion steps are as follows:

[0020] 3.2.1) Before calculation, the accuracy consistency of the two input data needs to be ensured. The precision conversion unit supports conversion from low precision to high precision, that is, converting INT8 data to FP16 floating point data to ensure high-precision calculation. The conversion formula is:

[0021] Value fp16 =Scale fp16 *Value int8

[0022] 3.2.2) The precision conversion unit also supports conversion from high precision to low precision, that is, converting FP16 floating point data to INT8 integer data, thereby reducing the computational overhead. The conversion formula is:

[0023] Value int8 =Clamp(Round(Value fp16 / Scale fp16 ),-128,127)

[0024] Among them, Round is the rounding function and Clamp is the truncation function, which can limit the rounded data to [-128, 127]. After the precision conversion is completed, the input data precision remains consistent, and the corresponding MAC calculation unit is called to perform multiplication and addition calculations to obtain the final calculation result.

[0025] The neural network algorithm processor design method for mixed precision quantization provided by the present invention reduces the computational complexity in the neural network reasoning process and improves the computational efficiency by reducing unnecessary high-precision calculations. At the same time, a layered precision allocation method is adopted, and high-precision calculations are used in the key layers to maintain the overall computational accuracy of the neural network. It can be applied to scenarios that have strict requirements on the neural network reasoning accuracy and performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is intended to represent the quantitative configuration of the present invention;

[0027] Figure 2 A schematic diagram of a mixed-precision computing embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The method steps of the specific embodiment of the present invention are as follows:

[0030] 1) Perform mixed precision quantization on the neural network model based on the calibration dataset;

[0031] 1.1) Quantize 32-bit floating-point data to INT8 integer data. The data conversion between FP32 floating-point data and INT8 integer data is achieved by the following formula:

[0032] Value fp32 =Scale fp32 *Value int8

[0033] Scale fp32 is a quantization parameter, in order to determine the effective Scale fp32 The relative entropy is usually used to describe the difference between the probability distribution of FP32 data and INT8 data. The smaller the relative entropy, the smaller the difference between the two probability distributions, and the closer the shape and value of the probability density function are. The relative entropy search method is used to determine the appropriate quantization parameter Scale fp32 , to minimize the difference between FP32 data distribution and INT8 data distribution, thereby reducing the accuracy error of INT8 quantization calculation;

[0034] 1.2) During the quantization process, the neural network model is analyzed layer by layer, and the activation distribution, weight distribution, gradient information and impact on the final result of each layer are comprehensively considered, so as to select different quantization precisions for each layer and generate a quantization configuration table, such as Figure 1 shown.

[0035] 2) The quantization scheduling module maps each layer of the neural network with its corresponding computing accuracy according to the quantization configuration table; specifically: the quantization scheduling module maps each layer of the neural network with its corresponding computing accuracy according to the quantization configuration table, and schedules the corresponding computing unit according to the accuracy requirements of each layer.

[0036] 3) The mixed precision computing module includes a precision conversion unit and a computing unit. The precision conversion unit in the mixed precision computing module performs precision conversion on the input data to ensure the precision consistency of the input data. After the precision conversion is completed, the corresponding computing unit is called to perform multiplication and addition calculations.

[0037] 3.1) The mixed precision computing module is specially designed for the mixed precision quantized inference computing task of neural network. Its two inputs are the feature data and parameter data in the neural network calculation. The feature data refers to the data generated in real time in the neural network calculation, and the parameter data refers to the weight data generated offline during the neural network training process.

[0038] 3.2) The precision conversion between FP32 data format and INT8 data format can be completed offline for parameter data, that is, it does not need to be completed by the precision conversion unit in the mixed precision computing module, while the precision conversion between FP32 data format and INT8 data format needs to be completed online for feature data according to actual needs.

[0039] The precision conversion is completed by the precision conversion unit in the mixed precision calculation module. The type of precision conversion and the scale quantization parameter are determined by the quantization configuration table, as follows:

[0040] When both input data of the computing unit of the mixed-precision computing module use FP16 precision, we say that the mixed-precision computing module is in FP16 working mode, and the output data of the computing unit is also expressed in FP16 floating point; when both input data of the computing unit are expressed in INT8 integer, we say that the mixed-precision operator is in INT8 working mode, and the output data of the computing unit is expressed in INT8 integer. The precision conversion steps are as follows:

[0041] 3.2.1) Before calculation, the accuracy consistency of the two input data needs to be ensured. The precision conversion unit supports conversion from low precision to high precision, that is, converting INT8 data to FP16 floating point data to ensure high-precision calculation. The conversion formula is:

[0042] Value fp16 =Scale fp16 *Value int8

[0043] 3.2.2) The precision conversion unit also supports conversion from high precision to low precision, that is, converting FP16 floating point data to INT8 integer data, thereby reducing the computational overhead. The conversion formula is:

[0044] Value int8 =Clamp(Round(Value fp16 / Scale fp16 ),-128,127)

[0045] Among them, Round is the rounding function and Clamp is the truncation function, which can limit the rounded data to [-128, 127]. After the precision conversion is completed, the input data precision remains consistent, and the corresponding MAC calculation unit is called to perform multiplication and addition calculations to obtain the final calculation result.

[0046] See also Figure 2 In actual neural network calculations, two consecutive mixed-precision operators can be configured as different precision calculation modes as needed, so as to achieve the effect of using INT8 to accelerate some operations in neural network calculations while ensuring higher precision through FP16. This embodiment consists of three mixed-precision operators, which are in INT8 mode (the two input data of the calculation unit are INT8 integers, and the output is INT8 integer), FP16 mode (the two inputs of the calculation unit are FP16 floating points, and the output is FP16 floating points) and INT8 mode (the two input data of the calculation unit are INT8 integers, and the output is INT8 integer).

[0047] In this embodiment, the mixed precision operator A is in INT8 mode, and its two inputs are FP16 floating-point feature data and INT8 integer parameter data respectively. Inside the mixed precision operator A, the FP16 floating-point feature data will be converted into INT8 integer data by the precision conversion module. After the conversion is completed, the INT8 integer data and the parameter data are multiplied and added. The output result of the mixed precision operator A is used as the input of the mixed precision operator B. The mixed precision operator B is in FP16 mode, and its two inputs are FP16 floating-point parameter data and INT8 integer feature data respectively. Internally, the INT8 integer feature data is converted into FP16 floating-point data by the precision conversion module, and then the multiplication and addition operation of the two inputs is completed through the computing unit. Finally, the output result of the mixed precision operator B is used as the input of the mixed precision operator C. The mixed precision operator C is in int8 mode, and its two inputs are FP16 floating-point feature data and INT8 integer parameter data. Inside the mixed precision operator C, the FP16 floating-point feature data will be converted into INT8 integer data by the precision conversion module. After the conversion is completed, the INT8 integer data and the parameter data will be multiplied and added to obtain the final output result.

[0048] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

[0049] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.

[0050] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.

Claims

1. A method for designing a neural network algorithm processor for mixed precision quantization, characterized in that: The method comprises the following steps: 1) Perform mixed precision quantization on the neural network model based on the calibration dataset; 2) The quantization scheduling module maps each layer of the neural network to its corresponding calculation accuracy according to the quantization configuration table; 3) The mixed precision computing module includes a precision conversion unit and a computing unit. The precision conversion unit in the mixed precision computing module performs precision conversion on the input data to ensure the precision consistency of the input data. After the precision conversion is completed, the corresponding computing unit is called to perform multiplication and addition calculations.

2. The method for designing a neural network algorithm processor for mixed precision quantization according to claim 1, characterized in that: The specific steps of step 1) are as follows: 1.1) Quantize 32-bit floating-point data to INT8 integer data. The data conversion between FP32 floating-point data and INT8 integer data is achieved by the following formula: Value fp32 =Scale fp32 *Value int8 Scale fp32 is a quantization parameter, in order to determine the effective Scale fp32 The relative entropy is usually used to describe the difference between the probability distribution of FP32 data and INT8 data. The smaller the relative entropy, the smaller the difference between the two probability distributions, and the closer the shape and value of the probability density function are. The relative entropy search method is used to determine the appropriate quantization parameter Scale fp32 , to minimize the difference between FP32 data distribution and INT8 data distribution, thereby reducing the accuracy error of INT8 quantization calculation; 1.2) During the quantization process, the neural network model is analyzed layer by layer, and the activation distribution, weight distribution, gradient information of each layer and its impact on the final result are comprehensively considered, so as to select different quantization precisions for each layer and generate a quantization configuration table.

3. The method for designing a neural network algorithm processor for mixed precision quantization according to claim 2, characterized in that: The specific steps of step 2) are: the quantization scheduling module maps each layer of the neural network to its corresponding calculation accuracy according to the quantization configuration table, and schedules the corresponding calculation unit according to the accuracy requirement of each layer.

4. The method for designing a neural network algorithm processor for mixed precision quantization according to claim 3, characterized in that: The specific steps of step 3) are as follows: 3.1) The mixed precision computing module is specially designed for the mixed precision quantized inference computing task of neural network. Its two inputs are the feature data and parameter data in the neural network calculation. The feature data refers to the data generated in real time in the neural network calculation, and the parameter data refers to the weight data generated offline during the neural network training process. 3.2) The precision conversion between FP32 data format and INT8 data format can be completed offline for parameter data, that is, it does not need to be completed by the precision conversion unit in the mixed precision computing module, while the precision conversion between FP32 data format and INT8 data format needs to be completed online for feature data according to actual needs.

5. The method for designing a neural network algorithm processor for mixed precision quantization according to claim 4, characterized in that: The precision conversion in step 3.2) is completed by the precision conversion unit in the mixed precision calculation module, and the type of precision conversion and the scale quantization parameter are determined by the quantization configuration table.

6. The method for designing a neural network algorithm processor for mixed precision quantization according to claim 5, characterized in that: In the step 3.2), when both input data of the computing unit of the mixed precision computing module adopt FP16 precision, the mixed precision computing module is in FP16 working mode, and the output data of the computing unit is also expressed in FP16 floating point; when both input data of the computing unit adopt INT8 integer expression, the mixed precision operator is in INT8 working mode, and the output data of the computing unit adopts INT8 integer expression. The precision conversion steps are as follows: 3.2.1) Before calculation, the accuracy consistency of the two input data needs to be ensured. The precision conversion unit supports conversion from low precision to high precision, that is, converting INT8 data to FP16 floating point data to ensure high-precision calculation. The conversion formula is: Value fp16 =Scale fp16 *Value int8 3.2.2) The precision conversion unit also supports conversion from high precision to low precision, that is, converting FP16 floating point data to INT8 integer data, thereby reducing the computational overhead. The conversion formula is: Value int8 =Clamp(Round(Value fp16 / Scale fp16 ),-128,127) Among them, Round is the rounding function and Clamp is the truncation function, which can limit the rounded data to [-128, 127]. After the precision conversion is completed, the input data precision remains consistent, and the corresponding MAC calculation unit is called to perform multiplication and addition calculations to obtain the final calculation result.