Neural network quantization method, device, equipment, storage medium and program product

By converting the quantization parameters of the neural network from a format suitable for CPU to a format suitable for other processing units (such as DSP), the problem that the existing technology can only perform inference on the CPU is solved, and the adaptation of the quantization parameters of the neural network among multiple processing units is realized, which improves the flexibility and efficiency of calculation.

CN114444688BActive Publication Date: 2025-05-13BIGO TECH PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210044661.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-13
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

The existing neural network quantization methods can only perform inference on the CPU and cannot adapt to the needs of multiple computing scenarios.

Method used

By obtaining the quantization parameters of multiple functional layers of the neural network and converting them from quantization parameters suitable for CPU to quantization parameters suitable for other processing units (such as DSP), the quantization parameters of the neural network can be adapted between different types of processing units.

Benefits of technology

The adaptation of neural network quantization parameters among multiple different types of processing units is achieved, which meets the computing needs of compatible with multiple processing units, and improves the applicability and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444688B_ABST
    Figure CN114444688B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a quantization method, device, equipment, storage medium and program product of a neural network, which relates to the field of artificial intelligence technology. The method includes: obtaining a first quantization parameter of multiple functional layers of a neural network, the first quantization parameter is used to quantize data in a floating point format into data in a first fixed point format, and the first fixed point format is a format adopted by a first processing unit; for a first functional layer among multiple functional layers, converting the first quantization parameter of the first functional layer into a second quantization parameter, the second quantization parameter is used to quantize data in a floating point format into data in a second fixed point format, and the second fixed point format is a format adopted by a second processing unit, and the first processing unit and the second processing unit are two different processing units; when the second processing unit is used to infer the first functional layer, quantizing the relevant data involved in the inference process of the first functional layer based on the second quantization parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a neural network quantization method, device, equipment, storage medium and program product. Background Art

[0002] Quantization of neural networks is a way to compress the model, which helps to reduce the model size and speed up the inference.

[0003] The general CPU (Central Processing Unit) supports fixed-point calculations with a minimum of 8 bits. When the CPU is used as the computing backend of a neural network, the parameters of the neural network are generally quantized to 8 bits.

[0004] However, this quantization method only supports neural network reasoning on the CPU and cannot adapt to the needs of more computing scenarios. Summary of the invention

[0005] The embodiment of the present application provides a neural network quantization method, device, equipment, storage medium and program product. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, a method for quantizing a neural network is provided, the method comprising:

[0007] Acquire first quantization parameters of multiple functional layers of the neural network, where the first quantization parameters are used to quantize data in a floating point format into data in a first fixed point format, where the first fixed point format is a format used by the first processing unit;

[0008] For a first functional layer among the multiple functional layers, convert a first quantization parameter of the first functional layer into a second quantization parameter, where the second quantization parameter is used to quantize the data in the floating point format into data in a second fixed point format, where the second fixed point format is a format adopted by a second processing unit, and the first processing unit and the second processing unit are two different processing units;

[0009] In case that the second processing unit is used to infer the first functional layer, relevant data involved in the inference process of the first functional layer is quantized based on the second quantization parameter.

[0010] According to one aspect of an embodiment of the present application, a quantization device for a neural network is provided, the device comprising:

[0011] A parameter acquisition module, used to acquire first quantization parameters of multiple functional layers of the neural network, wherein the first quantization parameters are used to quantize data in a floating point format into data in a first fixed point format, wherein the first fixed point format is a format adopted by the first processing unit;

[0012] a parameter conversion module, configured to convert, for a first functional layer among the multiple functional layers, a first quantization parameter of the first functional layer into a second quantization parameter, wherein the second quantization parameter is used to quantize the data in the floating point format into data in a second fixed point format, wherein the second fixed point format is a format adopted by a second processing unit, and the first processing unit and the second processing unit are two different processing units;

[0013] A data quantization module is used to quantize relevant data involved in the reasoning process of the first functional layer based on the second quantization parameter when the second processing unit is used to infer the first functional layer.

[0014] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned neural network quantization method.

[0015] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned neural network quantization method.

[0016] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes computer instructions stored in a computer-readable storage medium, and a processor reads and executes the computer instructions from the computer-readable storage medium to implement the above-mentioned neural network quantization method.

[0017] The technical solution provided in the embodiments of the present application can bring the following beneficial effects:

[0018] By converting the quantization parameters of the neural network from quantization parameters suitable for calculations by a first processing unit (such as a CPU) to quantization parameters suitable for calculations by a second processing unit (such as a DSP), the quantization parameters of the neural network can be adapted between a variety of different types of processing units, such as between the CPU and the DSP, to meet the computing requirements of being compatible with a variety of different types of processing units. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;

[0020] Figure 2 is a flow chart of a neural network quantization method provided by an embodiment of the present application;

[0021] Figure 3 is a flow chart of the reasoning process of a neural network provided by one embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of converting the uint8 format to the int8 format provided by an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of converting the uint8 format to the int8 format provided by another embodiment of the present application;

[0024] Figure 6 is a block diagram of a quantization device for a neural network provided by an embodiment of the present application;

[0025] Figure 7 It is a block diagram of a quantization device for a neural network provided in another embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0027] Before introducing the embodiments of the present application, some technical terms involved in the present application are defined and explained.

[0028] 1. Neural network: It is an algorithmic mathematical model that imitates the behavioral characteristics of animal neural networks and performs distributed parallel information processing. Neural networks rely on the complexity of the system to adjust the interconnected relationships between a large number of internal nodes to achieve the purpose of processing information. Neural networks can include multiple functional layers. Optionally, the functional layers include but are not limited to convolution layers, quasi-convolution layers (such as deconvolution layers, fully connected layers, etc.), pooling layers, concatenation layers, eltwise layers, binary layers, scale layers, activation (relu) layers, etc.

[0029] 2. Quantization: Quantization is the process of approximating continuous values ​​(or a large number of possible discrete values) to a finite number of (or fewer) discrete values. Quantization of neural networks is a way of model compression, which is to approximate the weight value (weight) or activation value (activation) represented by a high bit width (such as float32, 32-bit floating point number) with a lower bit width (such as int8, 8-bit signed integer number). The numerical manifestation is to discretize the continuous value. Dequantization is the inverse process of quantization, that is, the process of converting fixed-point numbers to floating-point numbers.

[0030] 3. Processing unit: Also known as a processor, it refers to the hardware unit in a computer device that is responsible for data calculation. Optionally, the processing unit includes but is not limited to a CPU, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an NPU (Neural-network Processing Unit), etc.

[0031] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment may include: an offline computing device 10 and an online computing device 20.

[0032] The quantization of neural networks can be divided into an offline stage and an online stage. The offline stage refers to the quantization of some network parameters (such as weight values ​​and other parameters) obtained from the training after the training of the neural network is completed. This process can be performed offline, that is, there is no need to consider the data actually used in the online reasoning stage. The online stage refers to the stage of online reasoning of input data using a trained neural network. In the online reasoning stage, input data and output data of each functional layer of the neural network are generated, which can also be called activation values, and the activation values ​​need to be quantized.

[0033] The offline computing device 10 is mainly used for some computing processes in the offline stage. The online computing device 20 is mainly used for some computing processes in the online stage. Both the offline computing device 10 and the online computing device 20 can be computer devices with data storage and computing capabilities such as computers, mobile phones, tablet computers, wearable devices, smart home devices, vehicle terminals, servers, etc., and this application does not limit this.

[0034] In addition, the above-mentioned offline computing device 10 and online computing device 20 can be two independent devices or the same device, and this application does not limit this.

[0035] For the sake of convenience, in the following method embodiments, only the execution subject of each step is described as a computer device. It is understandable that the computer device can be Figure 1 An online computing device 20 in the illustrated implementation environment.

[0036] Please refer to Figure 2 , which shows a flow chart of a neural network quantization method provided by an embodiment of the present application. The method may include at least one of the following steps (210-230):

[0037] Step 210, obtaining first quantization parameters of multiple functional layers of the neural network, the first quantization parameters are used to quantize data in a floating point format into data in a first fixed point format, and the first fixed point format is a format adopted by the first processing unit.

[0038] A neural network may include multiple functional layers. For example, a neural network may include an input layer, at least one hidden layer and an output layer, wherein the at least one hidden layer may include at least one of the following: at least one convolutional layer, at least one quasi-convolutional layer (such as a deconvolutional layer, a fully connected layer, etc.), at least one pooling layer, at least one cascade layer, at least one alignment operation layer, at least one binary layer, at least one scaling layer, at least one activation layer, and the like.

[0039] The first quantization parameters of the multiple functional layers of the neural network can be determined in the offline stage. In the online stage, the computer device obtains the first quantization parameters of the multiple functional layers of the neural network determined in the offline stage. Optionally, the neural network is fully quantized, that is, all functional layers of the neural network have corresponding first quantization parameters.

[0040] Optionally, the first processing unit is a CPU, and the first fixed-point format may be int8 (8-bit signed integer) format. The CPU usually stores and calculates data in int8 format, and the processing unit of the online computing device is usually mainly CPU, so we can first determine the first quantization parameter of each functional layer of the neural network in int8 format in the offline stage.

[0041] Exemplarily, the quantization formula from floating point format to fixed point format is as follows: Q=R / S+Z; the dequantization formula from fixed point format to floating point format is as follows: R=(QZ)*S; where R represents the real floating point value, Q represents the fixed point value after quantization, Z represents the quantized fixed point value corresponding to the floating point value of 0, and S represents the minimum scale that can be represented after fixed point quantization, also known as the scaling factor (scale). S=(R max -R min ) / (Q max -Q min ), Z=Q max -Rmax / S; where R max Represents the maximum floating point value, R min Indicates the smallest floating point value, Q max Indicates the maximum fixed-point value, Q min Represents the smallest fixed-point value.

[0042] In the embodiment of the present application, the quantization parameter includes a scaling factor (scale). Each functional layer of the neural network may have a corresponding quantization parameter. Moreover, since the value range of the parameters of each functional layer is different, the corresponding quantization parameters may also be different.

[0043] For the first type of functional layer, the first type of functional layer refers to a functional layer that has only an activation value but no weight value, such as a pooling layer, a cascade layer, a bitwise operation layer, a binary layer, a scaling layer, an activation layer, etc., and its quantization parameters include an input quantization parameter and an output quantization parameter. Among them, the input quantization parameter refers to the quantization parameter corresponding to the input data (or input activation value) of the functional layer, and the output quantization parameter refers to the quantization parameter corresponding to the output data (or output activation value) of the functional layer.

[0044] For the second type of functional layer, the second type of functional layer refers to a functional layer that has a weight value in addition to an activation value, such as a convolutional layer or a quasi-convolutional layer, and its quantization parameters may include: an input quantization parameter, an output quantization parameter, and a weight value quantization parameter. Among them, the input quantization parameter refers to the quantization parameter corresponding to the input data (or input activation value) of the functional layer, the output quantization parameter refers to the quantization parameter corresponding to the output data (or output activation value) of the functional layer, and the weight value quantization parameter refers to the quantization parameter corresponding to the weight value of the functional layer.

[0045] Optionally, the quantization parameter corresponding to the functional layer also includes a fused unified quantization parameter. For the first type of functional layer, the corresponding fused unified quantization parameter refers to the quantization parameter calculated based on the input quantization parameter and the output quantization parameter. For example, fused scale = output scale / input scale; wherein fused scale represents the fused unified quantization parameter, input scale represents the input quantization parameter, and output scale represents the output quantization parameter. For the second type of functional layer, the corresponding fused unified quantization parameter refers to the quantization parameter calculated based on the input quantization parameter, the output quantization parameter, and the weight value quantization parameter. For example, fused scale = output scale / (input scale * weight scale); wherein fused scale represents the fused unified quantization parameter, input scale represents the input quantization parameter, output scale represents the output quantization parameter, and weight scale represents the weight value quantization parameter. By calculating the fused unified quantization parameter in advance in the offline stage, the fused unified quantization parameter can be directly used in the subsequent online stage inference process to participate in the calculation, without the need to use the input quantization parameter, the output quantization parameter, the weight value quantization parameter, etc. to participate in the calculation, thereby helping to reduce the amount of calculation in the online stage and shorten the time consumption of online inference.

[0046] In the embodiment of the present application, the method for determining the quantization parameters of each functional layer is not limited, for example, it can be obtained by statistics according to the KL divergence (Kullback–Leibler divergence) method or by other methods.

[0047] Step 220, for a first functional layer among multiple functional layers, convert a first quantization parameter of the first functional layer into a second quantization parameter, the second quantization parameter is used to quantize data in a floating-point format into data in a second fixed-point format, the second fixed-point format is a format adopted by the second processing unit, and the first processing unit and the second processing unit are two different processing units.

[0048] The first functional layer can be any one of the multiple functional layers of the neural network. If the first functional layer needs to use a second processing unit for inference, and the second processing unit is another processing unit different from the first processing unit, such as the first processing unit is a CPU and the second processing unit is a DSP, then the first quantization parameter of the first functional layer needs to be converted into a second quantization parameter.

[0049] Optionally, the first processing unit is a CPU, and the first fixed-point format may be int8 (8-bit signed integer) format; the second processing unit is a DSP, and the second fixed-point format may be uint8 (8-bit unsigned integer) format. The DSP usually stores and calculates data in uint8 format. Since the first quantization parameter suitable for calculating data in int8 format is calculated in advance in the offline stage, the first quantization parameter needs to be converted into a second quantization parameter suitable for calculating data in uint8 format.

[0050] In addition, for the int8 format, the quantization value range is [-128, 127], and the center point is 0. Symmetric quantization uses the int8 format. For the uint8 format, the quantization value range is [0, 255], and the center point is 128. Asymmetric quantization uses the uint8 format.

[0051] Optionally, if the first functional layer belongs to the first type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter, and the first activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the first fixed point format. Step 220 may include: based on the conversion relationship between the first fixed point format and the second fixed point format, converting the first activation value quantization parameter of the first functional layer into a second activation value quantization parameter; wherein the second activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the second fixed point format. That is, for a functional layer that does not contain a weight value, only the activation value quantization parameter needs to be converted. Exemplarily, according to the first activation value quantization parameter and the value range corresponding to the first fixed point format, the value range corresponding to the activation value in the floating point format is determined, and then according to the value range corresponding to the activation value in the floating point format and the value range corresponding to the second fixed point format, the second activation value quantization parameter is determined.

[0052] Optionally, if the first functional layer belongs to the second type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter and a first weight value quantization parameter, the first activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the first fixed point format, and the first weight value quantization parameter is used to quantize the weight value in the floating point format into the weight value in the first fixed point format. Step 220 may include: based on the conversion relationship between the first fixed point format and the second fixed point format, converting the first activation value quantization parameter of the first functional layer into a second activation value quantization parameter; wherein the second activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the second fixed point format. Step 220 may also include: based on the conversion relationship between the first fixed point format and the second fixed point format, converting the first weight value quantization parameter of the first functional layer into a second weight value quantization parameter; wherein the second weight value quantization parameter is used to quantize the weight value in the floating point format into the weight value in the second fixed point format. That is, for a functional layer containing a weight value, in addition to converting the activation value quantization parameter, it is also necessary to convert the weight value quantization parameter. Exemplarily, according to the first activation value quantization parameter and the range of values ​​corresponding to the first fixed-point format, the range of values ​​corresponding to the activation value in the floating-point format is determined, and then according to the range of values ​​corresponding to the activation value in the floating-point format and the range of values ​​corresponding to the second fixed-point format, the second activation value quantization parameter is determined. Exemplarily, according to the first weight value quantization parameter and the range of values ​​corresponding to the first fixed-point format, the range of values ​​corresponding to the weight value in the floating-point format is determined, and then according to the range of values ​​corresponding to the weight value in the floating-point format and the range of values ​​corresponding to the second fixed-point format, the second weight value quantization parameter is determined.

[0053] Optionally, if the first functional layer belongs to the second type of functional layer, it is also necessary to convert the first weight value of the first functional layer into a second weight value based on the conversion relationship between the first fixed-point format and the second fixed-point format; wherein the first weight value refers to the weight value expressed in the first fixed-point format, and the second weight value refers to the weight value expressed in the second fixed-point format. In other words, when the weight value of the first functional layer has been calculated in the offline stage, and the weight value of the first functional layer is a first weight value expressed in the first fixed-point format suitable for calculation by the first processing unit, if the first functional layer needs to use the second processing unit for inference, it is also necessary to convert the first weight value into a second weight value expressed in the second fixed-point format suitable for calculation by the second processing unit.

[0054] Exemplarily, the formula for converting the first activation value quantization parameter (or the first weight value quantization parameter) into the second activation value quantization parameter (or the second weight value quantization parameter) is as follows:

[0055] real=scale1*(quantize-zero_point)

[0056] scale2=(f max -f min ) / (q max -q min )

[0057] Among them, scale1 represents the first activation value quantization parameter (or the first weight value quantization parameter), scale2 represents the second activation value quantization parameter (or the second weight value quantization parameter), quantize represents the quantized fixed-point value, zero_point represents the quantized fixed-point value corresponding to the 0 floating-point value, real represents the real floating-point value, and f max Represents the maximum floating point value, f min Represents the smallest floating point value, q max Represents the maximum fixed-point value, q min Indicates the minimum fixed-point value. For example, if the first fixed-point format is int8 format and the second fixed-point format is uint8 format, then zero_point = 0, the value of quantize is [-128, 127], and q max =255,q min =0.

[0058] Exemplarily, the calculation formula for converting the weight value in int8 format (ie, the first weight value described above) into the weight value in uint8 format (ie, the second weight value described above) is as follows:

[0059] Wuint8=Wint8^128

[0060] Among them, Wuint8 represents the weight value in uint8 format, and Wint8 represents the weight value in int8 format.

[0061] Step 230: When the second processing unit is used to infer the first functional layer, relevant data involved in the inference process of the first functional layer is quantized based on the second quantization parameter.

[0062] After obtaining the second quantization parameter corresponding to the first functional layer, when using the second processing unit to infer the first functional layer, the relevant data involved in the inference process can be quantized or dequantized based on the second quantization parameter to ensure the accuracy of the second processing unit's inference of the functional layer.

[0063] The technical solution provided in the embodiment of the present application converts the quantization parameters of the neural network from quantization parameters suitable for calculation by a first processing unit (such as a CPU) to quantization parameters suitable for calculation by a second processing unit (such as a DSP), thereby making it possible for the quantization parameters of the neural network to be adapted between a variety of different types of processing units, such as between a CPU and a DSP, to meet the computing requirements of being compatible with a variety of different types of processing units at the same time.

[0064] In addition, in addition to considering first-class functional layers such as convolutional layers or quasi-convolutional layers, the present application also considers first-class functional layers such as pooling layers, cascade layers, bit operation layers, binary layers, scaling layers, activation layers, etc., and provides a full-process symmetric quantization solution, and through the conversion of quantization parameters, it achieves compatibility with multiple different types of processing units.

[0065] Next, the reasoning process of the neural network is introduced and explained. Figure 3 As shown, the reasoning process may include at least one of the following steps (310-320):

[0066] Step 310: For a second functional layer among the multiple functional layers, determine an operation type of the second functional layer, where the operation type is used to indicate related characteristics of input data and output data of the second functional layer.

[0067] The second functional layer may be any functional layer of multiple functional layers of the neural network.

[0068] Optionally, the operation type includes: single-input single-output, multiple-input single-output and the operation rule is addition operation, multiple-input single-output and the operation rule is multiplication operation. The operation type of the second functional layer can be any one of the above operation types.

[0069] If the operation type of the second functional layer is single input single output, it means that the input data of the second functional layer has one and only one set of activation values, and the output data of the second functional layer also has one and only one set of activation values. In other words, the input data of the second functional layer is a feature map, and the output data is also a feature map. Exemplarily, the operation type of the pooling layer is single input single output.

[0070] If the operation type of the second functional layer is multi-input single output and the operation rule is addition operation, it means that the input data of the second functional layer includes multiple groups of activation values, the output data of the second functional layer has only one group of activation values, and the output data is obtained by adding the multiple groups of activation values ​​in the above input data. In other words, the input data of the second functional layer includes multiple feature maps, the output data is a feature map, and the output is obtained by adding the multiple feature maps in the above input data. Exemplarily, the bitwise operation layer that performs addition operation is multi-input single output and the operation rule is addition operation.

[0071] If the operation type of the second functional layer is multi-input single output and the operation rule is multiplication operation, it means that the input data of the second functional layer includes multiple groups of activation values, the output data of the second functional layer has only one group of activation values, and the output data is obtained by multiplying the multiple groups of activation values ​​in the above input data. In other words, the input data of the second functional layer includes multiple feature maps, the output data is a feature map, and the output is obtained by multiplying the multiple feature maps in the above input data. Exemplarily, the bitwise operation layer that performs multiplication operation is multi-input single output and the operation rule is multiplication operation.

[0072] Step 320 , based on the operation type of the second functional layer and according to the input data of the second functional layer, the second functional layer is inferred to obtain output data of the second functional layer.

[0073] Different processing methods can be used to reason about the functional layer for different operation types.

[0074] In some embodiments, when the operation type of the second functional layer is single-input single-output, the second functional layer is inferred according to the input data of the second functional layer in the target fixed-point format to obtain the output data of the second functional layer in the target fixed-point format.

[0075] The target fixed-point format may be the first fixed-point format or the second fixed-point format. Taking the case where the first processing unit is used to infer the second functional layer as an example, the target fixed-point format is the first fixed-point format. For example, if the first processing unit is a CPU, and the CPU is used to infer the second functional layer, the target fixed-point format is the int8 format. First, the input data of the second functional layer in the int8 format is obtained, and then the inference calculation of the second functional layer is directly performed using the data in the int8 format to obtain the output data of the second functional layer in the int8 format.

[0076] In some embodiments, when the operation type of the second functional layer is multiple-input single-output and the operation rule is addition operation, the multiple groups of input data are respectively dequantized according to the quantization parameters corresponding to the multiple groups of input data of the second functional layer in the target fixed-point format to obtain the multiple groups of input data in the floating-point format; the multiple groups of input data in the floating-point format are added to obtain the addition operation results in the floating-point format; the addition operation results in the floating-point format are quantized to obtain the output data of the second functional layer in the target fixed-point format.

[0077] The target fixed-point format may be the first fixed-point format or the second fixed-point format mentioned above. Taking the case of using the first processing unit to infer the second functional layer as an example, the target fixed-point format is the first fixed-point format. For example, the first processing unit is a CPU, and when the CPU is used to perform inference calculations on the second functional layer, the target fixed-point format is the int8 format. First, multiple sets of input data of the second functional layer in the int8 format are obtained. Since the quantization parameters corresponding to the multiple sets of input data may be different, they are expressed on different fixed-point scales and cannot directly participate in the addition operation. They need to be unified to the same scale before they can be calculated. Therefore, it is necessary to perform dequantization on the multiple sets of input data according to the quantization parameters corresponding to the multiple sets of input data, respectively, to obtain multiple sets of input data in floating-point format, and then perform addition operations on the multiple sets of input data in floating-point format to obtain the addition operation results in floating-point format, and finally quantize the addition operation results in floating-point format to obtain the output data of the second functional layer in int8 format.

[0078] Exemplarily, taking the second functional layer including 2 sets of input data as an example, the calculation formula of the output data of the second functional layer is as follows:

[0079] C=Clamp((reScaleA*A+reScaleB*B)*scaleC)

[0080] Among them, C represents the output data of the second functional layer, A represents one set of input data of the second functional layer, B represents another set of input data of the second functional layer, reScaleA represents the inverse quantization parameter corresponding to the input data A, reScaleB represents the inverse quantization parameter corresponding to the input data B, scaleC represents the quantization parameter corresponding to the output data C, and Clamp represents truncating the floating-point number to an integer.

[0081] In some embodiments, when the operation type of the second functional layer is multiple-input single-output and the operation rule is multiplication operation, multiplication operation is performed on multiple groups of input data of the second functional layer in the target fixed-point format to obtain the multiplication operation result in the target fixed-point format; based on the multiplication operation result in the target fixed-point format, the output data of the second functional layer in the target fixed-point format is generated.

[0082] The target fixed-point format may be the first fixed-point format or the second fixed-point format mentioned above. Taking the use of the first processing unit to perform inference on the second functional layer as an example, the target fixed-point format is the first fixed-point format. For example, the first processing unit is a CPU, and when the CPU is used to perform inference calculations on the second functional layer, the target fixed-point format is the int8 format. First, multiple sets of input data of the second functional layer in the int8 format are obtained. Although the quantization parameters corresponding to the multiple sets of input data may be different, they can be directly involved in the multiplication operation. Therefore, the multiple sets of input data in the int8 format are directly multiplied to obtain the multiplication result in the int8 format, and then the output data of the second functional layer in the int8 format is generated according to the multiplication result in the int8 format.

[0083] Exemplarily, taking the second functional layer including 2 sets of input data as an example, the calculation formula of the output data of the second functional layer is as follows:

[0084] C=Clamp((A*B)*fusedScale)

[0085] fusedScale=reScaleA*reScaleB*scaleC

[0086] Among them, C represents the output data of the second functional layer, A represents one set of input data of the second functional layer, B represents another set of input data of the second functional layer, reScaleA represents the inverse quantization parameter corresponding to the input data A, reScaleB represents the inverse quantization parameter corresponding to the input data B, scaleC represents the quantization parameter corresponding to the output data C, and Clamp represents truncating the floating-point number to an integer.

[0087] This embodiment divides the reasoning process of the neural network into a plurality of different operation types according to the operation characteristics of each functional layer, and then adopts different calculation methods to perform reasoning for different operation types, so as to improve the reasoning speed as much as possible and reduce the amount of calculation while ensuring the accuracy of reasoning.

[0088] In an exemplary embodiment, in order to improve the quantization accuracy, the present application proposes a re-quantization control strategy. Quantization reasoning is prone to precision loss, mainly from two aspects: one is the quantization representation method, that is, whether the calculation of the quantization parameter (i.e., scale) is reasonable, and the other is the truncation error from floating point to fixed point. The quantization parameter is obtained by the KL divergence method, which has been verified by the industry. If the truncation error can be reduced, it will also bring about precision improvement. For a single-input single-output type and a relatively simple calculation function layer, such as a pooling layer, the present application finds that its input quantization parameter and output quantization parameter are relatively close. When both are less than a certain threshold, the function layer does not need to be re-quantized, and the result can be directly calculated and output in a fixed-point manner, thereby omitting the re-quantization truncation step, which can not only improve the calculation accuracy, but also reduce some unnecessary calculations. Therefore, in the offline stage, a re-quantization parameter can be added to the neural network to indicate whether a certain function layer needs to be re-quantized. Exemplarily, for any function layer, it can be determined whether the function layer needs to be re-quantized based on the difference between the input quantization parameter and the output quantization parameter of the function layer, and the corresponding re-quantization parameter is obtained and written into the model structure. Exemplarily, if the difference between the input quantization parameter and the output quantization parameter of the functional layer is less than or equal to the threshold, it is determined that the functional layer does not need re-quantization, otherwise it needs re-quantization. The above threshold can be determined in combination with actual experience or experiments. For example, if the threshold is 0.01, if |input scale–output scale|≤0.01, re-quantization is not required, otherwise it needs re-quantization. Among them, input scale represents the input quantization parameter, and output scale represents the output quantization parameter.

[0089] Next, taking the third functional layer in the neural network as an example, the third functional layer can be any functional layer. When the requantization parameter corresponding to the third functional layer indicates that requantization is required, the third functional layer can be inferred in the following manner in the online stage:

[0090] In a possible implementation, the input data of the third functional layer in the target fixed-point format is dequantized to obtain the input data of the third functional layer in the floating-point format, and then the input data of the third functional layer in the floating-point format is used to infer the third functional layer to obtain the output data of the third functional layer in the floating-point format, and finally the output data of the third functional layer in the floating-point format is quantized to obtain the output data of the third functional layer in the target fixed-point format. Taking the target fixed-point format as the int8 format as an example, according to the input quantization parameter of the third functional layer, the input data of the third functional layer in the int8 format is dequantized to obtain the input data of the third functional layer in the floating-point format, and then the input data of the third functional layer in the floating-point format is used to perform inference calculation to obtain the output data of the third functional layer in the floating-point format, and finally the output quantization parameter corresponding to the third functional layer is used to quantize the output data of the third functional layer in the floating-point format to obtain the output data of the third functional layer in the int8 format.

[0091] In another possible implementation, the input data of the third functional layer in the target fixed-point format is used to infer the third functional layer to obtain the initial output data of the third functional layer in the target fixed-point format, and then based on the fused unified quantization parameter corresponding to the third functional layer, the initial output data of the third functional layer in the target fixed-point format is converted into the output data of the third functional layer in the target fixed-point format. Taking the target fixed-point format as int8 format as an example, the input data of the third functional layer in int8 format is used to perform inference calculation to obtain the initial output data of the third functional layer in int8 format, and then based on the fused unified quantization parameter corresponding to the third functional layer, the initial output data of the third functional layer in int8 format is calculated to obtain the output data of the third functional layer in int8 format. The fused unified quantization parameter corresponding to the third functional layer can be calculated based on the input quantization parameter and the output quantization parameter corresponding to the third functional layer.

[0092] In addition, when the re-quantization parameter corresponding to the third functional layer indicates that re-quantization is not required, the third functional layer can be inferred in the online stage in the following manner: the input data of the third functional layer in the target fixed-point format is used to infer the third functional layer to obtain the output data of the third functional layer in the target fixed-point format. That is, the target fixed-point format can be directly used for inference calculation without format conversion or post-processing.

[0093] The above-mentioned re-quantization control strategy proposed in the embodiment of the present application can not only improve the calculation accuracy, but also reduce some unnecessary calculation amounts.

[0094] In an exemplary embodiment, when the input data of the neural network is in uint8 format, for example, when the input data of the neural network is image data, its value range is uint8 format data in [0.255]. In order to be compatible with input data in uint8 format, the present application integrates the quantization parameters of the input layer of the neural network.

[0095] The input layer generally needs to be preprocessed, with the user inputting the mean and variance. The general operation steps are as follows: Figure 4 First, a preprocessing step is performed: the input data in uint8 format is converted to input data in float format according to the mean and variance; then, a quantization step is performed: the input data in float format is converted to input data in int8 format according to the quantization parameters.

[0096] like Figure 4 As shown, the conversion process is expressed by the following formula:

[0097] mid(float)=(input(uint8)–mean)*norm

[0098] output(int8)=Clamp(mid(float)*scale)

[0099] Among them, input(uint8) represents the input data in uint8 format, mid(float) represents the input data in float format obtained after the preprocessing step, output(int8) represents the input data in int8 format obtained by the final conversion, mean represents the mean of the input data, norm represents the variance of the input data, scale represents the quantization parameter, and clamp represents truncating the floating-point number to an integer.

[0100] The present application proposes that the above-mentioned mean and quantization parameters can be fused. After fusion, the preprocessing step can be omitted, and the quantization step can be directly executed to complete the conversion from the input data in uint8 format to the input data in int8 format. Figure 5 As shown, the formula is as follows:

[0101] output(int8)=Clamp((input(uint8)–zeroPoint)*scale')

[0102] Among them, input(uint8) represents the input data in uint8 format, output(int8) represents the input data in int8 format obtained by the final conversion, scale' represents the quantization parameter after fusion, scale'=norm*scale, zeroPoint=mean, mean represents the mean of the input data, norm represents the variance of the input data, scale represents the quantization parameter, and Clamp represents truncating the floating point number to an integer.

[0103] based on Figure 5 The corresponding formula can be obtained as follows: when the input data of the neural network is in the second fixed-point format, the translation factor corresponding to the second fixed-point format is determined according to the mean of the input data of the neural network; the fused quantization parameter is determined according to the variance of the input data of the neural network and the quantization parameter corresponding to the second fixed-point format; the input data of the neural network is converted from the second fixed-point format to the first fixed-point format according to the translation factor corresponding to the second fixed-point format and the fused quantization parameter. Exemplarily, the first fixed-point format is int8 format, and the second fixed-point format is uint8 format.

[0104] In an embodiment of the present application, when the neural network adopts the first fixed-point format for inference calculation, compatibility with input data in the second fixed-point format is achieved, so that input data in the second fixed-point format can also be input into the neural network to participate in inference calculation, thereby improving the applicability of the neural network.

[0105] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0106] Please refer to Figure 6 , which shows a block diagram of a neural network quantization device provided by an embodiment of the present application. The device has the function of implementing the above-mentioned neural network quantization method, and the function can be implemented by hardware, or by hardware executing corresponding software. The device can be a computer device, or it can be set in a computer device. The device 600 may include: a parameter acquisition module 610, a parameter conversion module 620 and a data quantization module 630.

[0107] The parameter acquisition module 610 is used to obtain first quantization parameters of multiple functional layers of the neural network, where the first quantization parameters are used to quantize data in a floating-point format into data in a first fixed-point format, where the first fixed-point format is the format used by the first processing unit.

[0108] The parameter conversion module 620 is used to convert a first quantization parameter of a first functional layer among the multiple functional layers into a second quantization parameter, wherein the second quantization parameter is used to quantize the data in the floating-point format into data in a second fixed-point format, and the second fixed-point format is a format adopted by the second processing unit, and the first processing unit and the second processing unit are two different processing units.

[0109] The data quantization module 630 is used to quantize relevant data involved in the reasoning process of the first functional layer based on the second quantization parameter when the second processing unit is used to infer the first functional layer.

[0110] In an exemplary embodiment, if the first functional layer belongs to the first type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter, and the first activation value quantization parameter is used to quantize the activation value in the floating-point format into the activation value in the first fixed-point format; the parameter conversion module 620 is used to convert the first activation value quantization parameter of the first functional layer into a second activation value quantization parameter based on the conversion relationship between the first fixed-point format and the second fixed-point format; wherein the second activation value quantization parameter is used to quantize the activation value in the floating-point format into the activation value in the second fixed-point format.

[0111] In an exemplary embodiment, if the first functional layer belongs to the second type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter and a first weight value quantization parameter, the first activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the first fixed point format, and the first weight value quantization parameter is used to quantize the weight value in the floating point format into the weight value in the first fixed point format; the parameter conversion module 620 is used to convert the first activation value quantization parameter of the first functional layer into a second activation value quantization parameter based on the conversion relationship between the first fixed point format and the second fixed point format; wherein the second activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the second fixed point format; based on the conversion relationship between the first fixed point format and the second fixed point format, the first weight value quantization parameter of the first functional layer is converted into a second weight value quantization parameter; wherein the second weight value quantization parameter is used to quantize the weight value in the floating point format into the weight value in the second fixed point format.

[0112] Optionally, the parameter conversion module 620 is also used to convert the first weight value of the first functional layer into a second weight value based on the conversion relationship between the first fixed-point format and the second fixed-point format; wherein the first weight value refers to the weight value expressed in the first fixed-point format, and the second weight value refers to the weight value expressed in the second fixed-point format.

[0113] In an exemplary embodiment, if Figure 7 As shown, the device 600 further includes: a type determination module 640 and an inference calculation module 650 .

[0114] The type determination module 640 is used to determine, for a second functional layer among the multiple functional layers, an operation type of the second functional layer, where the operation type is used to indicate relevant characteristics of input data and output data of the second functional layer.

[0115] The inference calculation module 650 is used to infer the second functional layer based on the operation type and according to the input data of the second functional layer to obtain the output data of the second functional layer.

[0116] Optionally, the inference calculation module 650 is used to infer the second functional layer according to the input data of the second functional layer in the target fixed-point format when the operation type is single-input single-output, so as to obtain the output data of the second functional layer in the target fixed-point format.

[0117] Optionally, the inference calculation module 650 is used to, when the operation type is multiple-input and single-output and the operation rule is addition operation, respectively dequantize the multiple groups of input data according to quantization parameters corresponding to the multiple groups of input data of the second functional layer in the target fixed-point format to obtain multiple groups of input data in floating-point format; perform addition operation on the multiple groups of input data in the floating-point format to obtain addition operation results in the floating-point format; and quantize the addition operation results in the floating-point format to obtain output data of the second functional layer in the target fixed-point format.

[0118] Optionally, the inference calculation module 650 is used to perform multiplication operations on multiple groups of input data of the second functional layer in the target fixed-point format when the operation type is multiple-input and single-output and the operation rule is multiplication operation, to obtain the multiplication operation results in the target fixed-point format; and generate the output data of the second functional layer in the target fixed-point format according to the multiplication operation results in the target fixed-point format.

[0119] In an exemplary embodiment, if Figure 7 As shown, the device 600 also includes: an inference calculation module 650.

[0120] The inference calculation module 650 is configured to, for a third functional layer among the plurality of functional layers, when a re-quantization parameter corresponding to the third functional layer indicates that re-quantization is required,

[0121] Dequantizing the input data of the third functional layer in the target fixed-point format to obtain the input data of the third functional layer in the floating-point format; inferring the third functional layer using the input data of the third functional layer in the floating-point format to obtain the output data of the third functional layer in the floating-point format; quantizing the output data of the third functional layer in the floating-point format to obtain the output data of the third functional layer in the target fixed-point format;

[0122] Alternatively, the input data of the third functional layer in the target fixed-point format is used to infer the third functional layer to obtain initial output data of the third functional layer in the target fixed-point format; based on the fused unified quantization parameter corresponding to the third functional layer, the initial output data of the third functional layer in the target fixed-point format is converted into output data of the third functional layer in the target fixed-point format.

[0123] In an exemplary embodiment, if Figure 7 As shown, the device 600 also includes: a data conversion module 660.

[0124] The data conversion module 660 is used to determine, when the input data of the neural network is in the second fixed-point format, a translation factor corresponding to the second fixed-point format according to the mean of the input data of the neural network; determine a fused quantization parameter according to the variance of the input data of the neural network and the quantization parameter corresponding to the second fixed-point format; and convert the input data of the neural network from the second fixed-point format to the first fixed-point format according to the translation factor corresponding to the second fixed-point format and the fused quantization parameter.

[0125] The embodiment of the present application converts the quantization parameters of the neural network from quantization parameters suitable for calculation by a first processing unit (such as a CPU) to quantization parameters suitable for calculation by a second processing unit (such as a DSP), thereby making it possible for the quantization parameters of the neural network to be adapted between a variety of different types of processing units, such as between a CPU and a DSP, to meet the computing requirements of being compatible with a variety of different types of processing units.

[0126] It should be noted that the device provided in the above embodiment, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0127] In an exemplary embodiment, a computer device is also provided, the computer device comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned neural network quantization method.

[0128] In an exemplary embodiment, a computer-readable storage medium is also provided, in which a computer program is stored, and the computer program is loaded and executed by a processor to implement the above-mentioned neural network quantization method. Optionally, the above-mentioned computer-readable storage medium can be ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk and optical data storage device, etc.

[0129] In an exemplary embodiment, a computer program product is also provided, the computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium, and the processor reading and executing the computer instructions from the computer-readable storage medium to implement the above-mentioned neural network quantization method.

[0130] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application are not limited to this.

[0131] The above are merely exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A neural network quantization method, characterized in that: The method comprises: Acquire first quantization parameters of multiple functional layers of the neural network, where the first quantization parameters are used to quantize data in a floating point format into data in a first fixed point format, where the first fixed point format is a format used by the first processing unit; For a first functional layer among the multiple functional layers, converting a first quantization parameter of the first functional layer into a second quantization parameter, where the second quantization parameter is used to quantize the data in the floating point format into data in a second fixed point format, where the second fixed point format is a format adopted by a second processing unit, where both the first processing unit and the second processing unit are hardware units for inferring the functional layer of the neural network, and the first processing unit and the second processing unit are two different processing units; In a case where the second processing unit is used to infer the first functional layer, quantizing relevant data involved in the inference process of the first functional layer based on the second quantization parameter; For a third functional layer among the multiple functional layers, when the difference between the input quantization parameter and the output quantization parameter of the third functional layer is greater than a threshold value, the input data of the third functional layer in the target fixed-point format is used to infer the third functional layer to obtain initial output data of the third functional layer in the target fixed-point format; based on the fused unified quantization parameter corresponding to the third functional layer, the initial output data of the third functional layer in the target fixed-point format is converted into the output data of the third functional layer in the target fixed-point format, wherein the fused unified quantization parameter corresponding to the third functional layer is obtained based on the input quantization parameter and the output quantization parameter of the third functional layer, and when the first processing unit is used to infer the third functional layer, the target fixed-point format is the first fixed-point format, and when the second processing unit is used to infer the third functional layer, the target fixed-point format is the second fixed-point format.

2. The method according to claim 1, characterized in that If the first functional layer belongs to the first type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter, and the first activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the first fixed point format; The converting the first quantization parameter of the first functional layer into a second quantization parameter includes: Based on the conversion relationship between the first fixed-point format and the second fixed-point format, the first activation value quantization parameter of the first functional layer is converted into a second activation value quantization parameter; wherein the second activation value quantization parameter is used to quantize the activation value in the floating-point format into the activation value in the second fixed-point format.

3. The method according to claim 1, characterized in that If the first functional layer belongs to the second type of functional layer, the first quantization parameter of the first functional layer includes a first activation value quantization parameter and a first weight value quantization parameter, the first activation value quantization parameter is used to quantize the activation value in the floating point format into the activation value in the first fixed point format, and the first weight value quantization parameter is used to quantize the weight value in the floating point format into the weight value in the first fixed point format; The converting the first quantization parameter of the first functional layer into a second quantization parameter includes: Based on the conversion relationship between the first fixed-point format and the second fixed-point format, converting the first activation value quantization parameter of the first functional layer into a second activation value quantization parameter; wherein the second activation value quantization parameter is used to quantize the activation value in the floating-point format into the activation value in the second fixed-point format; Based on the conversion relationship between the first fixed-point format and the second fixed-point format, the first weight value quantization parameter of the first functional layer is converted into a second weight value quantization parameter; wherein the second weight value quantization parameter is used to quantize the weight value in the floating-point format into the weight value in the second fixed-point format.

4. The method according to claim 3, characterized in that The method further comprises: Based on the conversion relationship between the first fixed-point format and the second fixed-point format, the first weight value of the first functional layer is converted into a second weight value; wherein the first weight value refers to the weight value expressed in the first fixed-point format, and the second weight value refers to the weight value expressed in the second fixed-point format.

5. The method according to claim 1, characterized in that The method further comprises: For a second functional layer among the plurality of functional layers, determining an operation type of the second functional layer, the operation type being used to indicate relevant characteristics of input data and output data of the second functional layer; Based on the operation type, the second functional layer is inferred according to the input data of the second functional layer to obtain output data of the second functional layer.

6. The method according to claim 5, characterized in that The inferring the second functional layer based on the operation type and according to the input data of the second functional layer to obtain the output data of the second functional layer includes: In the case where the operation type is single-input single-output, the second functional layer is inferred according to the input data of the second functional layer in the target fixed-point format to obtain the output data of the second functional layer in the target fixed-point format.

7. The method according to claim 5, characterized in that The inferring the second functional layer based on the operation type and according to the input data of the second functional layer to obtain the output data of the second functional layer includes: When the operation type is multiple-input single-output and the operation rule is addition operation, dequantizing the multiple groups of input data respectively according to the quantization parameters respectively corresponding to the multiple groups of input data of the second functional layer in the target fixed-point format to obtain the multiple groups of input data in the floating-point format; Performing an addition operation on the plurality of groups of input data in the floating point format to obtain an addition operation result in the floating point format; The addition operation result in the floating-point format is quantized to obtain output data of the second functional layer in the target fixed-point format.

8. The method according to claim 5, characterized in that The inferring the second functional layer based on the operation type and according to the input data of the second functional layer to obtain the output data of the second functional layer includes: When the operation type is multiple-input single-output and the operation rule is a multiplication operation, performing a multiplication operation on multiple groups of input data of the second functional layer in a target fixed-point format to obtain a multiplication operation result in the target fixed-point format; According to the multiplication result in the target fixed-point format, output data of the second functional layer in the target fixed-point format is generated.

9. The method according to claim 1, characterized in that: The method further comprises: When the input data of the neural network is in the second fixed-point format, determining a translation factor corresponding to the second fixed-point format according to a mean value of the input data of the neural network; Determining a fused quantization parameter according to the variance of the input data of the neural network and the quantization parameter corresponding to the second fixed-point format; The input data of the neural network is converted from the second fixed-point format to the first fixed-point format according to the translation factor corresponding to the second fixed-point format and the fused quantization parameter.

10. A quantization device for a neural network, characterized in that: The device comprises: A parameter acquisition module, used to acquire first quantization parameters of multiple functional layers of the neural network, wherein the first quantization parameters are used to quantize data in a floating point format into data in a first fixed point format, wherein the first fixed point format is a format adopted by the first processing unit; a parameter conversion module, configured to convert, for a first functional layer among the multiple functional layers, a first quantization parameter of the first functional layer into a second quantization parameter, wherein the second quantization parameter is used to quantize the data in the floating point format into data in a second fixed point format, wherein the second fixed point format is a format adopted by a second processing unit, wherein both the first processing unit and the second processing unit are hardware units for inferring the functional layer of the neural network, and the first processing unit and the second processing unit are two different processing units; a data quantization module, configured to quantize relevant data involved in the inference process of the first functional layer based on the second quantization parameter when the second processing unit is used to infer the first functional layer; An inference calculation module is used to, for a third functional layer among the multiple functional layers, use the input data of the third functional layer in a target fixed-point format to infer the third functional layer to obtain initial output data of the third functional layer in the target fixed-point format when the difference between the input quantization parameter and the output quantization parameter of the third functional layer is greater than a threshold value; based on the fused unified quantization parameter corresponding to the third functional layer, convert the initial output data of the third functional layer in the target fixed-point format into the output data of the third functional layer in the target fixed-point format, wherein the fused unified quantization parameter corresponding to the third functional layer is obtained based on the input quantization parameter and the output quantization parameter of the third functional layer, and when the first processing unit is used to infer the third functional layer, the target fixed-point format is the first fixed-point format, and when the second processing unit is used to infer the third functional layer, the target fixed-point format is the second fixed-point format.

11. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The computer program product comprises computer instructions, which are stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for converting weights of deep neural network

    CN112418391A

  • Data processing method and related product

    CN113222098A

  • FPGA-based AI chip neural network acceleration method

    CN113392973A