Quantization parameter determination method according to confidence threshold of neural network

By determining quantization parameters based on a confidence threshold, the method reduces quantization errors in neural networks, enhancing the accuracy and speed of operations on dedicated hardware by optimizing quantization and dequantization processes.

WO2025154912A1PCT designated stage expired Publication Date: 2025-07-24OPENEDGES TECH INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017213
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2024-11-04
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing neural networks face significant quantization errors when transitioning from high-precision data formats to low-precision formats for faster operation on dedicated hardware, particularly in layers further downstream, which are exacerbated by discarding quantized data without considering the range of values.

Method used

The method involves determining quantization parameters by considering a confidence threshold at the confidence output layer, using input/output conversion functions to adjust boundary values for quantization and dequantization layers, thereby reducing quantization errors by clipping values outside a defined range.

Benefits of technology

This approach minimizes quantization errors by optimizing quantization parameters, ensuring accurate and efficient neural network operations on dedicated hardware with reduced clipping errors and improved precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017213_24072025_PF_FP_ABST
    Figure KR2024017213_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a quantization parameter determination method for determining quantization parameters used in a neural network comprising: a confidence output layer; a mapping layer for outputting a third activation inputted to the confidence output layer; a dequantization layer for outputting a second activation inputted to the mapping layer; and a quantization layer for outputting a first activation inputted to the dequantization layer. The method comprises: a step of obtaining a confidence threshold assigned to the confidence output layer; a step of determining, on the basis of an input-output transformation function of the mapping layer, an input value for the mapping layer such that the mapping layer outputs the confidence threshold as the third activation; a step of determining, on the basis of the determined input value, a boundary value for determining quantization parameters commonly used in both the dequantization layer and the quantization layer; and a step of determining, on the basis of the determined boundary value, the quantization parameters commonly used in both the quantization layer and the dequantization layer.
Need to check novelty before this filing date? Find Prior Art

Description

A method for determining quantization parameters based on the confidence threshold of a neural network

[0001] The present invention relates to a technique for optimizing parameters of a neural network to be provided to a neural network processing unit (NPU) in order to reduce quantization errors occurring during neural network operations in a computing device including a neural network processing unit (NPU).

[0002] The present invention relates to a neural network operation executed in an NPU installed in a computing device. In a neural network, data (activations) can be calculated and transformed each time they encounter a layer while moving along one direction. This transformation and flow of data can be expressed by the term "stream." The neural network may include a first layer and a second layer. In this case, if the output activation output by the first layer is input to the second layer as is or after further transformation, the first layer may be referred to as a layer located further upstream than the second layer, and the second layer may be referred to as a layer located further downstream than the first layer. The terms "upstream" and "downstream" are introduced for the convenience of explaining the present invention.

[0003] Computing devices such as desktop computers, laptop computers, smartphones, and tablets may be equipped with a Neural Processing Unit (NPU). The NPU may have a structure suitable for neural network operations. In this case, for the NPU to execute neural network operations, a control unit within the NPU must execute certain commands for neural network operations to control resources within the NPU. The commands may be stored within the NPU during the manufacturing process of the user device, or may be provided to the NPU after the user device is manufactured. The commands may be generated and provided by a computing device other than the computing device including the NPU.

[0004] Meanwhile, a neural network that performs a specific function can be developed by the developer of the neural network using a general-purpose computer. The neural network developed in this way may include multiple layers, and the parameters assigned to each layer may be provided in a high-precision data format such as FP32. The neural network can be executed on a general-purpose computer, but in this case, the operation speed may be very slow or not fast enough. Therefore, to improve the operation speed of the neural network, it is necessary to execute the neural network on dedicated hardware including an NPU developed for neural network operations. However, command codes developed to operate on the general-purpose computer cannot be used on dedicated hardware other than the general-purpose computer. Therefore, a process of converting the command codes into command codes (commands) that can be executed on the dedicated hardware is necessary.

[0005] At this time, the computing resources of the dedicated hardware are limited compared to a general-purpose computer, and the tolerance of its operating method may not be flexible. Furthermore, in order to achieve a high computational speed on the dedicated hardware, the parameters assigned to each layer of the neural network must be converted to a low-precision data format and may need to be quantized through an additional quantization process. As a result of this quantization operation, quantization errors occur in the parameters assigned to each layer of the neural network operating on the dedicated hardware. As a result, quantization errors also occur in the neural network computation results using the parameters. The quantization errors that occur may become larger or unpredictable as they progress downstream of the neural network.

[0006] In particular, in the second layer further downstream from the first layer that performs quantization, there may be cases where some of the quantized data output by the first layer is discarded. In this case, if quantization is performed in the first layer without considering the range of values ​​of the discarded data, the quantization error may increase. The inventor of the present patent application came up with the idea that if quantization is performed in the first layer considering the range of values ​​of the discarded data, the quantization error can be reduced. As a representative example of the second layer, there is a confidence output layer that is located at the most downstream of the neural network. The confidence output layer is configured to receive a scalar value having a value within a limited first range, and to discard the scalar value if it has a value within a specific range within the first range without outputting it.

[0007] The present invention aims to provide a technique for reducing quantization errors resulting from quantization operations using quantization parameters by considering a confidence threshold assigned to the confidence output layer in the process of determining quantization parameters used in the last quantization layer located upstream from the confidence output layer of a neural network.

[0008] Quantization typically refers to mapping a continuous range of values ​​into a finite number of ranges. In one embodiment, quantization may mean converting values ​​represented as Floating Point (FP) types into values ​​represented as Integer (INT) types.

[0009] Quantization used in the AI ​​field can be presented as in [Equation 1].

[0010] [Formula 1]

[0011] x q = clip(round(x fp * s + z), -2 bit-1 , 2 bit-1-1)

[0012] Here x q are quantized values, and x fp is a value expressed in FP type (e.g. FP32), s is the quantization scale, bit is the quantization bit, and z is the zero point.

[0013] Among the quantization parameters used in [Formula 1], s and z can be presented as in [Formula 2] and [Formula 3], respectively.

[0014] [Formula 2]

[0015] s = (2 bit - 1) / (max(x fp ) - min(x fp ))

[0016] [Formula 3]

[0017] z = -min(x fp ) * s - 2 bit - 1 or

[0018] z = round(-min(x fp ) * s - 2 bit - 1 )

[0019] Here, max(x fp ) is x to which the above quantization parameters s and z are to be applied. fp It means the 'maximum boundary value', which is the largest value selected from the set of fields, and min(x fp ) is the above x fp It means the 'minimum boundary value' which is the smallest value selected from the set of fields. In one example, the candidates for the largest and smallest values ​​are x fp It can be restricted to values ​​that have a distance of less than a predetermined value from the representative value of the field.

[0020] In the neural network operation process, the above x fpA set of fields can be, for example, a tensor representing one activation or the weights assigned to one layer.

[0021] According to one aspect of the present invention, an NPU command generation method may be provided that generates an NPU command that causes a first computing device (1) including an NPU (Neural Processing Unit) (110) to execute a neural network (50) including a confidence output layer (514). The method comprises: generating an NPU command that causes a computing device (2) to execute a confidence threshold (Th) assigned to the confidence output layer of the neural network. c ) obtaining step; the computing device, based on the input / output conversion function (f) of the mapping layer (513) that outputs the third activation (713) input to the confidence output layer, outputs the input value (f) of the mapping layer so that the mapping layer outputs the confidence threshold value as the third activation -1 (Th c )); a step of determining, by the computing device, a boundary value for determining a quantization parameter (s1, z1) commonly used in a dequantization layer (512) that outputs a second activation (712) input to the mapping layer and a quantization layer (511) that outputs a first activation (711) input to the dequantization layer, based on the determined input value; a step of determining, by the computing device, the quantization parameter commonly used in the quantization layer and the dequantization layer based on the determined boundary value; and a command generation step of generating, by the computing device, a command set that causes the NPU to execute the neural network, wherein the command set includes the determined quantization parameter.

[0022] At this time, the third activation and the second activation may be scalars expressed in the FP (floating point) type, and the first activation may be a scalar expressed in the INT (integer) type.

[0023] At this time, the above-mentioned confidence output layer may be the output layer of the above-mentioned neural network.

[0024] At this time, the above-mentioned confidence output layer may be a layer located upstream from the output layer of the above-mentioned neural network.

[0025] At this time, the input value of the mapping layer that outputs the confidence threshold value as the third activation may be a value output by the inverse function of the input / output conversion function that receives the confidence threshold value as input.

[0026] At this time, the confidence output layer is configured to discard values ​​smaller than the confidence threshold value, and the input / output transformation function is a monotonically increasing function in which the value output by the mapping layer increases as the value input to the mapping layer increases, and the determined boundary value is a minimum boundary value that is a smaller value among two boundary values ​​used to determine the quantization parameter, and the determined minimum boundary value may be selected from values ​​smaller than the input value of the determined mapping layer. At this time, the determined minimum boundary value may be a maximum value among values ​​smaller than the input value of the determined mapping layer, or the determined minimum boundary value may be a value that satisfies a condition in which a difference value obtained by subtracting the determined minimum boundary value from the input value of the determined mapping layer is greater than 0 and smaller than a predetermined value.

[0027] At this time, the confidence output layer is configured to discard values ​​smaller than the confidence threshold value, and the input / output transformation function is a monotonically decreasing function in which the value output by the mapping layer decreases as the value input to the mapping layer increases, and the determined boundary value is a maximum boundary value that is a larger value among two boundary values ​​used to determine the quantization parameter, and the determined maximum boundary value may be selected from values ​​larger than the input value of the determined mapping layer. At this time, the determined maximum boundary value may be a minimum value among values ​​larger than the input value of the determined mapping layer, or the determined maximum boundary value may be a value that satisfies a condition in which a difference value obtained by subtracting the input value of the determined mapping layer from the determined maximum boundary value is larger than 0 and smaller than a predetermined value.

[0028] At this time, the confidence output layer is configured to discard values ​​greater than the confidence threshold value, and the input / output transformation function is a monotonically increasing function in which the value output by the mapping layer increases as the value input to the mapping layer increases, and the determined boundary value is a maximum boundary value that is a larger value among two boundary values ​​used to determine the quantization parameter, and the determined maximum boundary value may be selected from values ​​greater than the input value of the determined mapping layer. At this time, the determined maximum boundary value may be a minimum value among values ​​greater than the input value of the determined mapping layer, or the determined maximum boundary value may be a value that satisfies a condition in which a difference value obtained by subtracting the input value of the determined mapping layer from the determined maximum boundary value is greater than 0 and less than a predetermined value.

[0029] At this time, the confidence output layer is configured to discard values ​​greater than the confidence threshold value, and the input / output transformation function is a monotonically decreasing function in which the value output by the mapping layer decreases as the value input to the mapping layer increases, and the determined boundary value is a minimum boundary value which is a smaller value among two boundary values ​​used to determine the quantization parameter, and the determined minimum boundary value may be selected from values ​​smaller than the input value of the determined mapping layer. At this time, the determined minimum boundary value may be a maximum value among values ​​smaller than the input value of the determined mapping layer, or the determined minimum boundary value may be a value that satisfies a condition in which a difference value obtained by subtracting the determined minimum boundary value from the input value of the determined mapping layer is greater than 0 and smaller than a predetermined value.

[0030] At this time, the input / output conversion function is a function that outputs the same value as the value input to the mapping layer, and the input value of the determined mapping layer may be the same as the confidence threshold value.

[0031] In the claims of this patent application, the configuration that “the computing device determines an input value of the mapping layer that causes the mapping layer to output the confidence threshold value as the third activation based on an input / output conversion function of the mapping layer that outputs the third activation input to the confidence output layer” includes a configuration in which the input / output conversion function is a function that outputs the same value as the value input to the mapping layer.

[0032] At this time, when the number of quantization bits in the quantization layer (511) is given as n, and the maximum boundary value, which is the larger value among the two boundary values ​​used to determine the quantization parameter, is max(x), the minimum boundary value is min + (x) can be calculated as a solution to the simultaneous equations of Equations 5, 6, and 7 described later in this specification.

[0033] According to another aspect of the present invention, a method for determining a quantization parameter used in a neural network (50) including a confidence output layer (514), a mapping layer (513) outputting a third activation (713) input to the confidence output layer, a dequantization layer (512) outputting a second activation (712) input to the mapping layer, and a quantization layer (511) outputting a first activation (711) input to the dequantization layer may be provided. The method comprises the steps of: a computing device (2) determining a confidence threshold (Th) assigned to the confidence output layer; c ) obtaining a step of; the computing device, based on the input / output conversion function (f) of the mapping layer, outputs the confidence threshold value as the third activation, the input value (f) of the mapping layer -1 (Th c )) a step of determining; a step of determining, by the computing device, a boundary value for determining a quantization parameter (s1, z1) commonly used in the inverse quantization layer and the quantization layer based on the determined input value; and a step of determining, by the computing device, the quantization parameter commonly used in the quantization layer and the inverse quantization layer based on the determined boundary value.

[0034] At this time, the input / output conversion function is a monotonically increasing function in which the value output from the mapping layer increases as the value input to the mapping layer increases, and the determined boundary value is a minimum boundary value that is a smaller value among two boundary values ​​used to determine the quantization parameter, and the determined minimum boundary value can be selected from among values ​​smaller than the input value of the determined mapping layer.

[0035] At this time, the input / output conversion function is a monotonically decreasing function in which the value output from the mapping layer decreases as the value input to the mapping layer increases, and the determined boundary value is a maximum boundary value that is a larger value among two boundary values ​​used to determine the quantization parameter, and the determined maximum boundary value can be selected from among values ​​larger than the input value of the determined mapping layer.

[0036] According to another aspect of the present invention, a non-volatile computer-readable recording medium having recorded thereon a program including commands for executing a quantization parameter determination method for determining a quantization parameter used in a neural network (50) including a confidence output layer (514), a mapping layer (513) for outputting a third activation (713) input to the confidence output layer, a dequantization layer (512) for outputting a second activation (712) input to the mapping layer, and a quantization layer (511) for outputting a first activation (711) input to the dequantization layer may be provided. At this time, the quantization parameter determination method is configured to cause a computing device (2) to determine a confidence threshold value (Th) assigned to the confidence output layer. c ) obtaining a step of; the computing device, based on the input / output conversion function (f) of the mapping layer, outputs the confidence threshold value as the third activation, the input value (f) of the mapping layer -1 (Th c )) a step of determining; a step of determining, by the computing device, a boundary value for determining a quantization parameter (s1, z1) commonly used in the inverse quantization layer and the quantization layer based on the determined input value; and a step of determining, by the computing device, the quantization parameter commonly used in the quantization layer and the inverse quantization layer based on the determined boundary value.

[0037] According to another aspect of the present invention, a non-volatile computer-readable recording medium having recorded thereon a program including commands for executing an NPU command generation method that generates an NPU command that causes a first computing device (1) including an NPU (Neural Processing Unit) (110) to execute a neural network (50) including a confidence output layer (514), wherein the NPU command generation method causes the computing device (2) to generate a confidence threshold value (Th) assigned to the confidence output layer of the neural network. c ) obtaining step; the computing device, based on the input / output conversion function (f) of the mapping layer (513) that outputs the third activation (713) input to the confidence output layer, outputs the input value (f) of the mapping layer so that the mapping layer outputs the confidence threshold value as the third activation -1 (Th c )); a step of determining, by the computing device, a boundary value for determining a quantization parameter (s1, z1) commonly used in a dequantization layer (512) that outputs a second activation (712) input to the mapping layer and a quantization layer (511) that outputs a first activation (711) input to the dequantization layer, based on the determined input value; a step of determining, by the computing device, the quantization parameter commonly used in the quantization layer and the dequantization layer based on the determined boundary value; and a command generation step of generating, by the computing device, a command set that causes the NPU to execute the neural network, wherein the command set includes the determined quantization parameter.

[0038] According to the present invention, in the process of determining a quantization parameter used in the last quantization layer located upstream from the confidence output layer of a neural network, a technique for reducing a quantization error caused by a quantization operation using the quantization parameter can be provided by taking into account a confidence threshold assigned to the confidence output layer.

[0039] FIG. 1 illustrates an example of the structure of a neural network processed by a user computing device provided according to one embodiment of the present invention.

[0040] Figure 2 shows an example of the internal structure of the last quantization layer presented in Figure 1.

[0041] Figure 3 shows another example of the internal structure of the last quantization layer presented in Figure 1.

[0042] FIG. 4 illustrates an example of a histogram that the first tensor, the first activation, and the second activation described in FIGS. 1 to 3 may have.

[0043] Figure 5 shows the input / output relationship of the mapping layer and the restriction of the value of the neural network output data by the confidence output layer.

[0044] Figure 6 illustrates the process of tracing backwards from the downstream layer of the neural network to the upstream layer, the values ​​that can be discarded in the first tensor, the first activation, the second activation, and the third activation.

[0045] FIG. 7 illustrates a histogram of a range of values ​​that can be discarded in the first activation and the first tensor according to one embodiment of the present invention.

[0046] FIG. 8 illustrates a histogram of activations observed downstream of x when a correction minimum boundary value of x for determining a quantization parameter is determined by considering a confidence threshold value set in a confidence output layer according to one embodiment of the present invention.

[0047] FIG. 9 is a diagram for comparing the difference between the case where the minimum boundary value selected for determining the quantization parameter is determined according to a comparative example and the case where the minimum boundary value is determined according to an embodiment of the present invention.

[0048] Figures 10a to 10d are intended to explain various modified examples of the present invention.

[0049] FIG. 11 illustrates the main structure of a computing device for developers and a computing device for users that execute a method for neural network operations according to one embodiment of the present invention.

[0050] FIG. 12 illustrates a concept of a user computing device obtaining a command file executed by an NPU according to one embodiment of the present invention.

[0051] FIG. 13 is a flowchart illustrating a method for generating a command set for dedicated hardware including an NPU according to one embodiment of the present invention.

[0052] FIG. 14 is a flowchart illustrating a quantization parameter determination method for determining quantization parameters of a neural network according to one embodiment of the present invention.

[0053] Hereinafter, embodiments of the present invention will be described with reference to the attached drawings. However, the present invention is not limited to the embodiments described herein and may be implemented in various other forms. The terminology used herein is intended to aid understanding of the embodiments and is not intended to limit the scope of the present invention. Furthermore, the singular forms used below also include the plural forms, unless the context clearly indicates otherwise.

[0054] FIG. 1 illustrates an example of the structure of a neural network processed by a user computing device provided according to one embodiment of the present invention.

[0055] In the above user computing device, the part that processes the function of the neural network may have limited computing resources. For example, the part that processes the function of the neural network may be a dedicated hardware accelerator, such as an NPU, that executes neural network operations. In order to make computationally efficient, the NPU may limit the data input to a specific operation, such as a convolution operation, to INT type data. Accordingly, the NPU may be configured to execute an operation that quantizes FP type data into INT type data and an operation that dequantizes INT type data into FP type data. The NPU may obtain quantization parameters for executing the quantization and dequantization from a command file for executing a neural network provided to the NPU.

[0056] The above NPU may not limit data input to operations other than convolution operations to INT type.

[0057] The neural network (50) can receive neural network input data (70) and output neural network output data (o) (80).

[0058] The neural network (50) may include multiple layers.

[0059] At least some of the layers among the above-described plurality of layers may be configured to perform quantization and / or dequantization of data within the layer. In this specification, a layer that outputs quantized activations expressed as INT type may be referred to as a 'quantization layer (q) (510)'. At least some of the quantization layers (q) (510) may be configured to perform a convolution operation function within the layer.

[0060] In this specification, among two consecutive layers included in a neural network, the first layer that provides data may be referred to as an upstream layer, and the second layer that receives data provided by the first layer may be referred to as a downstream layer.

[0061] In this specification, the quantization layer (q) that is the most downstream among the plurality of quantization layers (q) (510) is indicated by reference numeral 511 and may be referred to as the 'last quantization layer (511)'.

[0062] The data input to the last quantization layer (511) may be referred to as input activation (710), and the data output from the last quantization layer (511) may be referred to as first activation (711). The first activation (711) may be a scalar expressed as an INT type. In this specification, the first activation (711) is denoted by the symbol x q can be displayed as

[0063] The quantization parameters {quantization scale, zero point} used by the last quantization layer (511) to generate the first activation (711) can be expressed as {s1, z1}, respectively.

[0064] The first activation (711) output by the last quantization layer (511) can be input to the inverse quantization layer (fp) (512). The inverse quantization layer (512) can obtain the values ​​of s1 and z1, and the inverse quantization layer (512) can inverse quantize the first activation (711) into the second activation (712) using the obtained quantization parameters {s1, z1}. The second activation (712) can be a scalar expressed as an FP type. In this specification, the second activation (712) is denoted by the sign x fp can be displayed as

[0065] The above computing device (1) can provide the quantization parameters {s1, z1} used by the last quantization layer (511) to the inverse quantization layer (512).

[0066] The inverse quantization layer (512) may be a layer that performs the function of converting (recovering) input data of the INT type into output data of the FP type. At this time, the recovery may be performed using the quantization parameters used in the process of quantizing the input data.

[0067] The second activation (712) may be input to a predetermined mapping layer (513). The mapping layer (513) may perform a function of converting the second activation (712) into a third activation (713). The third activation (713) may be a scalar expressed in the FP type. In this specification, the third activation (713) may be expressed by the symbol y.

[0068] In one embodiment, the mapping layer (513) can output the input data as output data. In this case, the mapping layer (513) can be considered to not exist.

[0069] In Fig. 1, the mapping layer (513) is represented as a single layer, but depending on the embodiment, the mapping layer (513) may include multiple consecutive layers. That is, the mapping layer (513) presented in Fig. 1 is an integrated representation of the layers existing between the inverse quantization layer (512) and the confidence output layer (514).

[0070] A function that defines the input-output relationship (input-output transfer function) between the second activation (712) input to the mapping layer (513) and the third activation (713) output from the mapping layer (513) can be represented by the symbol f. There is no limitation to the above function, but for the convenience of the following explanation, it is assumed that the function f is a sigmoid function.

[0071] The function f provided by the mapping layer (513) may be a function that maps the value of the second activation (712) to a value in a limited range.

[0072] For example, if the above function f is a sigmoid function, the value of the second activation (712) can be mapped to a value greater than 0 and less than 1. Accordingly, the value of the third activation (713) has a value greater than 0 and less than 1.

[0073] The third activation (713) can be input to the confidence output layer (514).

[0074] The confidence output layer (514) may be configured to receive the third activation (713) as input and output neural network output data (80).

[0075] In one embodiment, the confidence output layer (514) may not be the output layer of the neural network (80). That is, there may be other layers within the neural network (80) that are located downstream from the confidence output layer (514). However, for the convenience of the following description, the following description will focus on the case where the confidence output layer (514) is the output layer of the neural network (80).

[0076] In one embodiment, the confidence output layer (514) determines whether the value of the third activation (713) is greater than a certain confidence threshold Th c In this case, the neural network output data (80) is output, and the value of the third activation (713) is greater than the above confidence threshold Th c In the case where the value of the third activation (713) is less than a certain confidence threshold Th, the neural network output data (80) is not output. Conversely, in another embodiment, the confidence output layer (514) is configured to output the value of the third activation (713) less than a certain confidence threshold Th c In the case below, the neural network output data (80) is output, and the value of the third activation (713) is greater than the above confidence threshold Th cIn larger cases, the neural network output data (80) may not be output. However, for convenience of explanation below, the confidence output layer (514) is configured so that the value of the third activation (713) is a predetermined confidence threshold Th c This explanation is based on the case where neural network output data (80) is output only in the above case.

[0077] At this time, the value of the third activation (713) is a certain confidence threshold Th c In this case, the neural network output data (80) has the same value as the third activation (713). Therefore, the neural network output data (80) once output may be a scalar expressed in the FP type. In this specification, the neural network output data (80) may be expressed by the symbol t.

[0078] Figure 2 shows an example of the internal structure of the last quantization layer (511) presented in Figure 1.

[0079] The last quantization layer (511) may include a quantization unit (501) that quantizes a first tensor (701) expressed as an FP type with quantization parameters {s1, z1} to output a first activation (711) of an INT type. The first tensor (701) may be a scalar expressed as an FP type. In this specification, the first tensor (701) may be represented by the symbol x.

[0080] Figure 3 shows another example of the internal structure of the last quantization layer (511) presented in Figure 1.

[0081] The last quantization layer (511) may receive input activation (710) expressed as an INT type. The input activation (710) may be quantized by quantization parameters {s_in, z_in} in a layer upstream from the last quantization layer (511). The computing device (1) may be configured to provide the quantization parameters {s_in, z_in} to the last quantization layer (511).

[0082] The final quantization layer (511) may include a computation unit (503). The computation unit (503) may be configured to output a second tensor (702) expressed as an INT type. For example, the computation unit (503) may be a computation unit that convolves an activation expressed as an INT type and a weight expressed as an INT type.

[0083] The last quantization layer (511) may include a dequantization unit (502). The dequantization unit (502) may dequantize the second tensor (702) using the quantization parameters {s_in, z_in} to output a first tensor (701). The first tensor (701) may be a scalar expressed in the FP type. In this specification, the first tensor (701) may be represented by the symbol x.

[0084] The last quantization layer (511) may include a quantization unit (501) that quantizes the first tensor (701) expressed as an FP type by quantization parameters {s1, z1} to output a first activation (711) of an INT type. As described above, in this specification, the first activation (711) is a sign x q can be displayed as

[0085] Although FIG. 2 and FIG. 3 are different embodiments of the last quantization layer (511) presented in FIG. 1, they are the same in that they include a quantization unit (501) that quantizes the first tensor (701) with quantization parameters {s1, z1} and outputs a first activation (711) of INT type.

[0086] FIG. 4 illustrates an example of a histogram that the first tensor (701), the first activation (711), and the second activation (712) described in FIGS. 1 to 3 may have.

[0087] Since the values ​​of the neural network input data (70) input to the neural network (50) may vary, the values ​​of the first tensor (701), the first activation (711), and the second activation (712) observed when different neural network input data (70) are input to the neural network (50) may also vary. Accordingly, the histogram (H701) of the first tensor (701), the histogram (H711) of the first activation (711), and the histogram (H712) of the second activation (712) may be determined based on the various values ​​that the first tensor (701), the first activation (711), and the second activation (712) may have.

[0088] The upper part of Fig. 4 shows an example of a histogram (H701) of the first tensor (701), the middle part shows an example of a histogram (H711) of the first activation (711), and the lower part shows an example of a histogram (H712) of the second activation (712).

[0089] Based on the distribution of the first tensor (701), the maximum boundary value (max(x)) and the minimum boundary value (min(x)) of the first tensor (701) can be determined to determine s1, which is an s value following [Formula 2]. Then, based on the determined maximum boundary value (max(x)) and minimum boundary value (min(x)), the quantization parameters {s1, z1} following [Formula 2] and [Formula 3] can be determined.

[0090] By quantization performed based on the above maximum boundary value (max(x)) and minimum boundary value (min(x)), all values ​​of x smaller than the minimum boundary value (min(x)) are clipped to have the minimum boundary value (min(x)), and all values ​​of x larger than the maximum boundary value (max(x)) are clipped to have the maximum boundary value (max(x)). Although these clipped values ​​have a disadvantage in that they themselves have an error (clipping error), the number of clipped values ​​is small, and despite the clipping error, if the difference between the maximum boundary value (max(x)) and the minimum boundary value (min(x)) is reduced, the quantization scale s1 of [Formula 2] increases, so that the value x quantized by [Formula 1] q It also has the advantage of reducing round errors.

[0091] Looking at the histogram (H711) of the first activation (711), x is due to the clipped values ​​by the quantization process. q The minimum value of (min(x) q )) and maximum(max(x q )) can be confirmed to increase compared to the frequency of the minimum boundary value (min(x)) and maximum boundary value (max(x)) of x.

[0092] The second activation (712) is converted to FP type by dequantizing the first activation (711), so the histogram (H712) of the second activation (712) is identical to the histogram (H711) of the first activation (711). That is, x q The minimum value of (min(x) q )) and maximum(max(x q )) is x fp The minimum value of (min(x) fp )) and maximum(max(x fp )) is the same as

[0093] Figure 5 shows the input / output relationship of the mapping layer (513) and the limitation of the value of the neural network output data (80) by the confidence output layer (514).

[0094] Figure 5 shows an example in which the input / output conversion function (=transfer function) f of the mapping layer (513) is a sigmoid function.

[0095] In the graph of Figure 5, the horizontal axis represents the second activation (x) input to the mapping layer (513). fp )(712), and the vertical axis represents the value of the third activation (713) output by the mapping layer (513).

[0096] The confidence threshold set for the confidence output layer (514) is Th c In this case, the confidence output layer (514) is set to the above confidence threshold Th c The third activation (713) with a larger value than the above is output as neural network output data (80), but the confidence output layer (514) outputs the above confidence threshold Th c The third activation (713) with a smaller value is discarded without being output.

[0097] That is, the inverse function of the input / output transformation function f of the mapping layer (513) is f -1 When said, the computing device (1) outputs the mapping layer (513) Th c When , the input value f of the corresponding mapping layer (513) is given -1 (Th c ) with x fp can find the value of .

[0098] For example, the input / output transformation function f of the mapping layer (513) is a sigmoid function, and the confidence threshold Th c Assuming that f is 0.05, -1 (Th c ) = sigmoid -1 (0.05) = -2.945. Therefore, x fp If y is less than -2.945, then y=f(x fp ) is not output by the confidence output layer (514) and is discarded. Conversely, x fpIf y is greater than or equal to -2.945, then y=f(x fp ) is output as neural network output data (80) by the confidence output layer (514).

[0099] That is, f -1 (Th c ) has a value less than the value of x fp Since it is a value whose output is blocked by the confidence output layer (514), it can be considered as an unnecessary discarded value.

[0100] Figure 6 illustrates a process of tracing backwards from the downstream layer to the upstream layer of the neural network (50) the values ​​that can be discarded in the first tensor (701), the first activation (711), the second activation (712), and the third activation (713).

[0101] Figure 6 illustrates an example in which the input / output conversion function f of the mapping layer (513) is a sigmoid function.

[0102] The confidence threshold Th among the third activations (713) indicated by the symbol y c Smaller values ​​can be discarded.

[0103] So the reference symbol x fp Among the second activations (712) marked as f -1 (Th c ) can be discarded. This is because the input-output conversion function f is a monotonically increasing function.

[0104] Likewise, reference mark x q Among the first activations (711) marked as f -1 (Th c )*s1+z1 are values ​​that can be discarded.

[0105] Finally, the first tensor (701) denoted by reference symbol x is the above f -1 (Th c )*s1+z1 corresponding value x Th Smaller values ​​can be discarded.

[0106] The concept of values ​​that can be discarded among the first activation (711) and the first tensor (701) is presented in Fig. 7.

[0107] FIG. 7 illustrates a range of values ​​that can be discarded in the first activation (711) and the first tensor (701) on a histogram according to one embodiment of the present invention.

[0108] As shown in the upper part of Fig. 7, the symbol x q Among the first activations (711) marked as f -1 (Th c )*s1+z1 are values ​​that can be discarded.

[0109] Therefore, x expressed as a quantized INT type presented in the upper part of Fig. 7 q In the quantization process to generate x, q =f -1 (Th c )*s1+z1 corresponding value x=x Th For values ​​of x smaller than x=x Th The 'corrected minimum boundary (min)' is the minimum boundary value selected from the values ​​of x smaller than + There is no problem even if it is clipped to (x)).

[0110] The above 'Correction minimum boundary value (min + (x))' has a clipping error. Therefore, x=x Th When the minimum value that is not discarded by the confident output layer is said to be the 'correction minimum boundary value (min + (x))' is x Th Smaller is preferable.

[0111] In Fig. 4, the minimum boundary value (min(x)) of the first tensor (701) for determining s1, which is a quantization scale following [Equation 2], is determined based on the distribution of the first tensor (701). In contrast, according to one embodiment of the present invention, if the values ​​discarded in the confidence output layer (514) are further considered, the minimum boundary value of the first tensor (701) for determining s1, which is a quantization scale following [Equation 2], is determined as a corrected minimum boundary value (min(x)) greater than the minimum boundary value (min(x)) of Fig. 4. + I can understand that it can be changed to (x)).

[0112] The above correction minimum boundary value (min) + (x)) is the above x=x Th can be selected from values ​​of x smaller than, preferably, x=x Th The largest value among the smaller values ​​of x can be selected.

[0113] Now, according to a preferred embodiment, the correction minimum boundary value (min + Explains how to determine (x)).

[0114] Assume that asymmetric quantization is performed with n bits. In this case, the quantization layer (511) has (-2 n-1 ) ~ (2 n-1 -1) Integer values ​​in the range are output. At this time, the determined f output from the inverse quantization layer (512) -1 (Th c ) must be able to express values ​​smaller than f. -1 (Th c ) is the second smallest number that the quantization layer (511) can output, which is -2. n-1 can be assigned to +1. At this time, x fp =f -1 (Th c ) corresponding to x q The value of f -1 (Th c )*s1+z1 is assigned to -2 n-1To map to +1, s1 and z1 must satisfy [Formula 4] below.

[0115] [Formula 4]

[0116] -2 n-1 +1-0.5 <= f -1 (Th c )*s1+z1 < -2 n-1 +1+0.5

[0117] Considering Equation 4, in a preferred embodiment of the present invention, [Equation 5] can be satisfied.

[0118] [Formula 5]

[0119] -2 n-1 +0.5 = f -1 (Th c )*s1+z1

[0120] Now, Equations 2 and 3 can be rewritten as Equations 6 and 7 as follows:

[0121] [Formula 6]

[0122] s1 = (2 n - 1) / (max(x) - min + (x))

[0123] [Formula 7]

[0124] z1 = - min + (x) * s1 - 2 n-1

[0125] Now, if we look at equations 5, 6, and 7 together, the unknowns are s1, z, and min + (x) is 3, max(x), and n are already given values. Therefore, the values ​​of the three unknowns can be determined by three formulas (Formula 5, Formula 6, Formula 7). As a result, min + The value of (x) can be determined.

[0126] min determined using Equations 5, 6, and 7 +(x) may be an optimal parameter that can minimize quantization error by the quantization layer (511).

[0127] As an example, if the function f of the mapping layer (513) is a sigmoid function, and 8-bit quantization is performed in the quantization layer (511) and the inverse quantization layer (512), and the confidence threshold Th of the confidence output layer (514) c Describes the case where the confidence threshold Th of the confidence output layer (514) is 0.05. c If this is 0.05, the input value f of the corresponding mapping layer (513) -1 (Th c )=f -1 (0.05)=-2.945. Therefore, the confidence threshold Th c = The output of the inverse quantization layer (512) corresponding to 0.05 is -2.945. At this time, when 8-bit quantization is performed in the quantization layer (511), the range of the output value is an integer from -128 to 127. In order to execute the confidence threshold in the range of the output of the quantization layer (511), values ​​smaller than -2.945 must be expressible, so -2.945 can be assigned to -127, which is the second smallest number in the range of the output of the quantization layer (511). Accordingly, the quantization parameters s1 and z1 commonly used in the quantization layer (511) and the inverse quantization layer (512) can be limited to the range of -127.5 <= -2.945*s1 + z1 < -126.5. At this time, when -2.945*s1 + z1 = -127.5 is set, all values ​​less than -2.945 are set as thresholds.

[0128] At this time, max(x) is min + Assuming that (x) and (x) have different signs and the same absolute values, we can use Equations 5, 6, and 7 above. In this case, if we solve the simultaneous equations in Equation 8 below, we get min + (x) is determined as -2.956.

[0129] [Formula 8]

[0130] -2 7 +0.5 = f -1 (0.05)*s1+z1 = -127.5 = -2.945*s1+z1

[0131] s1 = (2 8 - 1) / (- 2*min + (x)) = 255 / (- 2*min + (x)),

[0132] z1 = -min + (x)*s1 - 2 7 = -min + (x)*s1 - 128

[0133] Figure 8 illustrates a confidence threshold (Th) set in a confidence output layer (514) according to one embodiment of the present invention. c ), the minimum bound of correction of x for determining the quantization parameter (min + When (x)) is determined, a histogram of the activities observed downstream of x is shown.

[0134] The histogram presented in Fig. 8 can be compared with the histogram presented in Fig. 4. The histogram presented in Fig. 8 represents an example in which selected quantization parameters are applied according to one embodiment of the present invention, and the histogram presented in Fig. 4 represents a comparative example.

[0135] The upper part of Fig. 8 shows an example of a histogram (H701) of the first tensor (701), the middle part shows an example of a histogram (H711) of the first activation (711), and the lower part shows an example of a histogram (H712) of the second activation (712).

[0136] The maximum and minimum boundary values ​​of the first tensor (701) for determining the quantization scale s1 according to the above [Formula 2] are ① the distribution of the first tensor (701) and ② the given confidence threshold value (Th c) and the input / output transfer function f of the mapping layer (513) can be determined.

[0137] Specifically, the minimum boundary value of the first tensor (701) is the correction minimum boundary value (min + (x)) is x q =f -1 (Th c )*s1+z1 corresponding value x=x Th It can be selected from the values ​​of x that are smaller than . And the maximum boundary value (max(x)), which is the minimum boundary value of the first tensor (701), can be determined based on the distribution of the first tensor (701).

[0138] The above determined minimum correction boundary value (min + Based on the (x)) and maximum boundary value (max(x)), the quantization parameters {s1, z1} following [Formula 2] and [Formula 3] can be determined.

[0139] The above maximum boundary value (max(x)) and the corrected minimum boundary value (min + By quantization based on (x)), the correction minimum boundary value (min) + All values ​​of x less than (x) are within the minimum correction boundary (min + (x)) and all values ​​of x greater than the maximum boundary value (max(x)) are clipped to the maximum boundary value (max(x)).

[0140] Looking at the histogram (H711) of the first activation (711), x is due to the clipped values ​​by the quantization process. q The minimum value of (min) + (x q )) and maximum(max(x q )) is the frequency of the corrected minimum boundary value of x (min + It can be seen that the frequency of (x)) and the maximum boundary value (max(x)) increases.

[0141] The second activation (712) is converted to FP type by dequantizing the first activation (711), so the histogram (H712) of the second activation (712) is identical to the histogram (H711) of the first activation (711). That is, x q The minimum value of (min) + (x q )) and maximum(max(x q )) is x fp The minimum value of (min) + (x fp )) and maximum(max(x fp )) is practically identical to.

[0142] x fp is the minimum value (min) + (x fp )) if we have x fp Although it has the disadvantage that it is very likely to have clipping errors due to the above quantization process, the minimum value (min + (x fp )) with x fp This is not a problem because it is a value discarded by the confidence output layer (514).

[0143] FIG. 9 is a diagram for comparing the difference between the case where the minimum boundary value selected for determining the quantization parameter is determined according to a comparative example and the case where the minimum boundary value is determined according to an embodiment of the present invention.

[0144] The upper part of Fig. 9 presents the clipping strategy presented in Fig. 4, which represents a comparative example, and the lower part of Fig. 9 presents the clipping strategy presented in Fig. 8, which represents an embodiment of the present invention. According to an embodiment of the present invention, a selected correction minimum boundary value (min + It can be confirmed that (x)) is greater than the minimum boundary value (min(x)) selected according to the comparative example.

[0145] As described above, among the quantization layers constituting the neural network (50), the last quantization layer (511) located at the lowest point and the dequantization layer (512) that dequantizes the INT type activation output by the last quantization layer (511) to output an FP type activation, the quantization parameters {s1, z1} commonly used can be determined in a developer computing device. The developer computing device may be a general-purpose computer.

[0146] The developer computing device must know in advance the structure of the neural network presented in Fig. 1 in order to determine the values ​​of the quantization parameters {s1, z1} used by the last quantization layer (511) and the inverse quantization layer (512). Specifically, the developer computing device must know the confidence threshold (Th) assigned to the confidence output layer (514). c ) and the input / output transfer function f of the mapping layer (513) must be obtained, and the values ​​of the quantization parameters {s1, z1} can be determined using this.

[0147] The determined quantization parameters {s1, z1} may be provided from the developer computing device to the user computing device. The user computing device may include the NPU described above.

[0148] Figures 10a and 10b are intended to explain various modified examples of the present invention.

[0149] In Figures 5 to 9, the confidence output layer (514) sets the confidence threshold (Th c ) is discarded, and an example is shown in which the input / output function of the mapping layer (513) is a monotonically increasing function.

[0150] In contrast, the confidence output layer (514) has a confidence threshold (Th c ) is also possible, and an embodiment in which the input / output function of the mapping layer (513) is a monotonically decreasing function is also possible.

[0151] Figure 10a shows that the confidence output layer (514) sets a confidence threshold (Th c ) is set to discard values ​​smaller than , and an example is shown in which the input / output function of the mapping layer (513) is a monotonically increasing function.

[0152] Figure 10b shows that the confidence output layer (514) sets a confidence threshold (Th c ) is set to discard values ​​smaller than , and an example is shown in which the input / output function of the mapping layer (513) is a monotonically decreasing function.

[0153] Figure 10c shows the confidence output layer (514) with a confidence threshold (Th c ) is set to discard values ​​greater than , and an example is shown in which the input / output function of the mapping layer (513) is a monotonically increasing function.

[0154] Figure 10d shows the confidence output layer (514) with a confidence threshold (Th c ) is set to discard values ​​greater than , and an example is shown in which the input / output function of the mapping layer (513) is a monotonically decreasing function.

[0155] Figure 10a is a re-presentation of the examples presented in Figures 5 to 9.

[0156] In the case of Fig. 10a and Fig. 10d, the minimum boundary value (min) is the smaller of the two boundary values ​​(min, max) used to determine the quantization parameters {s1, z1} commonly used in the last quantization layer (511) and the inverse quantization layer (512). + In the process of determining the confidence threshold (Th) c ) is used. The above min + (x) may also be referred to as the corrected minimum boundary value or the corrected minimum boundary value.

[0157] In contrast, in the case of FIGS. 10b and 10c, the maximum boundary value (max) is the larger value among the two boundary values ​​(min, max) used to determine the quantization parameters {s1, z1} commonly used in the last quantization layer (511) and the inverse quantization layer (512). - In the process of determining the confidence threshold (Th) c ) is used. The above max - (x) may also be referred to as the corrected maximum boundary value or the corrected maximum boundary value.

[0158] FIG. 11 illustrates the main structure of a computing device for developers and a computing device for users that execute a method for neural network operations according to one embodiment of the present invention.

[0159] The user computing device (1) presented in FIG. 11 may be, for example, a desktop computer, a laptop computer, a smartphone, and a tablet.

[0160] A computing device (1) may include a DRAM (Dynamic Random Access Memory) (130), an NPU (110), a bus (700) connecting the DRAM (130) and the NPU (110), and other hardware (99) connected to the bus (700), a main processor (160), and a storage unit (170).

[0161] The NPU (110) may also be referred to as a hardware accelerator.

[0162] In addition, the computing device (1) may further include a power supply unit, a communication unit, a user interface, and peripheral units that are not shown. The bus (700) may be shared by the NPU (110), other hardware (99), and the main processor (160).

[0163] The above storage unit (170) may be integrally connected to the computing device (1) or may be detachably connected.

[0164] The above NPU (110) may include a DMA (Direct Memory Access) part (20), a control unit (40), an internal memory (30), an input buffer (650), a data operation unit (610), and an output buffer (640).

[0165] Some or all of the data temporarily stored in the internal memory (30) may be provided from the DRAM (130) via the bus (700). At this time, in order to move the data stored in the DRAM (130) to the internal memory (30), the control unit (40) and the DMA unit (20) may control the internal memory (30) and the DRAM (130).

[0166] In this specification, DRAM (130) may also be referred to as external memory.

[0167] Data stored in the internal memory (30) can be provided to the data operation unit (610) through the input buffer (650).

[0168] The output values ​​generated by the operation of the above data operation unit (610) can be stored in the internal memory (30) via the output buffer (640). The output values ​​stored in the internal memory (30) can also be written to the DRAM (130) under the control of the control unit (40) and the DMA unit (20).

[0169] The control unit (40) can comprehensively control the operation of resources within the NPU (110), such as the DMA unit (20), internal memory (30), and the data operation unit (610).

[0170] In one implementation example, the data operation unit (610) may perform a first operation function during a first time period and a second operation function during a second time period. For example, the data operation unit (610) may perform a first operation function according to the operation rule of a first layer of a neural network during a first time period and a second operation function according to the operation rule of a second layer of a neural network during a second time period.

[0171] In Fig. 11, one data operation unit (610) is provided within the NPU (110). However, in a modified embodiment not shown, a plurality of data operation units (610) shown in Fig. 11 may be provided within the NPU (110) to perform operations requested by the control unit (40) in parallel.

[0172] In one implementation example, the data operation unit (610) may output the output data sequentially in a given order over time, rather than outputting the output data all at once.

[0173] The computing device (2) for developers presented in Fig. 11 may be, for example, a server, a desktop computer, or a laptop computer. The computing device (2) may include a DRAM (230), a bus (2700), other hardware (299), a main processor (260), and a storage unit (270).

[0174] FIG. 12 illustrates a concept of a user computing device obtaining a command file executed by an NPU according to one embodiment of the present invention.

[0175] In this specification, a computing device (1) for a user may be referred to as a first computing device, and a computing device (2) for a developer may be referred to as a second computing device.

[0176] In one example, a user computing device (1) can obtain a command file to be executed by an NPU from a developer computing device (2) through a predetermined communication channel.

[0177] In another example, a user computing device (1) can obtain a command file to be executed by an NPU from a developer computing device (2) through a predetermined communication channel via a relay device (3). The relay device (3) may be a production device that is used in the production process of the user computing device (1).

[0178] FIG. 13 is a flowchart illustrating a method for generating a command set for dedicated hardware including an NPU according to one embodiment of the present invention.

[0179] The method shown in Fig. 13 is an NPU command generation method that generates an NPU command that causes a first computing device (1) including an NPU (Neural Processing Unit) (110) to execute a neural network (50) including a confidence output layer (514).

[0180] In step (S10), the computing device (2) assigns a confidence threshold (Th) to the confidence output layer of the neural network. c ) can be obtained.

[0181] In step (S20), the computing device outputs the third activation (713) input to the confidence output layer based on the input / output conversion function (f) of the mapping layer (513), which outputs the third activation (713) input to the confidence output layer, an input value (f) of the mapping layer that causes the mapping layer to output the confidence threshold value as the third activation. -1 (Th c )) can be decided.

[0182] In step (S30), the computing device can determine boundary values ​​for determining quantization parameters (s1, z1) commonly used in the dequantization layer (512) that outputs the second activation (712) input to the mapping layer and the quantization layer (511) that outputs the first activation (711) input to the dequantization layer, based on the determined input values.

[0183] In step (S40), the computing device can determine the quantization parameter commonly used in the quantization layer and the inverse quantization layer based on the determined boundary value.

[0184] In step (S50), the computing device may generate a set of commands that cause the NPU to execute the neural network. At this time, the generated set of commands may include the determined quantization parameters.

[0185] FIG. 14 is a flowchart illustrating a quantization parameter determination method for determining quantization parameters of a neural network according to one embodiment of the present invention.

[0186] The above quantization parameter determination method is a method for determining a quantization parameter used in a neural network (50) including a confidence output layer (514), a mapping layer (513) that outputs a third activation (713) input to the confidence output layer, a dequantization layer (512) that outputs a second activation (712) input to the mapping layer, and a quantization layer (511) that outputs a first activation (711) input to the dequantization layer.

[0187] In step (S110), the computing device (2) assigns a confidence threshold (Th) to the confidence output layer. c ) can be obtained.

[0188] In step (S120), the computing device outputs the input value (f) of the mapping layer so that the mapping layer outputs the confidence threshold value as the third activation based on the input / output conversion function (f) of the mapping layer. -1 (Th c )) can be decided.

[0189] In step (S130), the computing device can determine a boundary value for determining a quantization parameter (s1, z1) commonly used in the inverse quantization layer and the quantization layer based on the determined input value.

[0190] In step (S140), the computing device can determine the quantization parameter commonly used in the quantization layer and the inverse quantization layer based on the determined boundary value.

[0191] At this time, the confidence output layer is configured to discard values ​​smaller than the confidence threshold value, the input / output transformation function is a monotonically increasing function in which the value output from the mapping layer increases as the value input to the mapping layer increases, and the determined boundary value is a minimum boundary value that is a smaller value among two boundary values ​​used to determine the quantization parameter, and the determined minimum boundary value may be selected from among values ​​smaller than the input value of the determined mapping layer. This case corresponds to Fig. 10a.

[0192] Alternatively, the confidence output layer may be configured to discard values ​​smaller than the confidence threshold value, the input / output transformation function may be a monotonically decreasing function in which the value output from the mapping layer decreases as the value input to the mapping layer increases, and the determined boundary value may be a maximum boundary value that is a larger value among two boundary values ​​used to determine the quantization parameter, and the determined maximum boundary value may be selected from among values ​​larger than the input value of the determined mapping layer. This case corresponds to FIG. 10b.

[0193] Alternatively, the confidence output layer may be configured to discard values ​​greater than the confidence threshold value, the input / output transformation function may be a monotonically increasing function in which the value output from the mapping layer increases as the value input to the mapping layer increases, and the determined boundary value may be a maximum boundary value that is a larger value among two boundary values ​​used to determine the quantization parameter, and the determined maximum boundary value may be selected from among values ​​greater than the input value of the determined mapping layer. This case corresponds to FIG. 10c.

[0194] Alternatively, the confidence output layer may be configured to discard values ​​greater than the confidence threshold value, and the input / output transformation function may be a monotonically decreasing function in which the value output from the mapping layer decreases as the value input to the mapping layer increases, and the determined boundary value may be a minimum boundary value that is a smaller value among two boundary values ​​used to determine the quantization parameter, and the determined minimum boundary value may be selected from among values ​​smaller than the input value of the determined mapping layer. This case corresponds to FIG. 10d.

[0195] According to another embodiment of the present invention, a non-volatile computer-readable recording medium having recorded thereon a program including commands for executing a quantization parameter determination method for determining a quantization parameter used in a neural network (50) including a confidence output layer (514), a mapping layer (513) for outputting a third activation (713) input to the confidence output layer, a dequantization layer (512) for outputting a second activation (712) input to the mapping layer, and a quantization layer (511) for outputting a first activation (711) input to the dequantization layer may be provided. At this time, the quantization parameter determination method is configured such that a computing device (2) determines a confidence threshold value (Th) assigned to the confidence output layer. c ) obtaining a step of; the computing device, based on the input / output conversion function (f) of the mapping layer, outputs the confidence threshold value as the third activation, the input value (f) of the mapping layer -1 (Th c)) a step of determining; a step of the computing device determining a boundary value for determining a quantization parameter (s1, z1) commonly used in the inverse quantization layer and the quantization layer based on the determined input value; and a step of the computing device determining the quantization parameter commonly used in the quantization layer and the inverse quantization layer based on the determined boundary value. The non-volatile computer-readable recording medium may be, for example, a storage unit (270) of the developer-use computing device (2) shown in FIG. 11, or a storage medium accessible by the developer-use computing device (2), or another storage medium recording data transmitted from the storage medium.

[0196] By utilizing the embodiments of the present invention described above, those skilled in the art will be able to easily implement various changes and modifications without departing from the essential characteristics of the present invention. The content of each claim may be combined with other claims that are not in a citation relationship within the scope of this specification, as long as it is understood.

[0197] <Sasa>

[0198] This invention is a research project titled 'Development of Variable Precision High-Speed-Multiple Object Recognition Deep Learning Processor Technology' conducted by OpenEdge Technology Co., Ltd. as part of the 'Next-Generation Intelligent Semiconductor Technology Development (Design)', a national research and development project managed by the Ministry of Science and ICT and the National IT Industry Promotion Agency. The research project identification number is 20200031, the project number is 2020-0-01080, and the research period is from April 1, 2020 to December 31, 2024.

Claims

1. An NPU command generation method for generating an NPU command that causes a first computing device including an NPU to execute a neural network including a confidence output layer and a quantization layer located upstream of the confidence output layer, A step of the computing device obtaining a confidence threshold assigned to the confidence output layer of the neural network; A step of the computing device determining an input value of the mapping layer that causes the mapping layer to output the confidence threshold value as the third activation based on an input / output conversion function of the mapping layer that outputs the third activation input to the confidence output layer; A step for determining a boundary value used to determine a quantization parameter commonly used in the dequantization layer outputting the second activation inputted to the mapping layer and the quantization layer outputting the first activation inputted to the dequantization layer, based on the determined input value; and A command generation step in which the computing device generates a set of commands that cause the NPU to execute the neural network, the command set including the quantization parameters determined using the boundary values; Including, How to create an NPU command.

2. In paragraph 1, The above third activation and the above second activation are scalars expressed in FP (floating point) type, The above first activation is a scalar expressed as an INT (integer) type. How to create an NPU command.

3. A method for generating an NPU command, wherein the confidence output layer in the first paragraph is an output layer of the neural network.

4. A method for generating an NPU command in the first paragraph, wherein the confidence output layer is a layer located upstream from the output layer of the neural network.

5. An NPU command generation method in the first paragraph, wherein the input value of the mapping layer that outputs the confidence threshold value as the third activation is a value output by an inverse function of the input / output conversion function that inputs the confidence threshold value.

6. In paragraph 1, The above confidence output layer is designed to discard values that are less than the confidence threshold. The above input / output conversion function is a monotonically increasing function in which the value output from the mapping layer increases as the value input to the mapping layer increases. The above determined boundary value is the minimum boundary value, which is the smaller value among the two boundary values used to determine the quantization parameter. The determined minimum boundary value is characterized in that it is selected from among values smaller than the first value corresponding to the input value of the determined mapping layer among the range of values input to the quantization layer. How to create an NPU command.

7. In paragraph 6, The determined minimum boundary value is a maximum value among values smaller than the first value corresponding to the input value of the determined mapping layer among the range of values input to the quantization layer, such that the input value of the mapping layer does not include a clipping error caused by quantization of the quantization layer, or The determined minimum boundary value is a value that satisfies the condition that the difference between the first value corresponding to the input value of the determined mapping layer and the determined minimum boundary value among the range of values input to the quantization layer is greater than 0 and less than a predetermined value. How to create an NPU command.

8. In paragraph 1, The above confidence output layer is designed to discard values that are less than the confidence threshold. The above input / output conversion function is a monotonically decreasing function in which the value output from the mapping layer decreases as the value input to the mapping layer increases. The above determined boundary value is the maximum boundary value, which is the larger value among the two boundary values used to determine the quantization parameter. The determined maximum boundary value is characterized in that it is selected from among values greater than the first value corresponding to the input value of the determined mapping layer among the range of values input to the quantization layer. How to create an NPU command.

9. In paragraph 1, The above confidence output layer is designed to discard values greater than the confidence threshold. The above input / output conversion function is a monotonically increasing function in which the value output from the mapping layer increases as the value input to the mapping layer increases. The above determined boundary value is the maximum boundary value, which is the larger value among the two boundary values used to determine the quantization parameter. The determined maximum boundary value is characterized in that it is selected from among values greater than the first value corresponding to the input value of the determined mapping layer among the range of values input to the quantization layer. How to create an NPU command.

10. In paragraph 1, The above confidence output layer is designed to discard values greater than the confidence threshold. The above input / output conversion function is a monotonically decreasing function in which the value output from the mapping layer decreases as the value input to the mapping layer increases. The above determined boundary value is the minimum boundary value, which is the smaller value among the two boundary values used to determine the quantization parameter. The determined minimum boundary value is characterized in that it is selected from among values smaller than the first value corresponding to the input value of the determined mapping layer among the range of values input to the quantization layer. How to create an NPU command.

11. A step of the computing device obtaining a confidence threshold value assigned to a confidence output layer of a given neural network; A step of determining a boundary value used to determine a quantization parameter used in the last quantization layer existing upstream of the confidence output layer based on the confidence threshold value; and A step of the computing device determining the quantization parameter based on the determined boundary value; Including, How to determine quantization parameters.

12. A method for determining quantization parameters, wherein in the 11th paragraph, the computing device further comprises a command generation step of generating a set of commands that cause a predetermined NPU to execute the neural network, the set of commands including the quantization parameters determined using the boundary values.

13. In paragraph 11, The step of determining the above boundary value is: A step of the computing device determining an input value of the mapping layer that causes the mapping layer to output the confidence threshold value based on an input / output transformation function of the mapping layer existing between the confidence output layer and the quantization layer; and A step of the computing device determining the boundary value based on the determined input value; Including, How to determine quantization parameters.

14. In paragraph 11, The neural network includes a confidence output layer, a mapping layer that outputs a third activation input to the confidence output layer, a dequantization layer that outputs a second activation input to the mapping layer, and a quantization layer that outputs a first activation input to the dequantization layer. The above quantization parameters are commonly used in the inverse quantization layer and the quantization layer. The step of determining the above boundary value is: The computing device determines an input value of the mapping layer that causes the mapping layer to output the confidence threshold value as the third activation based on the input / output conversion function of the mapping layer; and A step of the computing device determining the boundary value based on the determined input value; Including, How to determine quantization parameters.

15. A computer-readable non-volatile storage medium having recorded thereon a program including commands for executing a quantization parameter determination method for determining quantization parameters used in a neural network including a confidence output layer and a last quantization layer existing upstream of the confidence output layer, The above quantization parameter determination method is, A step of the computing device obtaining a confidence threshold assigned to the confidence output layer; A step of determining a boundary value used to determine a quantization parameter used in the last quantization layer existing upstream of the confidence output layer based on the confidence threshold value; and A step of the computing device determining the quantization parameter based on the determined boundary value; Including, Non-volatile computer-readable recording medium

Citation Information

Patent Citations

  • Handwriting keyboard for screens

    KR1020190052667A

  • Manufacturing Apparatus For Secondary Battery

    KR1020210019192A

  • Apparatus and method to provide broadcast service in wireless communication system having shared base station

    KR1020230139523A

  • Quantization calibration method, computing device and computer readable storage medium

    US20230133337A1

  • KR20230022093A