Threshold Determination Program, Threshold Determination Method, and Threshold Determination Device
By determining thresholds based on quantization error statistics, the method enhances neural network quantization accuracy and efficiency, addressing the challenge of inference accuracy loss during threshold selection.
Patent Information
- Application Number
- JP2021136804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-08-25
AI Technical Summary
The challenge in neural network quantization is the difficulty in selecting an appropriate threshold for clipping numerical values, leading to a decrease in inference accuracy.
A method to determine the threshold value based on quantization error for each numerical value, using statistical values such as average, median, or maximum quantization error to select the most suitable threshold, thereby improving the accuracy of quantized values.
This approach maintains high inference accuracy even with reduced bit widths, reducing power consumption and memory usage by efficiently compressing neural networks.
Smart Images

Figure 0007700577000003 
Figure 0007700577000004 
Figure 0007700577000005
Abstract
Description
Technical Field
[0001] The present invention relates to threshold determination technology.
Background Art
[0002] A neural network, which is a type of learned model generated by machine learning, is used to make inferences on input data in various fields such as image processing and natural language processing (see, for example, Non-Patent Document 1 and Non-Patent Document 2).
[0003] Due to the complex configuration of neural networks in recent years, the power consumption of computers performing inferences using neural networks has a tendency to increase. Therefore, in order to reduce power consumption, quantization of neural networks may be performed. Quantization of a neural network is a process of converting a numerical value to be quantized, which is represented by a predetermined bit width, into a quantized numerical value represented by a smaller bit width.
[0004] Quantization of a neural network is effective in reducing power consumption and memory usage, but degrades the accuracy of the numerical values to be quantized. For example, when converting a 32-bit single-precision floating-point number (FP32) to an 8-bit integer (INT8) by quantization, the inference accuracy significantly decreases (see, for example, Non-Patent Document 3).
[0005] Regarding quantization of neural networks, technologies for promoting improvement in the efficiency of neural networks are known (see, for example, Patent Document 1). There is also known a learning device for a neural network that enables appropriate operations while reducing the weight of a CNN (Convolutional Neural Network) by reducing the number of bits of operations (see, for example, Patent Document 2). A method of adjusting the accuracy related to some selected layers of a neural network to a lower number of bits is also known (see, for example, Patent Document 3).
[0006] A sequence conversion model based on an attention mechanism is also known (see, for example, Non-Patent Document 4).
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0008]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
[0009] In the quantization of neural networks, it is important to select an appropriate scaling factor for converting the numerical values to be quantized into the quantized numerical values. The numerical values to be quantized are, for example, the weights of each of a plurality of edges between two layers of a neural network, the output values of each of a plurality of nodes included in each layer of the neural network, and the like. The output value of each node is called an activation. The plurality of numerical values to be quantized and the plurality of quantized numerical values may be represented by tensors.
[0010] By performing clipping on the numerical values to be quantized, the accuracy of the quantized numerical values may be improved. Clipping is a process of converting numerical values outside the numerical range defined by a threshold into the quantized numerical values corresponding to the threshold. However, it is difficult to select an appropriate threshold for clipping.
[0011] Note that such problems occur not only in the quantization of weights or activations, but also in the quantization of various numerical values in neural networks.
[0012] In one aspect, the present invention aims to suppress a decrease in inference accuracy due to quantization of a neural network. Means for Solving the Problems
[0013] In one approach, the threshold determination program causes a computer to execute the following processing.
[0014] When a computer converts a numerical value outside the numerical range defined by a threshold value among a plurality of numerical values to be quantized into a quantized numerical value corresponding to the threshold value in the quantization of a neural network, the computer determines the threshold value. At this time, the computer determines the threshold value based on the quantization error for each of the plurality of numerical values.
Advantages of the Invention
[0015] According to one aspect, it is possible to suppress a decrease in inference accuracy due to quantization of a neural network.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0017] Hereinafter, embodiments will be described in detail with reference to the drawings.
[0018] In the quantization of Non-Patent Document 3, when converting FP32 to INT8, the numerical range of FP32 is restricted by performing clipping before applying the scaling factor. In this case, the upper limit of the numerical range is defined by the positive threshold +|T|, and the lower limit of the numerical range is defined by the negative threshold -|T|.
[0019] Therefore, by quantization, floating-point numbers less than or equal to -|T| are converted to the integer corresponding to -|T|, and floating-point numbers greater than or equal to +|T| are converted to the integer corresponding to +|T|. The integer corresponding to -|T| is -127, and the integer corresponding to +|T| is +127. Floating-point numbers smaller than -|T| and larger than +|T| are called outliers.
[0020] By performing clipping before applying the scaling factor, quantization noise can be reduced, and the accuracy of the quantized values is improved.
[0021] FIG. 1 is a flowchart showing an example of threshold determination processing of a comparative example based on Non-Patent Document 3. The threshold determination processing in FIG. 1 is performed for each layer of the neural network.
[0022] First, the computer sets an initial value to a variable X representing a candidate threshold indicating the lower or upper limit of the numerical range (step 101), and quantizes N (N is an integer of 2 or more) numerical values to be quantized using X (step 102). In step 102, the computer converts numerical values outside the numerical range defined by X to the quantized values corresponding to X, and converts numerical values within the numerical range to the quantized values using the scaling factor.
[0023] Next, the computer calculates the Kullback-Leibler divergence (KL divergence) using the probability distribution P of the N numerical values to be quantized and the probability distribution Q of the N quantized numerical values according to the following formula (step 103).
[0024]
Number
[0025] The KL(P||Q) in Equation (1) represents the KL information quantity between the probability distribution P and the probability distribution Q. P(i) represents the probability of the i-th (i = 1 to N) value to be quantized, and Q(i) represents the probability of the i-th value after quantization. log represents the binary logarithm or the natural logarithm. KL(P||Q) is used as an index representing the difference between the probability distribution P and the probability distribution Q.
[0026] Next, the computer checks whether it has calculated the KL information quantity for all candidates (step 104). If there are unprocessed candidates remaining (step 104, NO), the computer updates the value of X (step 106) and repeats the processing from step 102 for the next candidate.
[0027] When the KL information quantity has been calculated for all candidates (step 104, YES), the computer selects the candidate with the minimum KL information quantity as the threshold value (step 105).
[0028] Figure 2 shows an example of the update processing in step 106 of Figure 1. 0 to 2048 indicates the positions of the bins of the histogram representing the probability distribution P. In this case, the variable X represents a candidate for the threshold value indicating the upper limit of the numerical range, and the initial value of X is set to the position of the 128th bin.
[0029] In step 106, the computer increases X by the bin width by incrementing the position of the bin indicating the value of X by 1. By repeating the processing of step 106, the value of X changes from the position of the 128th bin to the position of the 2048th bin. In step 102, outliers larger than X are converted to the quantized numerical values corresponding to X.
[0030] By performing quantization using a threshold value with the minimum KL information amount, the probability distribution of the quantized values can be made closer to the probability distribution of the values to be quantized. However, the threshold determination process in FIG. 1 is only effective for quantization that converts the activations of the CNN into 8-bit values.
[0031] The KL information amount only contains information on the occurrence frequencies of each value to be quantized and the occurrence frequencies of each quantized value, and does not contain information on the values themselves. For this reason, when the bit width of the quantized values is small, even if quantization is performed using a threshold value with the minimum KL information amount, the inference accuracy may decrease significantly.
[0032] FIG. 3 shows an example of the experimental results when the quantization of Non-Patent Document 3 is applied. In this experiment, as the learned model, a Transformer, which is a sequence conversion model described in Non-Patent Document 4, is used. The Transformer used in the experiment includes an encoder and a decoder, and each of the encoder and the decoder includes 9 fully connected layers.
[0033] The values to be quantized are the weights of the linear layers in the multi-head attention blocks included in each layer of the encoder or the decoder, and are represented by FP32. The bit width of the quantized values is 2 bits.
[0034] As the dataset, the Multi30k German-English translation dataset is used. The training data is 29,000 sentences, the validation data is 1,014 sentences, and the input data for inference is 1,000 sentences.
[0035] No quantization represents the case where inference is performed without quantizing the weights represented by FP32, and quantization (KL) represents the case where inference is performed by applying quantization based on a threshold value with the minimum KL information amount.
[0036] Inference accuracy 1 represents the BLEU (bilingual evaluation understudy) score when quantization is applied to the nine fully connected layers of the encoder. Inference accuracy 2 represents the BLEU score when quantization is applied to the nine fully connected layers of each of the encoder and the decoder. The higher the BLEU score, the higher the inference accuracy.
[0037] The inference accuracy without quantization is 35.08. On the other hand, the inference accuracy 1 of quantization (KL) is 33.26, and the inference accuracy 2 of quantization (KL) is 11.88. In this case, it can be seen that the inference accuracy 2 of quantization (KL) has decreased significantly.
[0038] Figure 4 shows a functional configuration example of the threshold determination device according to the embodiment. The threshold determination device 401 in Figure 4 includes a determination unit 411. When converting a numerical value outside the numerical range defined by the threshold into a quantized numerical value corresponding to the threshold among a plurality of numerical values to be quantized in the quantization of the neural network, the determination unit 411 determines the threshold. At this time, the determination unit 411 determines the threshold based on the quantization error for each of the plurality of numerical values.
[0039] According to the threshold determination device 401 in Figure 4, it is possible to suppress a decrease in inference accuracy due to quantization of the neural network.
[0040] Figure 5 shows a functional configuration example of the inference device corresponding to the threshold determination device 401 in Figure 4. The inference device 501 in Figure 5 includes a determination unit 511, a quantization unit 512, an inference unit 513, and a storage unit 514. The determination unit 511 corresponds to the determination unit 411 in Figure 4.
[0041] The storage unit 514 stores an inference model 521 for performing inferences in image processing, natural language processing, etc., and input data 524 to be inferred. The inference model 521 is a learned model including a neural network, and is generated, for example, by supervised machine learning. The inference model 521 may be a transformer.
[0042] The determination unit 511 determines a threshold value 522 for clipping for each layer of the neural network included in the inference model 521, and stores it in the storage unit 514. The threshold value 522 indicates the lower limit and the upper limit of the numerical range of the numerical values to be quantized.
[0043] The determination unit 511 generates the quantized numerical values corresponding to the respective numerical values by quantizing each of the N (N is an integer of 2 or more) numerical values to be quantized based on the numerical ranges defined by each of the plurality of candidates of the threshold value 522.
[0044] In quantization for converting FP32 to INT8, for example, the upper limit of the numerical range is defined by a candidate TC of a positive threshold value, and the lower limit of the numerical range is defined by a candidate -TC of a negative threshold value. In this case, the determination unit 511 can convert the i-th (i = 1 to N) numerical value v(i) to be quantized into the i-th quantized numerical value q(i) by, for example, the following formula.
[0045] q(i)=round(v(i) / S) (2)
[0046] S in formula (2) represents a scaling factor, and round(v(i) / S) represents a value obtained by rounding v(i) / S. However, when v(i) is greater than or equal to TC, q(i) = 127, and when v(i) is less than or equal to -TC, q(i) = -127.
[0047] Next, the determination unit 511 calculates a quantization error using each numerical value to be quantized and the quantized numerical value corresponding to each numerical value to be quantized, and calculates a statistical value of the quantization error for each of the N numerical values to be quantized. Then, the determination unit 511 selects the threshold value 522 from among the plurality of candidates based on the statistical values calculated from each of the plurality of candidates.
[0048] As the statistical value, for example, an average value, a median value, a mode value, a maximum value, or a sum is used. As the threshold value 522, for example, a candidate having the minimum statistical value is selected. By using the statistical value of the quantization error, the threshold value 522 suitable for each layer of the neural network can be easily determined.
[0049] In quantization for converting FP32 to INT8, for example, the average value QE of the quantization error for each of the N values to be quantized is calculated by the following equation.
[0050] vq(i)=S*q(i) (3)
Equation
[0051] vq(i) in Equation (3) represents the value obtained by inverse quantization of q(i), and |vq(i)-v(i)| in Equation (4) represents the i-th quantization error. However, when q(i)=127, vq(i)=TC, and when q(i)=-127, vq(i)=-TC.
[0052] The quantization error includes the information of the values themselves, along with the information on the occurrence frequencies of each value to be quantized and the occurrence frequencies of each value after quantization. Therefore, by selecting a candidate with the minimum statistical value of the quantization error as the threshold 522, the accuracy of the quantized values is improved compared to the case of selecting a candidate with the minimum KL information amount. Thus, even when the bit width of the quantized values is small, it is possible to suppress the decrease in inference accuracy due to quantization and maintain high inference accuracy.
[0053] The quantization unit 512 generates a quantization inference model 523 by quantizing each of the N values to be quantized using the threshold 522 for each layer of the neural network, and stores it in the storage unit 514.
[0054] In quantizing the values to be quantized, the quantization unit 512 converts an outlier value outside the numerical range defined by the lower and upper limits indicated by the threshold 522 into a quantized value corresponding to the lower or upper limit. Then, the quantization unit 512 converts the values within the numerical range into quantized values using a scaling factor.
[0055] The quantization target is, for example, the weights, biases, or activations in each layer of a neural network. The bit width of the quantized values is smaller than the bit width of the values of the quantization target. By quantizing the weights, biases, or activations, the neural network can be efficiently compressed.
[0056] The inference unit 513 performs inference on the input data 524 using the quantized inference model 523 and outputs the inference result. By performing inference using the quantized inference model 523 instead of the inference model 521, the power consumption and memory usage are reduced, and the inference process is accelerated.
[0057] FIG. 6 is a flowchart showing an example of the threshold determination process performed by the inference device 501 in FIG. 5. The threshold determination process in FIG. 6 is performed for each layer of the neural network included in the inference model 521.
[0058] First, the determination unit 511 sets an initial value to a variable X representing candidates for the threshold 522 (step 601), and quantizes N values of the quantization target using X (step 602). In step 602, the determination unit 511 converts values outside the numerical range defined by X into quantized values corresponding to X, and converts values within the numerical range into quantized values using a scaling factor.
[0059] Next, the determination unit 511 calculates the quantization error using each value of the quantization target and each quantized value, and calculates a statistical value of the quantization error for each of the N values of the quantization target (step 603).
[0060] Next, the determination unit 511 checks whether the statistical values of the quantization error have been calculated for all candidates (step 604). If there are remaining unprocessed candidates (step 604, NO), the determination unit 511 updates the value of X (step 606) and repeats the processing from step 602 for the next candidate.
[0061] When the statistical values of quantization errors are calculated for all candidates (step 604, YES), the determination unit 511 selects the candidate having the minimum statistical value as the threshold value 522 (step 605).
[0062] According to the threshold determination process of FIG. 6, since the statistical values of quantization errors are calculated for each candidate of the threshold value 522, the accuracy of the quantized numerical values for each candidate can be estimated based on the calculated statistical values. Therefore, it becomes possible to select a candidate having higher accuracy from among a plurality of candidates.
[0063] Next, the threshold determination process when the quantization target is the weights in each layer of the neural network will be described.
[0064] FIG. 7 shows an example of the distribution of weights to be quantized in one layer of the neural network. The horizontal axis represents the weights, and the vertical axis represents the appearance frequency. The weights are represented by FP32. W represents a set of N weights in one layer. max(W) represents the maximum value of the N weights, and min(W) represents the minimum value of the N weights.
[0065] The weight distribution in FIG. 7 is represented by a histogram including M bins. In this case, the bin width B is calculated by the following formula.
[0066] B = (max(W) - min(W)) / M (5)
[0067] FIG. 8 is a flowchart showing an example of the threshold determination process for weights. The threshold determination process in FIG. 8 is performed for each layer of the neural network included in the inference model 521.
[0068] The control variable k is used as a hyperparameter that specifies a candidate for the threshold value 522. The lower limit of the numerical range of the weights to be quantized is represented by -TH(k), and the upper limit is represented by +TH(k). TH(k) is a positive numerical value that changes according to k and represents a candidate for the upper limit of the numerical range.
[0069] First, the determination unit 511 sets an initial value k0 to k (step 801), and calculates TH(k) according to the following formula (step 802).
[0070] TH(k)=max(abs(W))-k*B (6)
[0071] In formula (6), abs(W) represents the set of absolute values of each weight included in W, and max(abs(W)) represents the maximum value of the elements of abs(W).
[0072] Next, the determination unit 511 generates the quantized weight Q(i) by quantizing the N weights W(i) (i = 1 to N) to be quantized using TH(k) (step 803).
[0073] In step 803, the determination unit 511 converts W(i) less than or equal to -TH(k) into the quantized weight -THQ(k) corresponding to -TH(k), and converts W(i) greater than or equal to TH(k) into the quantized weight THQ(k) corresponding to TH(k). Also, the determination unit 511 converts W(i) greater than -TH(k) and less than TH(k) into Q(i) using a scaling factor. For example, when Q(i) is represented by INT8, THQ(k) may be 127.
[0074] Next, the determination unit 511 sets an initial value 1 to the control variable i (step 804), and compares the absolute value abs(W(i)) of the i-th weight with TH(k) (step 805).
[0075] When abs(W(i)) is less than TH(k) (step 805, YES), the determination unit 511 calculates the quantization error qe(i) for W(i) according to the following formula (step 806).
[0076] qe(i)=abs(WQ(i)-W(i)) (7)
[0077] WQ(i) in Equation (7) represents the value obtained by inverse quantization of Q(i), and abs(WQ(i) - W(i)) represents the absolute value of WQ(i) - W(i).
[0078] On the other hand, when abs(W(i)) is equal to or greater than TH(k) (step 805, NO), the determination unit 511 calculates the quantization error qe(i) for W(i) using the following equation (step 807).
[0079] qe(i) = abs(W(i)) - TH(k) (8)
[0080] Next, the determination unit 511 compares i with N (step 808). When i has not reached N (step 808, NO), the determination unit 511 increments i by 1 (step 812) and repeats the processing from step 805 onwards.
[0081] When i reaches N (step 808, YES), the determination unit 511 calculates the average value QE(k) of the N quantization errors qe(i) using the following equation (step 809).
[0082] QE(k) = ave(qe) (9)
[0083] qe in Equation (9) represents the set of qe(1) to qe(N), and ave(qe) represents the average value of qe(1) to qe(N).
[0084] Next, the determination unit 511 compares TH(k) with L * B (step 810). L represents a positive integer. When TH(k) is greater than L * B (step 810, YES), the determination unit 511 increments k by Δk (step 813) and repeats the processing from step 802 onwards. For example, in the weight distribution shown in FIG. 7, when M = 2048, k0 = 0, Δk = 0.2, and L = 127 may be used.
[0085] When TH(k) is less than or equal to L*B (step 810, NO), the determination unit 511 ends the calculation of QE(k) and selects the TH(k) having the minimum QE(k) among the calculated QE(k) (step 811). Then, the determination unit 511 determines the threshold 522 indicating the lower limit of the numerical range as -TH(k) and determines the threshold 522 indicating the upper limit of the numerical range as TH(k).
[0086] FIG. 9 shows an example of experimental results when quantization of the embodiment is applied. The learned model and the dataset are the same as those in the experiment shown in FIG. 3.
[0087] The inference accuracy without quantization, the inference accuracy 1 and the inference accuracy 2 of quantization (KL) are the same as the experimental results shown in FIG. 3. Quantization (QE) represents the case where inference is performed by applying quantization based on the threshold 522 having the minimum QE(k).
[0088] The inference accuracy 1 of quantization (QE) is 35.09, and the inference accuracy 2 of quantization (QE) is 34.93. In this case, it can be seen that the inference accuracy 1 and the inference accuracy 2 of quantization (QE) are almost the same as the inference accuracy without quantization. Therefore, by determining the threshold 522 using the average value of the quantization error instead of the KL information amount, the inference accuracy comparable to that before quantization is maintained.
[0089] The configuration of the threshold determination device 401 in FIG. 4 is merely an example, and the components may be changed according to the use or conditions of the threshold determination device 401. The configuration of the inference device 501 in FIG. 5 is merely an example, and some components may be omitted or changed according to the use or conditions of the inference device 501.
[0090] The flowcharts in FIGS. 1, 6, and 8 are merely examples, and some processes may be omitted or changed according to the use or conditions of the threshold determination process. For example, in the threshold determination process in FIG. 8, it is also possible to change the quantization target to bias or activation.
[0091] The update process shown in FIG. 2 is just an example, and the method for updating candidates for the threshold value changes according to the use or conditions of the threshold value determination process. The experimental results shown in FIGS. 3 and 9 are just examples, and the inference accuracy changes according to the inference model and the quantization target. The weight distribution shown in FIG. 7 is just an example, and the weight distribution changes according to the inference model.
[0092] Equations (1) to (9) are just examples, and the inference device 501 may determine the threshold value 522 using another calculation formula.
[0093] FIG. 10 shows an example of the hardware configuration of an information processing apparatus (computer) used as the threshold value determination device 401 in FIG. 4 and the inference device 501 in FIG. 5. The information processing apparatus in FIG. 10 includes a CPU (Central Processing Unit) 1001, a memory 1002, an input device 1003, an output device 1004, an auxiliary storage device 1005, a medium drive device 1006, and a network connection device 1007. These components are hardware and are connected to each other by a bus 1008.
[0094] The memory 1002 is, for example, a semiconductor memory such as a ROM (Read Only Memory) or a RAM (Random Access Memory), and stores programs and data used for processing. The memory 1002 may operate as the storage unit 514 in FIG. 5.
[0095] The CPU 1001 (processor) operates as the determination unit 411 in FIG. 4 by executing a program using, for example, the memory 1002. The CPU 1001 also operates as the determination unit 511, the quantization unit 512, and the inference unit 513 in FIG. 5 by executing a program using the memory 1002.
[0096] The input device 1003 is, for example, a keyboard, a pointing device, etc., and is used for inputting instructions or information from a user or an operator. The output device 1004 is, for example, a display device, a printer, etc., and is used for making inquiries or giving instructions to a user or an operator, and for outputting processing results. The processing result may be an inference result for the input data 524.
[0097] The auxiliary storage device 1005 is, for example, a magnetic disk device, an optical disk device, a magneto-optical disk device, a tape device, etc. The auxiliary storage device 1005 may be a hard disk drive. The information processing apparatus can store programs and data in the auxiliary storage device 1005 and load them into the memory 1002 for use.
[0098] The medium drive device 1006 drives the portable recording medium 1009 and accesses the recording content thereof. The portable recording medium 1009 is a memory device, a flexible disk, an optical disk, a magneto-optical disk, etc. The portable recording medium 1009 may be a CD-ROM (Compact Disk Read Only Memory), a DVD (Digital Versatile Disk), a USB (Universal Serial Bus) memory, etc. A user or an operator can store programs and data in the portable recording medium 1009 and load them into the memory 1002 for use.
[0099] As described above, the computer-readable recording medium for storing the programs and data used for processing is a physical (non-transitory) recording medium such as the memory 1002, the auxiliary storage device 1005, or the portable recording medium 1009.
[0100] The network connection device 1007 is a communication interface circuit that is connected to a communication network such as a LAN (Local Area Network) or WAN (Wide Area Network) and performs data conversion associated with communication. The information processing device can receive a program and data from an external device via the network connection device 1007, load them into the memory 1002, and use them.
[0101] Note that the information processing device does not necessarily need to include all the components in FIG. 10, and it is also possible to omit some components according to the use or conditions of the information processing device. For example, when an interface with a user or operator is not required, the input device 1003 and the output device 1004 may be omitted. When a portable recording medium 1009 or a communication network is not used, the medium drive device 1006 or the network connection device 1007 may be omitted.
[0102] Although the disclosed embodiments and their advantages have been described in detail, those skilled in the art will be able to make various changes, additions, and omissions without departing from the scope of the present invention clearly described in the claims.
[0103] Regarding the embodiments described with reference to FIGS. 1 to 10, the following additional remarks are further disclosed. (Additional Remark 1) In the quantization of a neural network, when converting a numerical value outside the numerical range defined by a threshold value among a plurality of numerical values to be quantized into a quantized numerical value corresponding to the threshold value, the threshold value is determined based on the quantization error for each of the plurality of numerical values. A threshold determination program for causing a computer to execute the process. (Additional Remark 2) The process of determining the threshold value includes a process of determining the threshold value based on a statistical value of the quantization error for each of the plurality of numerical values, according to the threshold determination program described in Additional Remark 1. (Additional Remark 3) The process of determining the threshold value based on the statistical value is A process of generating a quantized value corresponding to each of the plurality of numerical values by quantizing each of the plurality of numerical values based on a numerical range defined by each of the plurality of candidates for the threshold value; A process of calculating the statistical value based on each of the plurality of numerical values and the quantized value corresponding to each of the plurality of numerical values; A process of selecting the threshold value from among the plurality of candidates based on the statistical value calculated from each of the plurality of candidates; The threshold determination program according to appended note 2, characterized by including the above. (Appended note 4) The threshold determination program according to any one of appended notes 1 to 3, characterized in that the quantization target is a weight, bias, or activation in the neural network. (Appended note 5) In quantization of a neural network, when converting a numerical value outside the numerical range defined by a threshold value among a plurality of numerical values of a quantization target into a quantized value corresponding to the threshold value, the threshold value is determined based on the quantization error for each of the plurality of numerical values. A threshold determination method, characterized in that a computer executes the process. (Appended note 6) The threshold determination method according to appended note 5, characterized in that the process of determining the threshold value includes a process of determining the threshold value based on a statistical value of the quantization error for each of the plurality of numerical values. (Appended note 7) The process of determining the threshold value based on the statistical value includes: A process of generating a quantized value corresponding to each of the plurality of numerical values by quantizing each of the plurality of numerical values based on a numerical range defined by each of the plurality of candidates for the threshold value; A process of calculating the statistical value based on each of the plurality of numerical values and the quantized value corresponding to each of the plurality of numerical values; A process of selecting the threshold value from among the plurality of candidates based on the statistical value calculated from each of the plurality of candidates; The threshold determination method according to appended note 6, characterized by including the above. (Appended note 8) The quantization target is a weight, bias, or activation in the neural network, and the threshold determination method according to any one of Appendices 5 to 7.
Explanation of symbols
[0104] 401 Threshold determination device 411, 511 Determination unit 501 Inference device 512 Quantization unit 513 Inference unit 514 Memory unit 521 Inference model 522 Threshold 523 Quantized inference model 524 Input data 1001 CPU 1002 Memory 1003 Input device 1004 Output device 1005 Auxiliary storage device 1006 Media drive device 1007 Network connection device 1008 Bus 1009 Portable recording medium
Claims
For each of a plurality of different variables that are candidates for a threshold value, execute a process of quantizing a plurality of numerical values to be quantized using the variable, and a process of calculating a statistical value of quantization errors based on the quantization errors for each of the plurality of numerical values calculated using each numerical value to be quantized and each quantized numerical value, determine, as the threshold value, the variable having the smallest said statistical value among the plurality of different variables that are candidates for the threshold value, for each layer of the neural network, execute a process of generating the inference model by quantizing each of the plurality of numerical values to be quantized using the determined threshold value. In the process of quantizing the plurality of numerical values to be quantized using the variable, convert a numerical value outside the numerical range defined in the variable into a quantized numerical value corresponding to the variable, convert a numerical value within the numerical range defined in the variable into a quantized numerical value using a scaling factor, A threshold determination program for causing a computer to execute the process. The quantization target is a weight, bias, or activation in the neural network, according to the threshold determination program according to claim 1. For each of a plurality of different variables that are candidates for a threshold value, execute a process of quantizing a plurality of numerical values to be quantized using the variable, and a process of calculating a statistical value of quantization errors based on the quantization errors for each of the plurality of numerical values calculated using each numerical value to be quantized and each quantized numerical value, determine, as the threshold value, the variable having the smallest said statistical value among the plurality of different variables that are candidates for the threshold value, for each layer of the neural network, execute a process of generating the inference model by quantizing each of the plurality of numerical values to be quantized using the determined threshold value. In the process of quantizing the plurality of numerical values to be quantized using the variable, convert a numerical value outside the numerical range defined in the variable into a quantized numerical value corresponding to the variable, convert a numerical value within the numerical range defined in the variable into a quantized numerical value using a scaling factor, A threshold determination method, characterized in that the computer executes the process.
4. In the quantization of the neural network included in the inference model, for each of a plurality of different variables that are candidates for the threshold value, a process of quantizing a plurality of numerical values to be quantized using the variable, and a process of calculating a statistical value of the quantization error based on the quantization error for each of the plurality of numerical values calculated using each numerical value to be quantized and each numerical value after quantization are executed, among the plurality of different variables that are candidates for the threshold value, the variable having the minimum said statistical value is determined as the threshold value, for each layer of the neural network, by quantizing each of the plurality of numerical values to be quantized using the determined threshold value, a process of generating the inference model is executed, and in the process of quantizing a plurality of numerical values to be quantized using the variable, a numerical value outside the numerical range defined in the variable is converted into a quantized numerical value corresponding to the variable, a numerical value within the numerical range defined in the variable is converted into a quantized numerical value using a scaling factor, A threshold determination device having a processing unit.
Citation Information
Patent Citations
Image processing method, device and equipment and readable storage medium
CN112287986A
Quantification method of neural network model and device for quantifying neural network model
CN112580805A
Neural network learning device and learning method
JP2020009048A
Method and device for neural network quantization
JP2020113273A
Information processing method and information processing device
JP2021005211A