Neural network processing device and neural network processing method
Patent Information
- Application Number
- JP2025032359
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0008】 本発明によれば、複数の畳み込み処理を含むニューラルネットワーク処理による認識結果を向上することができる。
Smart Images

Figure 2026144829000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates, in general terms, to the processing of neural networks. [Background technology]
[0002] A first example of a neural network processing device is described in Patent Document 1. This neural network processing device can perform processing at high speed by converting floating-point data to fixed-point data for processing. When converting floating-point data to fixed-point data, this neural network processing device performs the conversion using a linear equation that maps the minimum and maximum values of the floating-point data to the minimum and maximum values of the fixed-point data.
[0003] A second conventional example of a neural network processing device is described in Non-Patent Document 1. This neural network processing device employs a method to improve recognition results when processing floating-point data by converting it to integer data. Specifically, this neural network processing device accumulates the occurrence frequency of floating-point data and converts the data values whose accumulated frequency reaches a specific value (99.9%, 99.99%, etc.) to the maximum value of the integer type. This neural network processing device processes multiple values of the accumulated frequency corresponding to the maximum value of the integer type (hereinafter referred to as the accumulated frequency threshold) and finds the accumulated frequency threshold that yields the best recognition results. By processing using the found accumulated frequency threshold, the recognition results are improved. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Special Publication No. 2022-507704 [Non-patent literature]
[0005] [Non-Patent Document 1] Hao Wu, et al., “Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation”,April,20,2020 [Overview of the project] [Problems that the invention aims to solve]
[0006] In the first conventional example described above, calculation errors may occur due to the conversion of floating-point data to fixed-point data, potentially degrading the recognition results. In the second conventional example described above, when multiple convolution operations are performed, the same cumulative frequency threshold is used for all convolution operations, which means that the optimal cumulative frequency threshold may not be used when focusing on a specific convolution operation. [Means for solving the problem]
[0007] To achieve the above objective, a neural network processing device according to one aspect of the present invention performs a plurality of convolution processes, and an integerization process for each convolution process, which performs integerization using coefficients generated using statistical information of the input data and a cumulative frequency threshold. The cumulative frequency threshold used to generate the coefficients used in the integerization process may differ for each convolution process. [Effects of the Invention]
[0008] According to the present invention, the recognition results obtained by neural network processing, which includes multiple convolution operations, can be improved. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing a neural network processing device according to Embodiment 1 of the present invention. [Figure 2] This figure shows a first example of the configuration of the neural network processing to be executed by the neural network processing unit shown in Figure 1. [Figure 3]Figure 2 shows the relationship between the input and output of the 8-bit integer conversion process 211. [Figure 4] Figure 2 shows the processing details of the 8-bit integer conversion process 211. [Figure 5] This figure shows the processing details of the convolution process 212 in Figure 2. [Figure 6] This figure shows the relationship between input and output when the weight kS8 in Figure 5 is obtained from the floating-point weight kF. [Figure 7] Figure 2 shows the processing details of the bias addition process 213. [Figure 8] Figure 7 shows the relationship between input and output when the bias value bS32 is obtained from the floating-point bias value bF. [Figure 9] Figure 2 shows the processing details of the activation function process 214. [Figure 10A] This figure shows a method for determining the coefficient (inU32-max) of the 8-bit integer conversion process in Figure 4 using a cumulative frequency threshold of the input value. [Figure 10B] This figure shows a method for determining the coefficient (inU32-max) of the 8-bit integer conversion process in Figure 4 using a cumulative frequency threshold of the input value. [Figure 11] This figure shows a first example of a method for calculating recognition results used to select a cumulative frequency threshold used to determine the coefficients for 8-bit integer conversion. [Figure 12] This figure shows how to select the cumulative frequency threshold used to determine the coefficients for the 8-bit integer conversion process, using the recognition results calculated in Figure 11. [Figure 13] This figure shows a second example of a method for calculating recognition results used to select a cumulative frequency threshold used to determine the coefficients for 8-bit integer conversion. [Figure 14] This figure shows a second example of the configuration of the neural network processing to be executed by the neural network processing unit shown in Figure 1. [Figure 15] Figure 14 shows the relationship between the input and output of the 8-bit integer conversion process 1411. [Figure 16]Figure 14 shows the processing details of the 8-bit integer conversion process 1411. [Figure 17] This figure shows the processing details of the floating-point conversion process 1413 in Figure 14. [Figure 18A] Figure 14 shows the processing details of the activation function process 1415. [Figure 18B] Figure 14 shows the processing details of the activation function process 1415. [Figure 19] This figure shows a third example of the configuration of the neural network processing to be executed by the neural network processing unit shown in Figure 1. [Modes for carrying out the invention]
[0010] In the following explanation, "interface device" may refer to one or more interface devices. These one or more interface devices may be at least one of the following: An I / O interface device is one or more I / O (Input / Output) interface devices. An I / O (Input / Output) interface device is an interface device to at least one of the following: an I / O device and a remote display computer. The I / O interface device to the display computer may be a communication interface device. The at least one I / O device may be either a user interface device, such as an input device like a keyboard and a pointing device, or an output device like a display device. A communication interface device consisting of one or more communication interface devices. These one or more communication interface devices may be one or more identical communication interface devices (for example, one or more NICs (Network Interface Cards)) or two or more different communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).
[0011] Furthermore, in the following explanation, "memory" refers to one or more memory devices, which are examples of one or more storage devices, and are typically main memory devices. At least one memory device in memory may be a volatile memory device or a non-volatile memory device.
[0012] Furthermore, in the following explanation, "persistent storage device" may refer to one or more persistent storage devices, which are examples of one or more storage devices. Persistent storage devices are typically non-volatile storage devices (e.g., auxiliary storage devices), and specifically, they may be, for example, HDDs (Hard Disk Drives), SSDs (Solid State Drives), NVME (Non-Volatile Memory Express) drives, or SCMs (Storage Class Memory).
[0013] Furthermore, in the following explanation, "storage device" may refer to at least memory, which is part of both memory and persistent storage.
[0014] Furthermore, in the following explanation, "processor" may refer to one or more processor devices. At least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may be a broad processor device such as a circuit that is a collection of gate arrays defined by a hardware description language that performs some or all of the processing (e.g., FPGA (Field-Programmable Gate Array), CPLD (Complex Programmable Logic Device), or ASIC (Application Specific Integrated Circuit)). [Examples]
[0015] Figure 1 is a block diagram showing a neural network processing device according to Embodiment 1 of the present invention.
[0016] The neural network processing unit 108 may be an example of a computer, and may include a microprocessor 101, a GPU 111, a ROM (Read Only Memory) 109, and RAM (Random Access Memory) 110 and 112. The microprocessor 101 includes interface circuits 103, 104, 105, 106, and 107, and a CPU 102. Interface circuit 103 is the interface circuit for the ROM 109. Interface circuit 104 is the interface circuit for the RAM 110. Interface circuit 105 is the interface circuit for the GPU 111. Interface circuit 106 is the interface circuit for cameras 113 and 114. Interface circuit 107 is the interface circuit for the personal computer 115. Interface circuits 103 to 107 may be an example of interface devices. ROM 109, RAM 110 and 112 may be an example of memory. CPU 102 and GPU 111 may be an example of processors. The processor is connected to memory and input / output devices (such as cameras and personal computers) via an interface device, enabling communication. The processor may also be connected to external devices such as servers on a communication network via an interface device, enabling communication. The neural network processing unit 108 can communicate with cameras 113 and 114 and personal computer 115. "Personal computer" may be an example of an information processing terminal having a user interface device such as a keyboard or display. The number of elements such as cameras and interface circuits is not limited to the example shown in Figure 1, and may be one or more.
[0017] The neural network processing unit 108 receives image data from cameras 113 and 114, performs recognition using a neural network, and then transmits the recognition results to the personal computer 115. For example, the neural network processing unit 108 may be installed (or installed in a manner that allows communication with the train) on a train consisting of n cars (where n is a natural number), and the image recognition results from the neural network processing unit 108 may be input to the train's control device, which may then perform train operation control (e.g., automatic driving control) based on the image recognition results. Alternatively, for example, the train may be just one example of a vehicle, and the neural network processing unit 108 may be installed (or installed in a manner that allows communication with the vehicle) on a vehicle other than a train (e.g., a passenger car or a truck), and the image recognition results from the neural network processing unit 108 may be input to the vehicle's control device, which may then perform vehicle operation control (e.g., automatic driving control) based on the image recognition results. In this embodiment, an example of image recognition is shown, but the present invention is not limited to this and can be applied to neural network processing units in general, such as speech recognition and text analysis.
[0018] Cameras 113 and 114 each capture images in different directions. This makes it possible to capture images of the entire area to be recognized and perform recognition processing. Alternatively, cameras with a high pixel count and wide shooting range may be used as cameras 113 and 114. This makes it possible to cover the entire area with a small number of cameras, even if the area to be recognized is large. In this case, the image captured by one camera may be divided into multiple areas and recognition processing may be performed separately. This makes it possible to exclude areas that do not need to be recognized, such as the sky, and perform recognition processing only on the necessary areas.
[0019] The microprocessor 101 is an LSI that integrates the CPU 102 and interface circuits 103, 104, 105, 106, and 107 onto a single chip. This configuration is just one example, and some or all of the ROM 109, RAM 110, GPU 111, and RAM 112 may be built into the microprocessor 101. RAM 110 is just one example of a memory that stores the calculation results of the microprocessor 101, and RAM 112 is just one example of a memory that stores the calculation results of the GPU 111.
[0020] The CPU 102 retrieves and executes the software stored in the ROM 109 via the interface circuit 103. Since the ROM 109 is often slower than the RAM 110, the software may be copied from the ROM 109 to the RAM 110 at startup, and then retrieved from the RAM 110 thereafter. The CPU 102 then performs the following series of processes according to the software retrieved from either the ROM 109 or the RAM 110.
[0021] The CPU 102 first acquires image data from cameras 113 and 114 via interface circuit 106 and stores it in RAM 110 via interface circuit 104.
[0022] The CPU 102 then reads the image data stored in the RAM 110 via the interface circuit 104 and transfers it to the GPU 111 via the interface circuit 105.
[0023] The CPU 102 then reads the software stored in the ROM 109 or RAM 110 via the interface circuit 103 or 104, transfers it to the GPU 111 via the interface circuit 105, and instructs the GPU 111 to start calculations.
[0024] Next, when the CPU 102 receives notification from the GPU 111 that the calculation has finished via the interface circuit 105, it retrieves the calculation result from the GPU 111 via the interface circuit 105 and stores it in the RAM 110 via the interface circuit 104.
[0025] The CPU 102 then retrieves the calculation result from the GPU 111 from the RAM 110 via the interface circuit 104, performs predetermined processing, and then transmits it to the PC 115 via the interface circuit 107.
[0026] When the GPU 111 receives image data from the CPU 102 via the interface circuit 105, it stores it in the RAM 112.
[0027] The GPU 111 also receives software from the CPU 102 via the interface circuit 105, executes that software, and stores the calculation results in the RAM 112.
[0028] When the GPU 111 receives a request from the CPU 102 to read the calculation result via the interface circuit 105, it reads the calculation result from the RAM 112 and outputs it to the interface circuit 105.
[0029] When the personal computer 115 receives the calculation result from the CPU 102 via the interface circuit 107, it converts it into a predetermined format and displays it.
[0030] In this embodiment, we have shown an example of performing neural network calculations on the GPU 111, but the present invention is not limited to this and can also be applied when calculations are performed using calculation circuits such as a CPU, FPGA (Field Programmable Gate Array), or ASIC (Application Specific, Integrated Circuit). Furthermore, neural network processing is not limited to the neural network processing device 108 having the configuration illustrated in Figure 1, but may be performed in a neural network processing device with other configurations. For the sake of simplicity, it will be assumed below that neural network processing is performed by a processor (typically including a GPU). For example, neural network processing is performed by the processor executing neural network processing software.
[0031] Figure 2 shows a first example of the configuration of neural network processing. Note that this figure describes the recognition process for a single input image; when performing recognition processing on multiple images, the same processing as in Figure 2 should be performed for each input image. In this case, the processing configuration may be the same for all images, or it may be different for each image according to the purpose of the recognition processing.
[0032] The process shown in Figure 2 consists of three layers (Layer-1 (201), Layer-2 (202), and Layer-3 (203)). Layer-1 (201) receives the input image and outputs the processing result to Layer-2 (202). Layer-2 (202) receives the output from Layer-1 (201) and outputs the processing result to Layer-3 (203). Layer-3 (203) receives data from Layer-2 (202) and outputs the processing result as the recognition result.
[0033] Layer-1 (201) consists of an 8-bit integer conversion process 211, a convolution process 212, a bias addition process 213, and an activation function process 214. Layers-2 and-3 are similar. That is, Layer-2 (202) consists of an 8-bit integer conversion process 221, a convolution process 222, a bias addition process 223, and an activation function process 224. Layer-3 (203) consists of an 8-bit integer conversion process 231, a convolution process 232, a bias addition process 233, and an activation function process 234. In each of the multiple layers in the neural network, the processor performs an 8-bit integer conversion process to convert the 32-bit integer of the data input to that layer into an 8-bit integer, a convolution process to perform a sum-of-accumulate operation between the 8-bit integer obtained by the 8-bit integer conversion process and the weight data, a bias addition process to add a bias value to the output of the convolution process, and an activation function process to calculate a predetermined function using the output of the bias addition process.
[0034] Note that, although the number of layers is set to 3 in the present embodiment, the present invention is not limited thereto, and can be applied to a neural network processing apparatus having any number of layers. Further, the configuration of each layer is not limited to the present embodiment, and the present invention can also be applied to cases where there is no bias addition or activation function. Further, even when layers that do not include convolution processing are mixed, the present invention can be applied as long as there are a plurality of layers that perform convolution processing.
[0035] FIG. 3 is a diagram showing the relationship between the input and output of the 8-bit quantization processing 211 in FIG. 2.
[0036] As shown in FIG. 3, the processor converts the input (in U32 ) 0 and in U32-max to output (in U8 ) 0 and 255 by using a linear expression for association. Further, when the input is larger than in U32-max , the processor sets the output to 255. For the value of in U32-max , a numerical value predetermined by a method described later is used. In FIG. 3, the input (in U32 ) is of a 32-bit unsigned integer type (unsigned means being 0 or positive), and the output (in U8 ) is of an 8-bit unsigned integer type, which is provided as an example; however, the present invention is not limited thereto, and can be applied to unsigned integer types having any number of bits. In other words, 8-bit quantization processing is an example of quantization processing, and any n-bit quantization processing may be employed.
[0037] The 8-bit quantization processing 221 and 231 in FIG. 2 are also the same as in FIG. 3, provided that the value of in U32-max differs for each layer.
[0038] FIG. 4 is a diagram showing the processing content of the 8-bit quantization processing 211 in FIG. 2.
[0039] According to line 403, the processor calculates a value corresponding to the slope of the linear expression shown in FIG. 3 (a F ). a F is a floating-point type variable. in used for the calculationU32-max The value can be a value that is already stored in memory.
[0040] According to line 404, there is a loop corresponding to the vertical coordinate (h) of the image. H represents the number of pixels in the vertical direction. The value of H can be a constant or a variable. If it is a variable, you can use a value that has been stored in memory beforehand, or you can pass it as a function argument from a higher-level function.
[0041] According to line 405, there is a loop corresponding to the horizontal coordinate (w) of the image. W represents the number of pixels in the horizontal direction. The value of W can be a constant or a variable. If it is a variable, you can use a value that has been stored in memory beforehand, or you can pass it as an argument to a function from a higher-level function.
[0042] According to line 406, there is a loop corresponding to different attribute information (ci) within the same pixel. CI indicates the number of types of attribute information. If the input is a color image, there are usually three attribute information types: red (R), green (G), and blue (B), so CI is 3. If the image format is different, CI may have a different value.
[0043] According to line 407, the processor takes the input (in U32 The element to be calculated in ) is a calculated according to line 403. F The value obtained by multiplying by and converting it to a 32-bit unsigned integer type is stored in the variable in1 U32 Substitute into the input (in U32 ) is a 3D matrix corresponding to h, w, and ci. According to line 407, the elements corresponding to h, w, and ci at runtime are used.
[0044] According to line 408, the processor is in1 U32 If it is greater than 255, then it is set to 255. This is because in Figure 3, the input (in U32 ) is in U32-max This is equivalent to setting the output to 255 if it is greater than the given value.
[0045] According to line 409, the processor is in1 U32 Output (inU8 Substitute into ). Output (in U8 ) is a 3D matrix corresponding to h, w, and ci. According to line 409, the processor assigns values to the elements corresponding to h, w, and ci at runtime. U32 is the output (in U8 Because it has more bits than ), the number of bits is reduced during assignment. Generally, the output (in U8 ) is in1 U32 It may not be the same, but the input (in U32 Since ) is 0 or greater, in1 U32 It is guaranteed to be 0 or greater, and the process according to line 408 will result in in1 U32 It is guaranteed to be 255 or less. Therefore, in1 U32 This will be a number that can be represented as an 8-bit unsigned integer, and the output (in U8 ) is always in1 U32 This will have the same value.
[0046] The processing details of the 8-bit integer conversion processes 221 and 231 in Figure 2 are the same as in Figure 4, U32-max As mentioned above, the value of will differ for each layer. Additionally, the values of H, W, and CI may also differ for each layer.
[0047] Figure 5 shows the processing details of the convolution process 212 in Figure 2.
[0048] in U8 This is the output of the 8-bit integer conversion process in Figure 4, and it is an 8-bit unsigned integer type. d1 S32 This is the result of the convolution process and is a 32-bit signed integer. Note that d1 S32 The number of bits shown is just one example; the present invention can be applied to integer types with any number of bits.
[0049] Lines 503 and 504 are the same as lines 404 and 405 in Figure 4.
[0050] According to line 505, the output (d1 S32There is a loop corresponding to different attribute information (co) within the same pixel. CO indicates the number of types of attribute information.
[0051] According to line 506, the processor stores the result of the convolution in a variable (d10 S32 Initialize ) to 0. d10 S32 is output (d1 S32 It is the same 32-bit signed integer type as ). Output (d1 S32 If the number of bits is different from 32 bits, the processor will use d10 S32 Output (d1 S32 ) should have the same number of bits.
[0052] According to line 507, the input (in U8 There is a loop corresponding to different attribute information (ci) within the same pixel. CI indicates the number of types of attribute information.
[0053] Lines 508 and 509 correspond to the convolution operation, where the processor calculates the weights (k S8 ) and input (in U8 ) The product of d10 S32 Add to the weight (k S8 ) is a two-dimensional matrix corresponding to co and ci. According to line 508, the processor uses the elements corresponding to co and ci at runtime. Input (in U8 ) is a 3D matrix corresponding to h, w, and ci. According to line 509, the processor uses the elements corresponding to h, w, and ci at runtime. Weight (k S8 ) can be a value obtained beforehand during the neural network training process and converted to an 8-bit signed integer type. The weights obtained during the neural network training process are often of floating-point type, and it is necessary to convert these values to integer type. The method for converting to an 8-bit signed integer type will be described later. Note that k S8 The number of bits shown is just one example; the present invention can be applied to integer types with any number of bits.
[0054] According to line 511, the processor is d10 S32 Output (d1S32 Substitute into ). Output (d1 S32 ) is a 3D matrix corresponding to h, w, and co. According to line 511, the processor uses the elements corresponding to h, w, and co at runtime.
[0055] The processing details of the convolution operations 222 and 232 in Figure 2 are the same as in Figure 5, but H, W, CO, CI, k S8 The value may differ for each layer.
[0056] Figure 6 shows the weights k of the 8-bit signed integer type in Figure 5. S8 The weight k is a floating-point number. F This diagram shows the relationship between the two when calculating from them.
[0057] According to Figure 6, the processor uses floating-point weights (k F ) of -k F-max and k F-max The weight (k) is an 8-bit signed integer. S8 The transformation is performed using a linear expression that maps -127 and 127 in ). F ga-k F-max If it is smaller, the processor will be k S8 Let k be -127. Also, k F ga k F-max If it is greater than k, the processor will S8 Let k be 127. F and k S8 This is a two-dimensional matrix corresponding to co and ci in Figure 5, but the processor uses k for elements where co is the same but ci is different. F-max The same value is assumed. For different values of co, the processor uses different k values. F-max You may use it.
[0058] Figure 7 shows the processing details of the bias addition process 213 in Figure 2.
[0059] d1 s32 This is the output of the convolution process in Figure 6, and is a 32-bit signed integer. d2 S32This is the result of the bias addition process and is a 32-bit signed integer type. Note that d2 S32 The number of bits shown is just one example; the present invention can be applied to integer types with any number of bits.
[0060] Lines 703-705 are the same as lines 503-505 in Figure 5.
[0061] Lines 706 and 707 correspond to the bias addition operation, and the processor takes the input (d1 S32 ) with bias value (b S32 ) is added to output (d2 S32 Substitute into ). Input (d1 S32 ) and output (d2 S32 ) is a 3D matrix corresponding to h, w, and co. According to lines 706 and 707, the processor uses the elements corresponding to h, w, and co at runtime. Bias value (b S32 ) is a one-dimensional vector corresponding to co. According to line 707, the processor uses the element corresponding to co at runtime. Bias value (b S32 ) can be a value obtained in advance during the neural network training process and converted to a 32-bit signed integer type. The bias value obtained during the neural network training process is often of floating-point type, and it is necessary to convert this value to an integer type. The method for converting to a 32-bit signed integer type will be described later. S32 The number of bits shown is just one example; the present invention can be applied to integer types with any number of bits.
[0062] The processing details of bias addition processes 223 and 233 in Figure 2 are the same as in Figure 7, but H, W, CO, b S32 The value may differ for each layer.
[0063] Figure 8 shows the bias value b of the 32-bit signed integer type in Figure 7. S32 The bias value b of the floating-point type F This diagram shows the relationship between the two when calculating from them.
[0064] According to FIG. 8, the processor converts the bias value (b F ) of floating-point type, which is -b F-max and b F-max , by using a linear expression associated with -b S32 and b S32-max and b S32-max of 32-bit signed integer type bias value (b F is smaller than -b F-max , the processor sets b S32 as -b S32-max . In addition, when b F is larger than b F-max , the processor sets b S32 as b S32-max . b F-max is the product of in U32-max in FIG. 3 and k F-max in FIG. 6. in U32-max does not depend on co in FIG. 7, but since k F-max depends on co in FIG. 7, b F-max also depends on co in FIG. 7. b S32-max is the product of the maximum value (255) of in U8 in FIG. 3 and the maximum value (127) of k S8 in FIG. 6.
[0065] FIG. 9 is a diagram showing the processing content of the activation function processing 214 in FIG. 2.
[0066] d2 s32 is the output of the bias addition processing in FIG. 7, and is of 32-bit signed integer type. out S32 is the processing result of the activation function processing, and is of 32-bit signed integer type. Note that the number of bits of out S32 is an example, and the present invention can be applied to integer types of any number of bits.
[0067] Lines 903 to 905 are the same as lines 503 to 505 in FIG. 5.
[0068] Lines 906 to 910 correspond to arithmetic processing of an activation function, and the processor, when the input (d2 S32 ) is 0 or more, assigns the input (d2 S32 ) to the output (outS32 ) is assigned, and if it is less than 0, the output (out S32 0 is assigned to ). This process guarantees that the output of the activation function process 214 will be a non-negative integer. The output of the activation function process 214 becomes the input to the 8-bit integer conversion process 221 of the next layer. The reason why the 8-bit integer conversion process 221 assumes that the input is non-negative is because the processing content was determined based on the properties of the activation function process 214 described above.
[0069] The activation function processes 224 and 234 in Figure 2 are similar to those in Figure 5, but the values of H, W, and CO may differ for each layer.
[0070] Figures 10A and 10B show the coefficients (in) of the 8-bit integer conversion process in Figure 4. U32-max This figure shows a method for determining the input value using a cumulative frequency threshold.
[0071] Input (in U32 The distribution of the input (in) depends on the characteristics of the neural network being implemented, but here the. U32 An example was shown where the frequency of occurrence decreases exponentially as the number of elements increases. Since the frequency of occurrence of inputs in each layer depends on the input image of the neural network, when applying the present invention, the frequency of occurrence of inputs in each layer is obtained when a test image (one or more) is input to the neural network, and that data is used.
[0072] Enter the frequency of occurrence (in U32 The cumulative frequency, calculated by accumulating from the smallest value, is the input (in U32 It increases with increasing ). In the example in Figure 10B, the input (in) corresponding to a specific value of the cumulative frequency (5 possibilities: 99%, 99.9%, 99.99%, 99.999%, 100%; hereafter referred to as the cumulative frequency threshold) increases. U32 ) value (in U32-max(99%) , in U32-max(99.9%) , in U32-max(99.99%) , in U32-max(99.999%) , in U32-max(100%) (5 ways) are converted to 8-bit integers using coefficients (in U32-maxThese are the candidates. From these five options, we select the one that yields the best recognition result in neural network processing.
[0073] Figure 11 shows a first example of a method for calculating recognition results used to select a cumulative frequency threshold used to determine the coefficients for 8-bit integer processing.
[0074] Layers 1101 to 1103 execute layers 1(201) to 3(203) in Figure 2 using floating-point operations, respectively. The processor inputs one or more test images to layer 1101 and saves the output of layer 1103 as the recognition result (F) to memory, for example.
[0075] The processor processes the input as an 8-bit integer, similar to layer-1 (201) in Figure 2, as part of layer 1111 processing. The coefficient for 8-bit integer processing (in U32-max The processor uses the five values shown in Figure 10 (in U32-max(99%) , in U32-max(99.9%) , in U32-max(99.99%) , in U32-max(99.999%) , in U32-max(100%) Use ).
[0076] On the other hand, layer 1112 corresponds to layer-2(202) in Figure 2, but the processor converts the input to a floating-point type for processing. Layer 1113 does the same.
[0077] The processor inputs one or more test images to layer 1111 and saves the output of layer 1113 as the recognition result (1I) to memory, for example. The recognition result (1I) includes five possible results corresponding to five different values of the coefficients used in the 8-bit integer conversion process of layer 1111.
[0078] Although not shown in Figure 11, for Layers 2 and 3, similar to Layer 1, the processor creates a neural network configured to process only those layers with 8-bit integer data, while processing the other layers with floating-point data, thereby generating recognition results (2I) and (3I).
[0079] Figure 12 shows how to select a cumulative frequency threshold to use for determining the coefficients in the 8-bit integer conversion process, using the recognition results saved in Figure 11.
[0080] The table illustrated in Figure 12 is stored in memory (e.g., RAM 110 or 112). In Figure 12, for the recognition result of Layer-1 (1I), the coefficient corresponding to the cumulative frequency threshold of 99% (in U32-max(99%) The result when using R 1I-1 , coefficient corresponding to 99.9% (in U32-max(99.9%) The result when using ) is R 1I-2 , coefficient corresponding to 99.99% (in U32-max(99.99%) The result when using R 1I-3 , coefficient corresponding to 99.999% (in U32-max(99.999%) The result when using R 1I-4 , coefficient corresponding to 100% (in U32-max(100%) The result when using R 1I-5 It is written as follows. The processor is R 1I-1 ~R 1I-5 Compare the results. The best result is R 1I-5 Therefore, the processor, with respect to Layer-1, R 1I-5 Select the corresponding cumulative frequency threshold of 100%.
[0081] Similarly, with respect to Layer 2, R 2I-1 Since this yielded the best result, the processor is R 2I-1 Select the corresponding cumulative frequency threshold of 99%. For Layer-3, R 3I-3 Since this yielded the best result, the processor is R 3I-3 Select the corresponding cumulative frequency threshold of 99.99%.
[0082] There are various methods for obtaining the best result. For example, one method is to use the comparison result with the recognition result (F) saved in Figure 11. Specifically, for example, if the content of recognition is object detection, the processor may select the recognition result with the largest number of detections of the same object as recognition result (F), or the recognition result with the smallest number of detections of different objects, as the best result. With this method, learning is possible using recognition result (F) as the correct answer, so learning is possible without having to prepare correct answer data in advance. Also, with this method, even after the neural network processing unit has been shipped, it is possible to update the cumulative frequency threshold for at least one layer in the neural network processing unit. Specifically, for example, the processor may record one or more input images and the recognition result for each image in memory during the operating period (e.g., the period during which a train is running), which is a period during which inference including neural network processing is required. During periods other than the operating period (e.g., at night), for each recorded input image, the recognition result (1I), (2I), and (3I) may be compared with the recognition result (F), and the corresponding cumulative frequency threshold may be updated based on the result of this comparison (for example, if there is a layer where the difference between recognition result (F) and recognition result (I) is greater than a predetermined value, the cumulative frequency threshold corresponding to that layer may be updated).
[0083] Furthermore, for example, if the correct recognition result is available, the processor may select the recognition result with the highest number of correct answers or the recognition result with the lowest number of incorrect answers as the best result. In this method, the recognition result (F) saved in Figure 11 is not needed to find a good result, so it is not necessary to execute layers 1101 to 1103 in Figure 11.
[0084] The coefficients for 8-bit integer conversion corresponding to the cumulative frequency threshold selected in Figure 12 are stored in the memory circuit. When executing the 8-bit integer conversion processes 211, 221, and 231 in Figure 2, the coefficients stored in the memory circuit are used.
[0085] Note that the processing shown in Figure 12 does not necessarily have to be performed by the neural network processing unit 108 in Figure 1. For example, the processing in Figure 12 may be performed by the personal computer 115 in Figure 1, and the 8-bit integer conversion coefficients corresponding to the selected cumulative frequency threshold may be sent from the personal computer 115 to the neural network processing unit 108. In that case, the neural network processing unit 108 stores the 8-bit integer conversion coefficients received from the personal computer 115 in a memory circuit, and uses the coefficients stored in the memory circuit when executing the 8-bit integer conversion processes 211, 221, and 231 in Figure 2.
[0086] Figure 13 shows a second example of a method for calculating recognition results used to select a cumulative frequency threshold used to determine the coefficients for 8-bit integer processing.
[0087] The difference from Figure 11 is that the processor processes the data that has been converted to 8-bit integers in layers 1312 and 1313. In this case, in layer 1111, the processor calculates the coefficient (in) of the 8-bit integer conversion process. U32-max ) as shown in Figure 10, the five possible values in U32-max(99%) , in U32-max(99.9%) , in U32-max(99.99%) , in U32-max(99.999%) , in U32-max(100%) This is used. On the other hand, in layers 1312 and 1313, the processor uses the coefficient (in) of the 8-bit integer processing. U32-max As such, only one coefficient value is used. Similar to Figure 11, the recognition result (1I) includes five results corresponding to the five values of the coefficient in the 8-bit integer conversion process of layer 1111.
[0088] Similarly, for Layer-2 and Layer-3, the processor uses five different coefficient values only for those layers, and only one coefficient for the other layers, thereby obtaining recognition results (2I) and recognition results (3I).
[0089] The method for selecting the cumulative frequency threshold using recognition results (1I) to (3I) is the same as in Figure 12. [Examples]
[0090] Example 2 will now be described. In this description, the differences from the previously mentioned example will be explained primarily, and the similarities with the previously mentioned example will be omitted or simplified (this also applies to subsequent examples).
[0091] Figure 14 shows a second example of a neural network processing configuration.
[0092] The difference from Figure 2 is that a floating-point conversion process is added to convert the output of the convolution process to a floating-point type, and the bias addition process and activation function process are performed using floating-point data. This makes it possible to handle cases where the activation function process cannot be performed with integer data. The processing content of the convolution process is the same as in Figure 2, but the 8-bit integer conversion process differs from Figure 2 because the input is floating-point data.
[0093] Figure 15 shows the relationship between the input and output of the 8-bit integer conversion process 1411 in Figure 14.
[0094] According to Figure 15, the processor receives input (in F ) in F-min and in F-max Output (in S8 The conversion is performed using a linear equation that maps -128 and 127 of ). Also, the input is in F-min If the input is less than -128, the processor will set the output to -128. F-max If it is greater than, the processor sets the output to 127. F-max Regarding the value of , similar to Figure 10, the processor determines five values corresponding to the cumulative frequency threshold and uses the value corresponding to the cumulative frequency threshold selected in the same way as in Figure 12. Also, in F-min The processor uses the smallest input value with a non-zero frequency of occurrence. The frequency of occurrence is determined by inputting a test image, so when an image different from the test image is input, the input F-min A smaller value may be entered. According to Figure 15, the input is inF-min If the input is smaller, the output is set to -128. F-min It can also handle smaller cases. In Figure 15, the output (in S8 Although it is assumed that ) is an 8-bit signed integer type, the present invention is not limited thereto and can be applied to any number of signed or unsigned integer types.
[0095] The 8-bit integer conversion processes 1421 and 1431 in Figure 14 are the same as in Figure 15. However, in F-min , and in F-max The value differs for each layer.
[0096] Figure 16 shows the processing details of the 8-bit integer conversion process 1411 in Figure 14.
[0097] According to line 1603, the processor uses a coefficient (a) that corresponds to the slope of the linear equation shown in Figure 15. F ) calculate a F is a floating-point variable. The variable used in the calculation is `in`. F-min , and in F-max The value used is one that is pre-stored in memory.
[0098] Lines 1604-1606 are the same as lines 404-406 in Figure 4.
[0099] According to line 1607, the processor takes the input (in F From the elements to be calculated in F-min Subtracting this, a is calculated according to line 1603. F The value obtained by multiplying by and converting it to a 32-bit signed integer type is stored in the variable in S32 Substitute into the input (in F The matrix is a three-dimensional matrix corresponding to h, w, and ci, and according to line 1607, the processor uses the elements corresponding to h, w, and ci at runtime.
[0100] According to line 1608, the processor is in S32 If the value is less than 0, then S32Let this be 0. This is the input (in) in Figure 15. F ) is in F-min This is equivalent to setting the output to -128 when the value is smaller.
[0101] According to line 1609, the processor is in S32 If it is greater than 255 in S32 Let this be 255. This is because the input (in) in Figure 15 F ) is in F-max This is equivalent to setting the output to 127 if it is greater than the given value.
[0102] According to line 1610, the processor is in S32 Output the value obtained by subtracting 128 from (in S8 Substitute into ). Output (in S8 ) is a 3D matrix corresponding to h, w, and ci, and according to line 1610, the processor assigns values to the elements corresponding to h, w, and ci at runtime. S32 is the output (in S8 Because it has more bits than ), the number of bits is reduced during assignment, and generally the output (in S8 ) is in S32 It may not be the same as the value obtained by subtracting 128 from it. However, by processing according to lines 1608 and 1609 in S32 Since it is guaranteed to be between 0 and 255, in S32 The value obtained by subtracting 128 from this value is a number that can be represented as an 8-bit signed integer, and the output (in S8 ) is always in S32 This value is the same as the value obtained by subtracting 128 from that value.
[0103] The processing details of the 8-bit integer conversion processes 1421 and 1431 in Figure 14 are the same as those in Figure 16, F-min , and in F-max As mentioned above, the value of will differ for each layer. Additionally, the values of H, W, and CI may also differ for each layer.
[0104] Figure 17 shows the processing details of the floating-point conversion process 1413 in Figure 14.
[0105] Lines 1703-1705 are the same as lines 503-505 in Figure 5.
[0106] According to lines 1706 and 1707, the processor receives input (d1 S32 ) r F Multiply by s F Output the result after adding (d1 F Substitute into r. F , and s F The value will be described later. Input (d1 S32 ) and output (d1 F ) is a 3D matrix corresponding to h, w, and co, and according to lines 1706 and 1707, the processor uses the elements corresponding to h, w, and co at runtime. F , and s F This is a one-dimensional vector dependent on co, and according to line 1707, the processor uses the element corresponding to co at runtime.
[0107] The floating-point conversion processes 1423 and 1433 in Figure 14 are the same as those in Figure 17, but r F , and s F The value of will differ for each layer. Additionally, the values of H, W, and CO may also differ for each layer.
[0108] The constant (r) used in the floating-point conversion process shown in Figure 17. F , s F One example of a method for demonstrating how to calculate ) is to follow the methods described in Numbers 1 through 5.
number
[0109] Equation 1 is a mathematical expression of the integer conversion process shown in Figure 15. The symbol "=~" means equal except for the error. The error here is the difference between the input and the input. F-min If smaller, and in F-maxThis is due to the fact that the relationship between input and output deviates from a linear equation when the value is greater than the given value, and because the decimal part is truncated when converting floating-point numbers to integers.
number
[0110] Equation 2 is a mathematical expression of the integer conversion process shown in Figure 6. The meaning of the symbol "=~" is the same as in Equation 1.
number
[0111] Equation 3 is the result of substituting equations 1 and 2 into the mathematical expression (left side) of the convolution process shown in Figure 5.
number
[0112] Equation 4 is obtained by rearranging Equation 3 to find the result (left side) of the convolution operation calculated in floating-point format. The floating-point conversion process in Figure 17 is a process that converts the result of the convolution operation in Figure 5 to the result of the convolution operation calculated in floating-point format, so Equation 4 is the mathematical expression of the floating-point conversion process in Figure 17.
number
[0113] For number 5, from the correspondence between Figure 17 and number 4, r F and s F This is the value obtained.
[0114] In Math 1 to Math 5, F-min and in F-max This does not depend on h, w, and co in Figure 17. On the other hand, k F-max It does not depend on h and w in Figure 17, but it does depend on co. Therefore, r F , and s F This does not depend on h and w in Figure 17, but it does depend on co.
[0115] The processing details of bias addition processes 1414, 1424, and 1434 in Figure 14 are as follows: input (d1) S32 ), output (d2 S32 ) and bias value (b S32 This is a floating-point type representation of ).
[0116] Figures 18A and 18B show the processing details of the activation function processing 1415 in Figure 14.
[0117] Lines 1903-1905 are the same as lines 503-505 in Figure 5.
[0118] According to lines 1906 and 1907, the processor receives input (d2 F The function f is calculated for ) and the output (out F Substitute the value into (d2). Details of function f are shown in Figure 18B. Because function f uses an exponential function, it is not possible to calculate function f with integer data, and calculation must be done using floating-point type. Since the output may take a negative value, the integer conversion process in Figure 15 is designed to handle negative values as input. Input (d2 F ), and output (out F The matrix is a three-dimensional matrix corresponding to h, w, and co, and according to lines 1906 and 1907, the processor uses the elements corresponding to h, w, and co at runtime.
[0119] The activation function processes 1425 and 1435 in Figure 14 are similar to those in Figure 17, but the values of H, W, and CO may differ for each layer. [Examples]
[0120] Figure 19 shows a third example of a neural network processing configuration.
[0121] The neural network processing in Figure 19 is almost the same as in Figure 14, except that the convolution operation in Layer 2 (2002) is performed using floating-point numbers. Performing the convolution operation using integer numbers can reduce processing time compared to using floating-point numbers, but it may worsen the recognition results. In Figure 19, performing the convolution operation in Layer 2 (2002) using floating-point numbers may improve the recognition results compared to using integer numbers.
[0122] In the neural network processing shown in Figure 19, the 8-bit integer conversion processes 1411 and 1431 use coefficients corresponding to the cumulative frequency threshold selected using the method shown in Figure 12.
[0123] Although several embodiments have been described above, these are merely illustrative examples for explaining the present invention and are not intended to limit the scope of the invention to these embodiments only. The present invention can be implemented in various other forms. For example, the above embodiments can be summarized as follows. The following summary may include supplementary explanations for at least one embodiment, or may include explanations of modifications.
[0124] From the first perspective, there is a neural network processing device (108) comprising interface devices (103-107), memory devices (109, 110, and 112), and processors (102 and 111) connected to the interface devices and memory devices, wherein the processor performs neural network processing on data (e.g., input images) input via the interface devices. The neural network processing includes multiple convolution operations and integerization operations for each convolution operation. For example, there are multiple layers, and each layer includes integerization operations and convolution operations. For each of the multiple convolution operations, the integerization operation is an operation that performs integerization using coefficients. The coefficients are generated using statistical information of the input data (e.g., distribution of application frequency) and a cumulative frequency threshold for the integerization operation. For each convolution operation, the cumulative frequency threshold used to generate the coefficients used in the integerization operation may differ. Although the statistical information differs depending on the convolution process, an optimal cumulative frequency threshold corresponding to that statistical information can be used in the integer conversion process for each convolution, thus improving the recognition results obtained by neural network processing.
[0125] According to the second perspective, in the first perspective, for each of the multiple convolution operations, the processor may perform integerization on each of the multiple different cumulative frequency thresholds with respect only to that convolution operation, and not perform integerization on the convolution operations other than that one (for example, by processing floating-point data), and select the cumulative frequency threshold that produced the best result among the multiple different cumulative frequency thresholds used for integerization of that convolution operation. Furthermore, according to the third perspective, in the first perspective, for each of the multiple convolution operations, the processor may perform integerization on each of the multiple different cumulative frequency thresholds with respect to that convolution operation, and for the convolution operations other than that one, perform integerization using coefficients generated with a fixed cumulative frequency threshold regardless of which of the multiple different cumulative frequency thresholds is used, and select the cumulative frequency threshold that produced the best result among the multiple different cumulative frequency thresholds. In both the second and third perspectives, for each of the multiple integerization operations, the cumulative frequency threshold used to generate the coefficients may be the cumulative frequency threshold selected above for that integerization operation. If multiple different cumulative frequency thresholds can be applied simultaneously to each of the convolutional operations, then the number of neural network operations required for training the neural network will follow a number of X to the power of Y, where X is the number of cumulative frequency thresholds and Y is the number of convolutional operations (e.g., the number of layers). However, from a second perspective, since integer conversion is not performed for convolutional operations other than the one being trained, the number of neural network operations can be reduced from a number following X to the power of Y to a number following the product of X and Y. When the number of convolutional operations is large (e.g., several hundred to several thousand), it is expected that training can be achieved that reduces the time and processing load required for training while improving the recognition results.
[0126] According to the fourth viewpoint, in any of the first to third viewpoints, the processor may store the coefficients generated using the cumulative frequency threshold of the integerization process for each convolution process in a memory device. According to the fifth viewpoint, in any of the first to third viewpoints, the processor may receive the coefficients generated using the cumulative frequency threshold of the integerization process for each convolution process from a device outside the neural network processing unit and store them in a memory device. In either of the fourth and fifth viewpoints, the processor may, for each of a plurality of integerization processes, identify the coefficients corresponding to that integerization process from the memory device and perform the integerization process using the identified coefficients. In this case, since the integerization process can be performed using the coefficients in the memory device, there is no need to generate coefficients each time an integerization process is performed.
[0127] According to the sixth perspective, in the second or third perspective, the processor may, for each of the multiple convolution operations, process the input data through multiple other convolution operations that, unlike the multiple convolution operations, do not involve integer conversion. It may then compare the processing results obtained through the multiple convolution operations (1I, 2I, 3I) with the processing results obtained through the other multiple convolution operations (F), and based on the results of this comparison, select the cumulative frequency threshold that yielded the best processing result. This allows learning to be performed using the recognition result (F) as the correct answer, thus eliminating the need to prepare correct answer data in advance.
[0128] According to the seventh perspective, in the second or third perspective, the processor may, for each of the multiple convolution operations, compare the result of processing the input data through multiple convolution operations with the correct result of processing the input data, and based on the result of this comparison, select the cumulative frequency threshold that yielded the best result. This enables learning without the need for multiple other convolution operations to produce the recognition result (F).
[0129] According to the eighth perspective, in any of the first to seventh perspectives, the processor may store the data input during a first period (e.g., during operation) in memory, which is the period during which inference including neural network processing is performed. During a second period (e.g., at night) in which no inference including neural network processing is performed, the processor may, for each of the multiple convolutional operations, process the input data stored in memory during the first period through multiple other convolutional operations, in addition to the multiple convolutional operations, each of which does not involve integerization. The processor may then compare the processing results obtained through the multiple convolutional operations with the processing results obtained through the other multiple convolutional operations, and based on the results of this comparison, select the cumulative frequency threshold that yielded the best processing result. This allows the system to determine a cumulative frequency threshold for each convolutional operation during learning, and then, for a particular convolutional operation, identify when the determined cumulative frequency threshold has become inappropriate (for example, by identifying that the discrepancy between the recognition result (F) and the recognition result (I) for that particular convolutional operation is greater than a predetermined discrepancy), and update the cumulative frequency threshold to an appropriate value. In other words, it is possible to maintain good recognition results in neural network processing. The cumulative frequency threshold may be stored in memory after each convolutional operation. [Explanation of Symbols]
[0130] 108 Neural Network Processing Units 211, 221, 231 8-bit integer conversion process 212, 222, 232 Convolution 1411, 1421, 1431 8-bit integer conversion process
Claims
1. A neural network processing device comprising an interface device, a storage device, and a processor connected to the interface device and the storage device, wherein the processor performs neural network processing of data input via the interface device, The aforementioned neural network processing includes multiple convolution operations and integerization operations for each convolution operation. For each of the aforementioned convolution processes, the integerization process is a process that performs integerization using a coefficient generated with respect to the statistical information of the input data and the cumulative frequency threshold for the integerization process. For each convolution operation, the cumulative frequency threshold used to generate the coefficients used in the integer conversion operation may differ. A neural network processing device characterized by the following:
2. In the neural network processing device according to claim 1, For each of the aforementioned convolution processes, the processor: For the convolution process only, integer conversion is performed for each of the multiple different cumulative frequency thresholds. For convolution operations other than the convolution operation in question, the integer conversion operation is not performed. Of the multiple different cumulative frequency thresholds used in the integerization process of the convolution, the cumulative frequency threshold that yielded the best processing result is selected. For each of the aforementioned integerization processes, the cumulative frequency threshold used to generate the coefficient is the cumulative frequency threshold selected for that integerization process. A neural network processing device characterized by the following:
3. In the neural network processing device according to claim 1, For each of the aforementioned convolution processes, the processor: Regarding the convolution process, integer conversion is performed for each of the multiple different cumulative frequency thresholds. With respect to convolution operations other than the convolution operation in question, regardless of which of the multiple different cumulative frequency thresholds is used, an integer conversion operation is performed using coefficients generated with a fixed cumulative frequency threshold. Select the cumulative frequency threshold that yielded the best processing result from among the aforementioned multiple different cumulative frequency thresholds. For each of the aforementioned integerization processes, the cumulative frequency threshold used to generate the coefficient is the cumulative frequency threshold selected for that integerization process. A neural network processing device characterized by the following:
4. In the neural network processing device according to claim 1, The processor stores in the memory a coefficient generated using the cumulative frequency threshold of the integerization process for each convolution process. The processor, in each of the plurality of integerization processes, identifies a coefficient corresponding to the integerization process from the storage device and performs the integerization process using the identified coefficient. A neural network processing device characterized by the following:
5. In the neural network processing device according to claim 1, The processor receives coefficients generated using the cumulative frequency threshold of the integerization process for each convolution process from a device outside the neural network processing device and stores them in the storage device. The processor, in each of the plurality of integerization processes, identifies a coefficient corresponding to the integerization process from the storage device and performs the integerization process using the identified coefficient. A neural network processing device characterized by the following:
6. In the neural network processing apparatus according to claim 2 or 3, The processor performs the following for each of the plurality of convolution processes: The input data is processed not only through the aforementioned multiple convolution processes, but also through a plurality of other convolution processes that differ from the aforementioned multiple convolution processes in that they do not involve integer conversion. The processing results obtained through the aforementioned multiple convolution processes are compared with the processing results obtained through the aforementioned other multiple convolution processes. Based on the results of this comparison, select the cumulative frequency threshold that yielded the best processing results. A neural network processing device characterized by the following:
7. In the neural network processing apparatus according to claim 2 or 3, The processor performs the following for each of the plurality of convolution processes: The processing result of the input data through the multiple convolution processes is compared with the correct processing result of the input data. Based on the results of this comparison, select the cumulative frequency threshold that yielded the best processing results. A neural network processing device characterized by the following:
8. In the neural network processing device according to claim 1, The processor stores the data input during the first period, which is the period during which inference including the neural network processing is performed, in the storage device. During a second period in which no inference, including the neural network processing, is performed, the processor performs the following for each of the plurality of convolutional processes: The input data stored in the storage device during the first period is processed not only through the aforementioned convolution processes, but also through a plurality of other convolution processes that differ from the aforementioned convolution processes in that they do not involve integer conversion. The processing results obtained through the aforementioned multiple convolution processes are compared with the processing results obtained through the aforementioned other multiple convolution processes. Based on the results of this comparison, select the cumulative frequency threshold that yielded the best processing results. A neural network processing device characterized by the following:
9. For each of the multiple convolutional operations in a neural network, The cumulative frequency threshold, which may vary depending on the convolution process, is determined and used to generate the coefficients used in the integer conversion process for the convolution. The data is integerized using coefficients generated from statistical information of the input data and a cumulative frequency threshold for the integerization process applied to the convolution. A neural network processing method characterized by performing the following using a computer.
Citation Information
Patent Citations
Adaptive quantization method, apparatus, device, and medium
JP2022507704A