Learning method and learning device

The learning apparatus and method address errors in converting CNN functional models to neural network circuits by using a dual forward propagation approach with reduced bit operations and backpropagation to minimize loss, resulting in reduced errors and improved accuracy.

JP2025096942APending Publication Date: 2025-06-30MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023212953
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

Existing methods for converting convolutional neural network (CNN) functional models and learned parameters into operations executable in lightweight neural network circuits often result in errors due to differences in operation accuracy and data format.

Method used

A learning apparatus and method that includes a learning data acquisition unit, a first forward propagation unit, a second forward propagation unit with reduced bit operations, a calculation unit to determine a threshold value, a backpropagation unit, and a parameter determination unit to minimize loss and reduce errors between functional model and neural network circuit results.

Benefits of technology

The proposed solution effectively reduces the error between the calculation results of the CNN functional model and the neural network circuit, ensuring more accurate operations in lightweight neural network circuits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096942000001_ABST
    Figure 2025096942000001_ABST
Patent Text Reader

Abstract

To reduce an error between a result of calculation by a functional model and a result of calculation by a network circuit.SOLUTION: A learning device includes: a learning data acquisition unit which acquires learning data; a first forward propagation unit which obtains loss by forward propagating the information acquired by the learning data acquisition unit in a neural network being a learning object; a second forward propagation unit which is generated on the basis of the first forward propagation unit, and obtains loss by performing calculation with a bit number smaller than the number of calculation bits of the first forward propagation unit; a calculation unit which calculates difference between the loss obtained by the first forward propagation unit and the loss obtained by the second forward propagation unit; a back propagation unit which adds the difference calculated by the calculation unit, and makes it back propagation in the neural network; and a parameter determination unit which determines the parameter of the neural network so as to reduce the loss being the result of the back propagation performed by the back propagation unit.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning method and a learning device.

Background Art

[0002] In recent years, as a method for recognizing some pattern from an image, a convolutional neural network (CNN) is known. In embedded devices such as IoT devices, a lightweight neural network circuit that can be embedded in an edge device is used (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] On the other hand, in order to determine the configuration and specifications of a convolutional neural network, generate a functional model of the convolutional neural network, and generate learned parameters learned using the functional model, known libraries and platforms are used. When converting the functional model of the neural network and the learned parameters generated in such a library or platform into operations that can be performed in a lightweight neural network circuit and performing the operations, an error may occur in the operation result due to differences in operation accuracy and data format.

[0005] In view of such circumstances, an object of the present invention is to provide a learning apparatus and a learning method in which, when a functional model of a neural network and learned parameters learned using the functional model are converted into operations executable in a neural network circuit and the operations are performed, an error hardly occurs between the operation result by the functional model and the operation result by the neural network circuit.

Means for Solving the Problems

[0006] [1] To solve the above problems, one aspect of the present invention is a learning data acquisition unit that acquires learning data, and a first forward propagation unit that obtains a loss by forward propagating the information acquired by the learning data acquisition unit through a neural network to be learned, a second forward propagation unit that is generated based on the first forward propagation unit and obtains a loss by performing an operation with a number of bits smaller than the number of operation bits of the first forward propagation unit, a calculation unit that calculates a threshold value from the loss obtained by the first forward propagation unit and the loss obtained by the second forward propagation unit, a backpropagation unit that adds the threshold value calculated by the calculation unit and backpropagates the neural network, and a parameter determination unit that determines the parameters of the neural network so that the loss decreases as a result of the backpropagation performed by the backpropagation unit.

[0007] [2] Further, one aspect of the present invention is the learning apparatus according to [1] above, wherein the operation order of the first forward propagation unit and the operation order of the second forward propagation unit are different from each other.

[0008] [3] Further, one aspect of the present invention is the learning apparatus according to [2] above, wherein the parameters determined by the parameter determination unit are those implemented on an accelerator, and the operation order of the second forward propagation unit is different depending on the accelerator to be implemented.

[0009] [4] Further, one aspect of the present invention is the learning device described in [1] above, wherein the parameters determined by the parameter determination unit are implemented on an accelerator, and the arithmetic expression of the second forward propagation unit is different depending on the implemented accelerator.

[0010] [5] Further, one aspect of the present invention is the learning device described in [4] above, wherein an arithmetic expression corresponding to the implemented accelerator is prepared in advance, and the arithmetic expression of the second forward propagation unit is selected according to the implemented accelerator.

[0011] [6] Further, one aspect of the present invention is a learning method including a learning data acquisition step of acquiring learning data, a first forward propagation step of obtaining a loss by forward propagating the information acquired in the learning data acquisition step through a neural network to be learned, a second forward propagation step of obtaining a loss using a second model generated based on the first model used in the first forward propagation step, wherein the number of arithmetic bits of the second model is smaller than that of the first model, a calculation step of calculating a threshold value from the loss obtained in the first forward propagation step and the loss obtained in the second forward propagation step, a backpropagation step of adding the threshold value calculated in the calculation step and backpropagating the neural network, and a parameter determination step of determining the parameters of the neural network so that the loss decreases as a result of the backpropagation performed in the backpropagation step.

Effect of the Invention

[0012] According to the present invention, it is possible to reduce the error between the calculation result by the functional model and the calculation result by the neural network circuit.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0014] [Embodiment] Hereinafter, a preferred embodiment of the learning device and the learning method according to an aspect of the present invention will be described in detail with reference to the accompanying drawings. Note that the aspects of the present invention are not limited to these embodiments, but also include those with various modifications or improvements. That is, the components described below include those that can be easily assumed by those skilled in the art and substantially the same components, and the components described below can be combined as appropriate. Also, various omissions, substitutions, or changes of the components can be made without departing from the gist of the present invention. Also, in the following drawings, in order to make each configuration easy to understand, the scale and number, etc. in each structure may be different from the scale and number, etc. in the actual structure.

[0015] FIG. 1 is a diagram showing a neural network learning apparatus according to an embodiment. First, the neural network learning apparatus 300 will be described with reference to this figure. The neural network learning apparatus 300 generates and learns a convolutional neural network 200 (hereinafter also referred to as "CNN 200" or "NN functional model 200") which is a neural network functional model, and generates software 500 for operating a neural network circuit 100 (hereinafter also referred to as "NN circuit 100") that can be incorporated into embedded devices such as IoT devices. The operations executed by the NN circuit 100 are at least a part of the inference operations executed by the CNN 200 (NN functional model 200).

[0016] The neural network learning apparatus 300 is a program-executable apparatus (computer) including a processor such as a CPU (Central Processing Unit) and hardware such as a memory. The functions of the neural network learning apparatus 300 are realized by executing a neural network learning program and a software generation program in the neural network learning apparatus 300. The neural network learning apparatus 300 includes a storage unit 310, an arithmetic unit 320, a data input unit 330, a data output unit 340, a display unit 350, and an operation input unit 360.

[0017] The storage unit 310 stores network information NW1, inference network information NW2, a learning data set DS, and learned parameters PM. The learning data set DS and the inference network information NW2 are input data input to the neural network learning apparatus 300. The learned parameters PM are output data output by the neural network learning apparatus 300. Note that the "learned NN circuit 100" includes the NN circuit 100 and the learned parameters PM.

[0018] The network information (learning network information) NW1 is information regarding the CNN200 (NN function model 200). The network information NW1 includes, for example, information defining the functions of the CNN200 (NN function model 200). The network information NW1 is, for example, the network configuration of the CNN200, input data information, output data information, quantization information, and the like. The input data information includes the input data type such as an image or voice, and the input data size and the like.

[0019] The inference network information NW2 is information regarding the inference operation executed by the NN circuit 100. The inference network information NW2 includes, for example, information defining the functions of the inference operations of the neural network executable by the NN circuit 100. The inference network information NW2 is, for example, the circuit configuration of the NN circuit 100, the functions of the arithmetic units, the data bit width, and the like.

[0020] The learning dataset DS has learning data D1 used for learning and test data D2 used for inference testing.

[0021] The arithmetic unit 320 includes a learning unit 322, an inference unit 323, a software generation unit 325, and a function model generation unit 326. Network information NW is input to the arithmetic unit 320, and a learned parameter PM is output. The network information NW input to the arithmetic unit 320 may be generated by a device other than the neural network learning device 300.

[0022] FIG. 2 is a diagram for explaining the input and output of the arithmetic unit of the neural network learning device according to the present embodiment. Next, with reference to the same figure, the configuration of the arithmetic unit 320 and the information input to and output from the configuration will be described.

[0023] The learning unit 322 generates a learned parameter PM using the network information NW1, the inference network information NW2, and the learning data D1. The inference unit 323 performs an inference test using the network information NW and the test data D2.

[0024] The software generation unit 325 generates software 500 that operates the NN circuit 100 based on the network information NW1 and the inference network information NW2. The software 500 includes software that transfers the learned parameters PM to the NN circuit 100 as necessary.

[0025] The function model generation unit 326 generates (configures) the CNN 200 (NN function model 200) based on an input from the user, and outputs network information NW1, which is information regarding the CNN 200 (NN function model 200).

[0026] Returning to FIG. 1, the data input unit 330 receives hardware information HW, network information NW, etc. necessary for generating the learned NN circuit 100. The hardware information HW, network information NW, etc. are input as data described in a predetermined data format, for example. The input hardware information HW, network information NW, etc. are stored in the storage unit 310. The hardware information HW, network information NW, etc. may be input or changed by the user from the operation input unit 360.

[0027] The generated learned NN circuit 100 is output to the data output unit 340. For example, the generated NN circuit 100 and the learned parameters PM are output to the data output unit 340.

[0028] The display unit 350 has a known monitor such as an LCD display. The display unit 350 can display a console screen or the like for receiving GUI (Graphical User Interface) images, commands, etc. generated by the arithmetic unit 320. Further, when the arithmetic unit 320 requires information input from the user, the display unit 350 can display a message prompting the user to input information from the operation input unit 360 and a GUI image necessary for information input.

[0029] The operation input unit 360 is a device for the user to input instructions to the arithmetic unit 320 and the like. The operation input unit 360 is a known input device such as a touch panel, a keyboard, a mouse, etc. The input of the operation input unit 360 is transmitted to the arithmetic unit 320.

[0030] All or part of the functions of the arithmetic unit 320 are realized by one or more processors such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) executing a program stored in the program memory. However, all or part of the functions of the arithmetic unit 320 may be realized by hardware (for example, a circuit unit; circuity) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a PLD (Programmable Logic Device). Also, all or part of the functions of the arithmetic unit 320 may be realized by a combination of software and hardware.

[0031] All or part of the functions of the arithmetic unit 320 may be realized by using an external accelerator such as a CPU, a GPU, or hardware provided in an external device such as a cloud server. The arithmetic unit 320 can improve the arithmetic speed of the arithmetic unit 320, for example, by using a GPU with high arithmetic performance or dedicated hardware on a cloud server in combination.

[0032] The storage unit 310 is realized by a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), a ROM (Read-Only Memory), a RAM (Random Access Memory), or the like. All or part of the storage unit 310 may be provided in an external device such as a cloud server and connected to the arithmetic unit 320 and the like via a communication line.

[0033] Note that the neural network learning device 300 is composed of a plurality of devices (computers), and the functional blocks of the arithmetic unit 320 may be distributed among the plurality of devices. For example, the neural network learning device 300 may be separated into a first device (computer) having a function model generation unit 326, a second device (computer) having a learning unit 322 and an inference unit 323, and a third device (computer) having a software generation unit 325.

[0034] FIG. 3 is a diagram showing an example of the layer configuration of the convolutional neural network according to the present embodiment. With reference to this figure, an example of the layer configuration of the CNN 200 will be described. The CNN 200 is a multi-layer network including a convolutional layer 210 that performs a convolutional operation, a quantization operation layer 220 that performs a quantization operation, and an output layer 230. In at least a part of the CNN 200, the convolutional layer 210 and the quantization operation layer 220 are alternately connected. The CNN 200 is a model widely used for image recognition and video recognition. The CNN 200 may further include a layer (layer) having other functions such as a fully connected layer.

[0035] FIG. 4 is a diagram for explaining the input and output of each quantization convolutional operation block included in the convolutional neural network according to the present embodiment. With reference to this figure, the convolutional operation performed by the convolutional layer 210 will be described. The convolutional layer 210 performs a convolutional operation on the input data a using the weight w. The convolutional layer 210 performs a sum-of-products operation taking the input data a and the weight w as inputs.

[0036] The input data a (also referred to as activation data or feature map) to the convolutional layer 210 is multi-dimensional data such as image data. In the present embodiment, the input data a is a three-dimensional tensor composed of elements (x, y, c). The convolutional layer 210 of the CNN 200 performs a convolutional operation on the low-bit input data a. In the present embodiment, the elements of the input data a are 2-bit unsigned integers (0, 1, 2, 3). The elements of the input data a may be, for example, 4-bit or 8-bit unsigned integers.

[0037] When the input data input to CNN200 has a different format from the input data a to the convolutional layer 210, such as 32-bit floating-point type, etc., CNN200 may further have an input layer that performs type conversion or quantization before the convolutional layer 210.

[0038] The weights w (also referred to as filters or kernels) of the convolutional layer 210 are multi-dimensional data having elements that are learnable parameters. In the present embodiment, the weights w are 4D tensors consisting of elements (i, j, c, d). The weights w have d 3D tensors (hereinafter referred to as "weights wo") consisting of elements (i, j, c). The weights w in the learned CNN200 are learned data. The convolutional layer 210 of CNN200 performs a convolutional operation using low-bit weights w. In the present embodiment, the elements of the weights w are 1-bit signed integers (0, 1), where the value "0" represents +1 and the value "1" represents -1.

[0039] The convolutional layer 210 performs a predetermined convolutional operation and outputs output data f to the quantization operation layer 220.

[0040] The quantization operation layer 220 performs quantization operations and the like on the output of the convolutional operation output by the convolutional layer 210. The quantization operation layer 220 includes an pooling layer 221, a Batch Normalization layer 222, an activation function layer 223, and a quantization layer 224.

[0041] The pooling layer 221 compresses the output data f of the convolutional layer 210 by performing operations such as average pooling or MAX pooling on the output data f of the convolutional operation output by the convolutional layer 210.

[0042] The Batch Normalization layer 222 normalizes the data distribution by performing a predetermined operation on the output data of the quantization operation layer 220 and the pooling layer 221.

[0043] The activation function layer 223 performs operations of activation functions such as ReLU (Rectified Linear Unit) on the outputs of the quantization operation layer 220, the pooling layer 221, and the Batch Normalization layer 222.

[0044] The quantization layer 224 performs a quantization operation on the outputs of the pooling layer 221 and the activation function layer 223 based on quantization parameters. The quantization operation is, for example, one that reduces the input tensor u to 2 bits.

[0045] The output layer 230 is a layer that outputs the result of the CNN 200 by an identity function, a softmax function, or the like. The layer before the output layer 230 may be the convolutional layer 210 or the quantization operation layer 220.

[0046] In the CNN 200, since the output data of the quantized quantization layer 224 is input to the convolutional layer 210, the load of the convolutional operation in the convolutional layer 210 is smaller compared to other convolutional neural networks that do not perform quantization.

[0047] Returning to FIG. 1, the function model generation unit 326 sets the network structure and the specifications for each layer in the CNN 200 (NN function model 200) based on the input from the user via the operation input unit 360. For example, the user changes the network structure of the NN function model 200 by reconnecting the visually diagrammed layer connections displayed as a GUI image. Also, the user changes the specifications (input data information, output data information, quantization information, etc.) for each layer visually diagrammed and displayed as a GUI image. For example, the user can reconnect the connections between the pooling layer 221, the Batch Normalization layer 222, the activation function layer 223, and the quantization layer 224 in the quantization operation layer 220.

[0048] Note that the network structure and specifications for each layer in the CNN200 (NN function model 200) do not have to be described in a visually diagrammatic form as illustrated in FIG. 4. The network structure and specifications for each layer in the CNN200 may be described by a programming language, XML, or the like.

[0049] The CNN200 (NN function model 200) generated by the function model generation unit 326 is a neural network function model capable of performing learning and inference operations in the arithmetic unit 320 (learning unit 322 and inference unit 323) of the neural network learning device 300. The arithmetic unit 320 of the neural network learning device 300 includes an arithmetic circuit with higher performance than the arithmetic circuit provided in the NN circuit 100, such as a CPU, GPU, or dedicated hardware. Therefore, the CNN200 generated by the function model generation unit 326 may include an arithmetic block (hereinafter also referred to as a "convertible arithmetic block") that can be converted into an operation executable in the NN circuit 100 and an arithmetic block (hereinafter also referred to as a "non-convertible arithmetic block") that cannot be converted into an operation executable in the NN circuit 100. Here, an arithmetic block is a plurality of consecutive operations in the CNN200.

[0050] In order for the CNN200 (NN function model 200) to be efficiently inference-operated by the NN circuit 100, it is desirable for the function model generation unit 326 to generate more arithmetic blocks (convertible arithmetic blocks) that can be converted into operations executable in the NN circuit 100.

[0051] As shown in FIG. 4, an arithmetic block from a convolution operation to a quantization operation in the CNN200 (NN function model 200) is defined as a "quantized convolution operation block QC". At least a part of the CNN200 is configured by connecting a plurality of quantized convolution operation blocks QC.

[0052] FIG. 5 is a diagram showing an example of the configuration of an inference operation block in the neural network circuit according to the present embodiment. With reference to this figure, an example of the configuration of the inference operation block EB in the NN circuit 100 will be described. Note that the inference operation block EB is an operation block included in hardware such as an edge device, and is an example of an operation environment used at the time of inference.

[0053] "C" shown in this figure represents the sum-of-products operation in the convolutional operation circuit. "AW" shown in this figure is data obtained by multiplying the input vector A and the weight matrix W, and is vector data of 16-bit integers per element. Also, "Q" shown in this figure represents the quantization operation. "U" shown in this figure is data obtained by quantizing AW, and is vector data of 2-bit integers per element.

[0054] Note that the operation environment in the inference operation block EB (that is, the operation environment at the time of inference) is less accurate compared to the operation environment in the quantization convolutional operation block QC (that is, the operation environment at the time of learning) described later. The operation environment includes, in addition to the operation accuracy, data format, operation order, and the like.

[0055] FIG. 6 is a diagram showing an example of the configuration of a quantization convolutional operation block in the convolutional neural network according to the present embodiment. With reference to this figure, an example of the configuration of the quantization convolutional operation block QC in the NN circuit 100 will be described. Note that the quantization convolutional operation block QC is an operation block used at the time of learning, and is an example of an operation environment used at the time of learning. Also, it can be said that the quantization convolutional operation block QC shows the configuration on the NDK (Network Development Kit).

[0056] The quantization convolution operation block QC shown in the figure is configured as a convertible operation block. An input vector A and a weight matrix W are input, and an output vector U quantized to 2 bits per element is output. "X1" shown in the figure is an operation that performs an affine transformation operation with the first scaling coefficients (first scaling factors) Sa1 and Sb1 in floating-point format as coefficients on the input vector A (Sa1 × A + Sb1), and outputs vector data As in floating-point format (post-scaler after quantization). When the quantization convolution operation block QC is configured as a convertible operation block, the input data is limited to an input vector A of 2 bits per element. Even in this case, by performing an affine transformation on the input vector A with the first scaling coefficients Sa1 and Sb1 as coefficients, a decrease in the accuracy of the input data can be suppressed.

[0057] Also, "X2" shown in the figure is an operation that performs an affine transformation operation with the second scaling coefficients (second scaling factors) Sa2 and Sb2 in floating-point format as coefficients on the weight matrix W (Sa2 × W + Sb2), and outputs matrix data Ws in floating-point format. When the quantization convolution operation block QC is configured as a convertible operation block, the weights are limited to a weight matrix W of 1 bit per element. Even in this case, by performing an affine transformation on the weight matrix W with the scaling coefficients Sa2 and Sb2 as coefficients, a decrease in the accuracy of the weights can be suppressed.

[0058] Also, "Cf" shown in the figure is a convolution operation that multiplies As and Ws and outputs vector data AWs in floating-point format. Also, "X3" shown in the figure is an operation that performs an affine transformation operation with the third scaling coefficients (third scaling factors) Sa3 and Sb3 in floating-point format as coefficients on the vector data AWs (Sa3 × AWs + Sb3), and outputs vector data AWss in floating-point format (pre-scaler before quantization). For example, "X3" is a pre-scaler corresponding to "X1" (post-scaler after quantization).

[0059] Also, "Qf" shown in the same figure indicates a quantization operation that quantizes vector data AWss in floating-point format based on quantization parameter qf(thf0, thf1, thf2) and outputs vector data U of 2-bit integers per element. The quantization parameter qf is the threshold (thf0, thf1, thf2) in floating-point format. When the quantization convolution operation block QC is configured as a convertible operation block, the output data is limited to vector data U of 2 bits per element.

[0060] The quantization convolution operation block QC can be treated as a convertible operation block that can be converted into an operation executable in the inference operation block EB by incorporating and aggregating the scaling factors (Sa1, Sb1, Sa2, Sb2, Sa3, Sb3) into the quantization parameter qf(thf0, thf1, thf2) in the quantization operation Qf. For example, when Sa1 is 1.5, Sa2 is 2.0, and Sa3 is 1.1, by updating the quantization parameter qf(thf0, thf1, thf2) in the quantization operation Q to a value 1 / 3.3 times the original quantization parameter, the scaling factors are aggregated into the quantization parameter qf.

[0061] Even when other types of operations P are added, the quantization convolution operation block QC can be configured as a convertible operation block according to the type for operation P. For example, as described above, operations such as Batch Normalization and activation functions can be incorporated and aggregated into the quantization parameter qf(thf0,thf1,thf2). Also, the addition of a bias value to the convolution operation result can be incorporated and aggregated into the quantization parameter qf by subtracting the bias value from the quantization parameter qf. Therefore, even when other types of operations P such as Batch Normalization, activation functions, and addition of bias values are added, the quantization convolution operation block QC can be configured as a convertible operation block. If operation P is an operation that cannot be incorporated and aggregated into the quantization parameter qf, the operation block including operation P becomes a non-convertible operation block.

[0062] When operation P includes a plurality of floating-point operations, it is desirable that the plurality of floating-point operations be performed in an order in which rounding errors are less likely to occur. This is because if rounding errors are likely to occur, errors are likely to occur between the operation result by the quantization convolution operation block QC and the operation result by the inference operation block EB due to the variation of the rounding errors described later.

[0063] FIG. 7 is a diagram for explaining the learning stage of the neural network learning apparatus according to the present embodiment. With reference to this figure, an example of the learning process performed by the learning apparatus 1 will be described. Here, the computing environment in the learning stage and the computing environment in the inference stage are different, and it is generally known that the computing environment in the learning stage has higher computing accuracy. Conventionally, even if learning is performed in a computing environment with higher accuracy, there has been a problem that the inference accuracy decreases when the inference environment is incorporated into an edge device or the like with low computing accuracy. In the present embodiment, in the learning stage, by performing learning in consideration of the computing environment in the inference stage, the inference accuracy can be improved. The learning apparatus 1 includes a learning data acquisition unit 10, a first forward propagation unit 11, a second forward propagation unit 12, a calculation unit 13, and a backpropagation unit 14.

[0064] The learning data acquisition unit 10 acquires learning data to be used for learning. When the purpose is object detection, the learning data may be, for example, image data and information in which the position and class of an object reflected in the image data are associated.

[0065] The first forward propagation unit 11 obtains the output of the neural network from the final layer as a result by forward-propagating the information acquired by the learning data acquisition unit 10 through the neural network to be learned.

[0066] The second forward propagation unit 12 obtains the output of the neural network from the final layer as a result by forward-propagating the information acquired by the learning data acquisition unit 10 through the neural network to be learned.

[0067] Here, the first forward propagation unit 11 and the second forward propagation unit 12 share the value of the parameter updated by the parameter determination unit 15. When the parameter is updated by the parameter determination unit 15, the parameters referred to by the first forward propagation unit 11 and the second forward propagation unit 12 are both updated. The parameter referred to by the second forward propagation unit 12 is the parameter referred to by the first forward propagation unit 11 with the precision reduced to the number of operation bits of the second forward propagation unit 12. That is, the first forward propagation unit 11 and the second forward propagation unit 12 have different numbers of operation bits. Specifically, the number of operation bits of the first forward propagation unit 11 is larger than the number of operation bits of the second forward propagation unit 12. More specifically, the first forward propagation unit 11 may be Fixed (fixed-point), and the second forward propagation unit 12 may be Float (floating-point). By sharing the parameter between the first forward propagation unit 11 and the second forward propagation unit 12, two forward propagation units with different numbers of operation bits but the same essence can be created.

[0068] Note that the second forward propagation unit 12 is assumed for an edge device on which the neural network is implemented. The second forward propagation unit 12 may be generated based on the first forward propagation unit 11. The operation order and operation formula of the second forward propagation unit 12 may be determined assuming an edge device on which the neural network is implemented. That is, the first forward propagation unit 11 and the second forward propagation unit 12 may have different operation orders and operation formulas from each other. In the following description, the model performed by the first forward propagation unit 11 may be described as the first model, and the model performed by the second forward propagation unit 12 may be described as the second model.

[0069] The calculation unit 13 calculates a predetermined threshold value from the output obtained by the first forward propagation unit 11 and the output obtained by the second forward propagation unit 12. As an example of the threshold value, the difference between the output obtained by the first forward propagation unit 11 and the output obtained by the second forward propagation unit 12 is calculated. This difference can also be described as Threshold. This difference can also be described as a constant term. Note that this difference is calculated for applying this embodiment to frameworks such as PyTorch and TensorFlow. Therefore, in the case of the configuration having the second forward propagation unit 12 as shown in FIG. 7, the difference does not necessarily have to be calculated. In this case, the calculation may be performed by the backpropagation unit 14 based on the loss calculated by the second forward propagation unit 12.

[0070] The backpropagation unit 14 adds the difference (Threshold) calculated by the calculation unit 13 to the first forward propagation unit 11. As a result, the value of the first forward propagation unit 11 becomes equal to the value of the second forward propagation unit 12. This means that the inference result is equal to the second forward propagation unit 12. In frameworks such as PyTorch and TensorFlow, by performing this addition, the update of the parameters can be performed based on the first forward propagation unit 11. By adopting such a configuration, learning in which the inference result is based on the second forward propagation unit 12 and the parameter update is based on the first forward propagation unit 11 can also be realized in frameworks such as PyTorch and TensorFlow. After the difference calculated by the calculation unit 13 is added, the neural network is backpropagated. Note that the neural network is defined by the first forward propagation unit 11. The backpropagation unit 14 may be of type Float (floating point). The addition of the difference (Threshold) may be, for example, directly adding the coordinate value obtained as the difference (shifting the coordinates). Also, in the case of a model for performing object detection, the addition of the difference (Threshold) may be an operation of changing the likelihood value (weighting the likelihood).

[0071] The parameter determination unit 15 determines the parameters of the neural network so that the loss decreases as a result of the backpropagation performed by the backpropagation unit 14. Based on the parameters determined by the parameter determination unit 15, both the parameters of the first forward propagation unit 11 and the second forward propagation unit 12 are updated.

[0072] In addition, when there are a plurality of accelerators that may be implemented, arithmetic expressions corresponding to the accelerators may be prepared in advance. In this case, the arithmetic expression, and the arithmetic expression of the second forward propagation unit 12, may be selected according to the implemented accelerator.

[0073] FIG. 8 is a diagram for explaining the relationship between the first model and the second model used by the neural network learning device according to the present embodiment in the learning stage. With reference to this figure, the relationship between the first model and the second model will be described. The first model is a model for performing learning in a computing environment having rich resources. The second model is a model assumed to perform inference on an accelerator mounted on an edge device. Although the second model is actually executed in a computing environment having rich resources similar to the first model, since it is assumed to perform inference on an accelerator mounted on an edge device, the number of arithmetic bits is deliberately reduced to Fixed (fixed-point).

[0074] Here, conventionally, when performing learning in a computing environment having rich resources, forward propagation was performed using a Float (floating-point) model, and similarly, backpropagation was performed using a float model. Learning was performed by repeating such steps. On the other hand, in the present embodiment, forward propagation is performed using a Float (floating-point) model and a Fixed (fixed-point) model, respectively, a difference (threshold or constant term) is added to the calculation result of the Float (floating-point) model, and backpropagation using the Float (floating-point) model is performed. By repeating such steps, learning is performed assuming inference on an accelerator mounted on an edge device.

[0075] Note that the addition of the difference (threshold or constant term) may be the difference between the outputs of the last layers (in the example shown, the quantization convolutional operation block QC1n and the quantization convolutional operation block QC2n). However, this embodiment is not limited to this example, and the difference may be calculated for each quantization convolutional operation block QC, and the addition of the difference may be performed.

[0076] FIG. 9 is a block diagram showing an example of the internal configuration of the learning device 1 according to the present embodiment. At least some functions of the learning device 1 can be realized using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be realized using existing technologies. The central processing unit 901 executes instructions included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic operations and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory". The input / output port 903 is a port for the central processing unit 901 to exchange data with external input / output devices and the like. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data in the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906. Also, all or part of each functional unit included in the learning device 1 may be realized using hardware (for example, a circuit part; circuitry) such as an ASIC, a PLD, or an FPGA. Also, all or part of each functional unit may be realized by a combination of software and hardware.

[0077] [Summary of Embodiment] According to the embodiments described above, the learning device 1 acquires learning data by including the learning data acquisition unit 10. Further, the learning device 1 includes the first forward propagation unit 11, and obtains a loss by forward propagating the information acquired by the learning data acquisition unit 10 through the neural network to be learned. Also, the learning device 1 includes the second forward propagation unit 12 generated based on the first forward propagation unit 11, and obtains a loss by forward propagating the information acquired by the learning data acquisition unit 10 through the neural network to be learned. Note that the number of operation bits of the second forward propagation unit 12 is smaller than the number of operation bits of the first forward propagation unit 11. Furthermore, the learning device 1 includes the calculation unit 13 to calculate the difference between the loss obtained by the first forward propagation unit 11 and the loss obtained by the second forward propagation unit 12, and includes the backpropagation unit 14 to add the difference calculated by the calculation unit 13 and backpropagate the neural network. Also, the learning device 1 includes the parameter determination unit 15 to determine the parameters of the neural network so that the loss decreases as a result of the backpropagation performed by the backpropagation unit 14. By adopting such a configuration, the learning device 1 can reduce the error between the operation result by the functional model and the operation result by the neural network circuit.

[0078] Note that all or part of the functions of each unit included in the learning device according to the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, reading the program recorded on this recording medium into a computer system, and executing it. Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices.

[0079] The "computer-readable recording medium" refers to a portable medium such as a flexible disk, magneto-optical disk, ROM, CD-ROM, etc., or a storage unit such as a hard disk built into a computer system. Further, the "computer-readable recording medium" also includes those that dynamically hold a program for a short time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and those that hold a program for a certain period of time, such as volatile memory inside a computer system serving as a server or client in that case. Also, the above program may be for realizing a part of the functions described above, and may further be capable of realizing the functions described above in combination with a program already recorded in a computer system.

[0080] As described above, the embodiments for implementing the present invention have been described using embodiments. However, the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the spirit of the present invention.

Explanation of Reference Numerals

[0081] 300... Neural network learning device, 310... Storage unit, 320... Arithmetic unit, 322... Learning unit, 323... Inference unit, 325... Software generation unit, 326... Functional model generation unit, NW1... Network information, NW2... Inference network information, DS... Learning data set, PM... Learned parameters, 330... Data input unit, 340... Data output unit, 350... Display unit, 360... Operation input unit, 500... Software, 200... CNN, 210... Convolution operation layer, 220... Quantization operation layer, 230... Output layer, 221... Pooling layer, 222... Batch Normalization layer, 223... Activation function layer, 224... Quantization layer, QC... Quantization convolution operation block, EB... Inference operation block, 1... Learning device, 10... Learning data acquisition unit, 11... First forward propagation unit, 12... Second forward propagation unit, 13... Calculation unit, 14... Backpropagation unit, 15... Parameter determination unit

Claims

1. A learning data acquisition unit that acquires learning data, A first forward propagation unit that obtains a loss by forward propagating the information acquired by the learning data acquisition unit through a neural network to be learned, A second forward propagation unit that is generated based on the first forward propagation unit and obtains a loss by performing calculations with a number of bits smaller than the number of calculation bits of the first forward propagation unit, A calculation unit that calculates a threshold value from the loss obtained by the first forward propagation unit and the loss obtained by the second forward propagation unit, A backpropagation unit that adds the threshold value calculated by the calculation unit and backpropagates the neural network, A parameter determination unit that determines the parameters of the neural network so that the loss decreases as a result of the backpropagation performed by the backpropagation unit, A learning device comprising the above.

2. The calculation order of the first forward propagation unit and the calculation order of the second forward propagation unit are different from each other. The learning device according to Claim 1.

3. The parameters determined by the parameter determination unit are implemented on an accelerator, The calculation order of the second forward propagation unit varies depending on the accelerator to be implemented, The learning device according to Claim 2.

4. The parameters determined by the parameter determination unit are implemented on an accelerator, The calculation formula of the second forward propagation unit varies depending on the accelerator to be implemented. The learning device according to Claim 1.

5. A calculation formula corresponding to the accelerator to be implemented is prepared in advance, The calculation formula of the second forward propagation unit is selected according to the accelerator to be implemented. The learning device according to Claim 4.

6. A learning data acquisition step of acquiring learning data, A first forward propagation step of obtaining a loss by forward propagating the information acquired in the learning data acquisition step through a neural network to be learned, A step of obtaining a loss using a second model generated based on the first model used in the first forward propagation step, wherein the number of calculation bits of the second model is smaller than the number of calculation bits of the first model (second forward propagation step), A calculation step of calculating a threshold value from the loss obtained in the first forward propagation step and the loss obtained in the second forward propagation step, A backpropagation step of adding the threshold value calculated in the calculation step and backpropagating the neural network, A parameter determination step of determining parameters of the neural network so that the loss decreases as a result of the backpropagation performed in the backpropagation step; A learning method having the above.

Citation Information

Patent Citations

  • Neural network circuit, edge device, and neural network operation method

    JP6896306B1