Training method and training device
The learning device and method address the issue of errors between neural network functional models and lightweight circuit operations by using a dual forward propagation approach with reduced arithmetic bits and backpropagation to minimize loss, resulting in reduced operational errors.
Patent Information
- Application Number
- PCT/JP2024/043261
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for converting functional models of neural networks into operations for lightweight neural network circuits often result in errors due to differences in operation accuracy and data format between the functional model and the neural network circuit.
A learning device and method that includes a learning data acquisition unit, a first forward propagation unit, a second forward propagation unit with reduced arithmetic bits, a calculation unit to determine a threshold value, a backpropagation unit, and a parameter determination unit to minimize loss and reduce errors between the functional model and the neural network circuit.
The proposed solution effectively reduces the error between the operation results of the functional model and the neural network circuit, ensuring more accurate and reliable operations.
Smart Images

Figure JP2024043261_26062025_PF_FP_ABST
Abstract
Description
Learning method and learning device
[0001] This application claims priority to Japanese Patent Application No. 2023-212953, filed on December 18, 2023, the contents of which are incorporated herein by reference.
[0002] In recent years, convolutional neural networks (CNNs) have become known as a method for recognizing patterns in images. Lightweight neural network circuits that can be incorporated into edge devices are used in embedded devices such as IoT devices (see, for example, Patent Document 1).
[0003] Patent No. 6896306
[0004] On the other hand, known libraries and platforms are used to determine the configuration and specifications of a convolutional neural network, generate a functional model of the convolutional neural network, and generate trained parameters trained using the functional model.When the functional model and trained parameters of the neural network generated in such a library or platform are converted into operations that can be performed in a lightweight neural network circuit, errors in the calculation results may occur due to differences in calculation precision and data format.
[0005] In light of these circumstances, an object of the present invention is to provide a learning device and a learning method that, when a functional model of a neural network and learned parameters learned using the functional model are converted into operations that can be performed in a neural network circuit, are less likely to produce errors between the results of operations using the functional model and the results of operations using a neural network circuit.
[0006] [1] In order to solve the above problem, one aspect of the present invention is a learning device comprising: a training data acquisition unit that acquires training data; a first forward propagation unit that obtains a loss by forward propagating information acquired by the training data acquisition unit through a neural network to be trained; a second forward propagation unit that is generated based on the first forward propagation unit and obtains the loss by performing an operation with a number of bits smaller than the number of operation bits of the first forward propagation unit; a calculation unit that calculates a threshold from the loss obtained by the first forward propagation unit and the loss obtained by the second forward propagation unit; a backpropagation unit that adds the threshold calculated by the calculation unit and backpropagates the result through the neural network; and a parameter determination unit that determines parameters of the neural network so that the loss is reduced as a result of the backpropagation performed by the backpropagation unit.
[0007] [2] In one aspect of the present invention, in the learning device described in [1] above, the operation order of the first forward propagation unit and the operation order of the second forward propagation unit are different from each other.
[0008] [3] Furthermore, one aspect of the present invention is the learning device described in [2] above, wherein the parameters determined by the parameter determination unit are implemented in an accelerator, and the order of operations of the second forward propagation unit differs depending on the accelerator to be implemented.
[0009] [4] Furthermore, one aspect of the present invention is the learning device described in [1] above, wherein the parameters determined by the parameter determination unit are implemented in an accelerator, and the arithmetic expression of the second forward propagation unit differs depending on the accelerator to be implemented.
[0010] [5] Furthermore, one aspect of the present invention is the learning device described in [4] above, in which an arithmetic expression corresponding to the accelerator to be implemented is prepared in advance, and the arithmetic expression of the second forward propagation unit is selected according to the accelerator to be implemented.
[0011] [6] Also, one aspect of the present invention is a training method including: a training data acquisition step of acquiring training data; a first forward propagation step of acquiring a loss by forward propagating information acquired in the training data acquisition step through a neural network to be trained; a second forward propagation step of acquiring a loss using a second model generated based on a first model used in the first forward propagation step, wherein the number of operation bits of the second model is smaller than the number of operation bits of the first model; a calculation step of calculating a threshold from the loss obtained in the first forward propagation step and the loss obtained in the second forward propagation step; a backpropagation step of adding the threshold calculated in the calculation step and backpropagating the result through the neural network; and a parameter determination step of determining parameters of the neural network so that the loss is reduced as a result of the backpropagation performed in the backpropagation step.
[0012] According to the present invention, it is possible to reduce the error between the calculation result based on the functional model and the calculation result based on the neural network circuit.
[0013] FIG. 1 is a diagram showing a neural network learning device according to an embodiment. FIG. 2 is a diagram for explaining input / output of a calculation unit of the neural network learning device according to the embodiment. FIG. 3 is a diagram showing an example of a layer configuration of a convolutional neural network according to the embodiment. FIG. 4 is a diagram for explaining input / output of each quantization convolution calculation block included in the convolutional neural network according to the embodiment. FIG. 5 is a diagram showing an example of the configuration of an inference calculation block in a neural network circuit according to the embodiment. FIG. 6 is a diagram showing an example of the configuration of a quantization convolution calculation block in the convolutional neural network according to the embodiment. FIG. 7 is a diagram for explaining the learning stage of the neural network learning device according to the embodiment. FIG. 8 is a diagram for explaining the relationship between a first model and a second model used in the learning stage by the neural network learning device according to the embodiment. FIG. 9 is a block diagram showing an example of the internal configuration of a learning device 1 according to the embodiment.
[0014] [Embodiments] Preferred embodiments of a learning device and a learning method according to aspects of the present invention will be described in detail below with reference to the accompanying drawings. Note that aspects of the present invention are not limited to these embodiments and include various modifications or improvements. In other words, the components described below include those that would be easily conceivable to a person skilled in the art or that are substantially identical, and the components described below can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the present invention. Furthermore, in the drawings below, the scale and number of components may differ from the scale and number of the actual structures to make each configuration easier to understand.
[0015] FIG. 1 is a diagram illustrating a neural network training device according to an embodiment. First, a neural network training device 300 will be described with reference to the diagram. The neural network training device 300 is a device that generates and trains a convolutional neural network 200 (hereinafter also referred to as "CNN 200" or "NN functional model 200"), which is a neural network functional model, and generates software 500 that operates a neural network circuit 100 (hereinafter also referred to as "NN circuit 100") that can be incorporated into embedded devices such as IoT devices. The calculations performed by the NN circuit 100 are at least a part of the inference calculations performed by the CNN 200 (NN functional model 200).
[0016] Neural network training device 300 is a programmable device (computer) equipped with a processor such as a CPU (Central Processing Unit) and hardware such as a memory. The functions of neural network training device 300 are realized by executing a neural network training program and a software generation program in neural network training device 300. Neural network training device 300 includes a storage unit 310, a calculation unit 320, a data input unit 330, a data output unit 340, a display unit 350, and an operation input unit 360.
[0017] The storage unit 310 stores network information NW1, inference network information NW2, a training data set DS, and trained parameters PM. The training data set DS and the inference network information NW2 are input data input to the neural network training device 300. The trained parameters PM are output data output by the neural network training device 300. Note that the "trained NN circuit 100" includes the NN circuit 100 and the trained parameters PM.
[0018] The network information (learning network information) NW1 is information relating to the CNN 200 (NN function model 200). The network information NW1 includes, for example, information defining the function of the CNN 200 (NN function model 200). The network information NW1 includes, for example, the network configuration of the CNN 200, input data information, output data information, quantization information, etc. The input data information includes the input data type such as image or audio, and the input data size, etc.
[0019] The inference network information NW2 is information relating to the inference operation executed by the NN circuit 100. The inference network information NW2 includes, for example, information defining the functions of the neural network inference operation that can be executed by the NN circuit 100. The inference network information NW2 includes, for example, the circuit configuration of the NN circuit 100, the functions of the arithmetic unit, the data bit width, etc.
[0020] The training data set DS includes training data D1 used for training and test data D2 used for inference testing.
[0021] The calculation unit 320 includes a learning unit 322, an inference unit 323, a software generation unit 325, and a function model generation unit 326. Network information NW is input to the calculation unit 320, and the calculation unit 320 outputs learned parameters PM. The network information NW input to the calculation unit 320 may be generated by a device other than the neural network learning device 300.
[0022] 2 is a diagram for explaining the input and output of the calculation unit of the neural network learning device according to this embodiment. Next, the configuration of the calculation unit 320 and the information input and output to and from this configuration will be described with reference to the same figure.
[0023] The learning unit 322 generates learned parameters PM using the network information NW1, inference network information NW2, and learning data D1. The inference unit 323 performs an inference test using the network information NW and test data D2.
[0024] The software generation unit 325 generates software 500 for operating the NN circuit 100 based on the network information NW1 and the inference network information NW2. The software 500 includes software for transferring the learned parameters PM to the NN circuit 100 as necessary.
[0025] The functional model generation unit 326 generates (configures) the CNN 200 (NN functional model 200) based on input from the user, and outputs network information NW1, which is information related to the CNN 200 (NN functional model 200).
[0026] Returning to FIG. 1 , the hardware information HW, network information NW, etc. required to generate the trained NN circuit 100 are input to the data input unit 330. The hardware information HW, network information NW, etc. are input as data written in a predetermined data format, for example. The input hardware information HW, network information NW, etc. are stored in the storage unit 310. The hardware information HW, network information NW, etc. may be input or changed by the user via the operation input unit 360.
[0027] The generated trained NN circuit 100 is output to the data output unit 340. For example, the generated NN circuit 100 and the trained parameters PM are output to the data output unit 340.
[0028] The display unit 350 has a known monitor such as an LCD display. The display unit 350 can display GUI (Graphical User Interface) images generated by the calculation unit 320, a console screen for receiving commands, etc. Furthermore, when the calculation unit 320 requires information input from the user, the display unit 350 can display a message prompting the user to input information from the operation input unit 360 and a GUI image required for information input.
[0029] The operation input unit 360 is a device through which a user inputs instructions to the calculation unit 320, etc. The operation input unit 360 is a known input device such as a touch panel, a keyboard, a mouse, etc. The input of the operation input unit 360 is transmitted to the calculation unit 320.
[0030] All or part of the functions of the calculation unit 320 are realized by one or more processors, such as a central processing unit (CPU) or a graphics processing unit (GPU), executing programs stored in a program memory. However, all or part of the functions of the calculation unit 320 may also be realized by hardware (e.g., a circuit unit; circuitry) such as a large-scale integration (LSI), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a programmable logic device (PLD). Furthermore, all or part of the functions of the calculation unit 320 may also be realized by a combination of software and hardware.
[0031] All or part of the functions of the calculation unit 320 may be realized using an external accelerator such as a CPU, GPU, or hardware provided in an external device such as a cloud server. The calculation speed of the calculation unit 320 can be improved by using, for example, a GPU or dedicated hardware with high calculation performance on a cloud server.
[0032] The storage unit 310 is realized by a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), a ROM (Read-Only Memory), a RAM (Random Access Memory), etc. All or part of the storage unit 310 may be provided in an external device such as a cloud server, and may be connected to the calculation unit 320, etc. via a communication line.
[0033] Note that neural network training device 300 may be configured with multiple devices (computers), and the functional blocks of calculation unit 320 may be distributed across the multiple devices. For example, neural network training device 300 may be separated into a first device (computer) having functional model generation unit 326, a second device (computer) having learning unit 322 and inference unit 323, and a third device (computer) having software generation unit 325.
[0034] FIG. 3 is a diagram showing an example of the layer configuration of a convolutional neural network according to this embodiment. An example of the layer configuration of the CNN 200 will be described with reference to the same figure. The CNN 200 is a network with a multi-layer structure including a convolution layer 210 that performs convolution operations, a quantization operation layer 220 that performs quantization operations, and an output layer 230. In at least a part of the CNN 200, the convolution layer 210 and the quantization operation layer 220 are alternately connected. The CNN 200 is a model widely used in image recognition and video recognition. The CNN 200 may further include layers having other functions, such as a fully connected layer.
[0035] 4 is a diagram for explaining the input and output of each quantization convolution operation block included in the convolutional neural network according to this embodiment. The convolution operation performed by the convolution layer 210 will be explained with reference to the diagram. The convolution layer 210 performs a convolution operation on input data a using a weight w. The convolution layer 210 performs a product-sum operation using the input data a and the weight w as input.
[0036] The input data a (also called activation data or feature map) to the convolutional layer 210 is multidimensional data such as image data. In this embodiment, the input data a is a three-dimensional tensor consisting of elements (x, y, c). The convolutional layer 210 of the CNN 200 performs a convolution operation on the low-bit input data a. In this embodiment, the elements of the input data a are 2-bit unsigned integers (0, 1, 2, 3). The elements of the input data a may be, for example, 4-bit or 8-bit unsigned integers.
[0037] If the input data input to the CNN 200 has a different format from the input data a to the convolutional layer 210, such as a 32-bit floating-point format, the CNN 200 may further include an input layer before the convolutional layer 210 that performs type conversion and quantization.
[0038] The weights w (also called filters or kernels) of the convolutional layer 210 are multidimensional data having elements that are learnable parameters. In this embodiment, the weights w are four-dimensional tensors consisting of elements (i, j, c, d). The weights w have d three-dimensional tensors (hereinafter referred to as "weights wo") consisting of elements (i, j, c). The weights w in the trained CNN 200 are trained data. The convolutional layer 210 of the CNN 200 performs convolution operations using low-bit weights w. In this embodiment, the elements of the weights w are 1-bit signed integers (0, 1), where a value "0" represents +1 and a value "1" represents -1.
[0039] The convolution layer 210 performs a predetermined convolution operation and outputs output data f to the quantization operation layer 220 .
[0040] The quantization operation layer 220 performs quantization operation and the like on the output of the convolution operation output by the convolution layer 210. The quantization operation layer 220 includes a pooling layer 221, a batch normalization layer 222, an activation function layer 223, and a quantization layer 224.
[0041] The pooling layer 221 performs operations such as average pooling and MAX pooling on the output data f of the convolution operation output by the convolution layer 210, thereby compressing the output data f of the convolution layer 210.
[0042] The batch normalization layer 222 performs a predetermined operation on the output data of the quantization operation layer 220 and the pooling layer 221 to normalize the data distribution.
[0043] The activation function layer 223 calculates an activation function such as ReLU (Rectified Linear Unit) on the outputs of the quantization operation layer 220 , the pooling layer 221 , and the batch normalization layer 222 .
[0044] The quantization layer 224 performs a quantization operation on the output of the pooling layer 221 and the activation function layer 223 based on the quantization parameter. The quantization operation is, for example, a bit reduction of the input tensor u to 2 bits.
[0045] The output layer 230 is a layer that outputs the results of the CNN 200 using an identity function, a softmax function, etc. The layer preceding the output layer 230 may be the convolution layer 210 or the quantization operation layer 220.
[0046] In the CNN 200, the quantized output data of the quantization layer 224 is input to the convolution layer 210, so the load of the convolution calculation in the convolution layer 210 is smaller than in other convolutional neural networks that do not perform quantization.
[0047] Returning to FIG. 1 , the functional model generation unit 326 sets the network structure and layer-by-layer specifications in the CNN 200 (NN functional model 200) based on user input from the operation input unit 360. For example, the user changes the network structure of the NN functional model 200 by rearranging the connections of the visually schematic layers displayed as GUI images. The user also changes the specifications (input data information, output data information, quantization information, etc.) of each layer visually schematicized and displayed as GUI images. For example, the user can rearrange the connections between the pooling layer 221, the batch normalization layer 222, the activation function layer 223, and the quantization layer 224 in the quantization operation layer 220.
[0048] The network structure and the specifications for each layer in the CNN 200 (NN function model 200) do not have to be described in a visually schematic form as shown in Fig. 4. The network structure and the specifications for each layer in the CNN 200 may be described in a programming language, XML, or the like.
[0049] The CNN 200 (NN functional model 200) generated by the functional model generation unit 326 is a neural network functional model capable of performing learning and inference operations in the operation unit 320 (learning unit 322 and inference unit 323) of the neural network training device 300. The operation unit 320 of the neural network training device 300 includes an operation circuit with higher performance than the operation circuit provided in the NN circuit 100, such as a CPU, GPU, or dedicated hardware. Therefore, the CNN 200 generated by the functional model generation unit 326 may include operation blocks that can be converted into operations that can be performed in the NN circuit 100 (hereinafter also referred to as "convertible operation blocks") and operation blocks that cannot be converted into operations that can be performed in the NN circuit 100 (hereinafter also referred to as "unconvertible operation blocks"). Here, an operation block refers to a plurality of consecutive operations in the CNN 200.
[0050] In order for the CNN 200 (NN functional model 200) to be efficiently subjected to inference operations by the NN circuit 100, it is desirable for the functional model generation unit 326 to generate as many operation blocks (convertible operation blocks) as possible that can be converted into operations that can be performed by the NN circuit 100.
[0051] 4, a part of the CNN 200 (NN functional model 200) that is an operation block from the convolution operation to the quantization operation is defined as a "quantization convolution operation block QC." At least a part of the CNN 200 is configured by connecting a plurality of quantization convolution operation blocks QC.
[0052] 5 is a diagram showing an example of the configuration of an inference operation block in a neural network circuit according to this embodiment. An example of the configuration of the inference operation block EB in the NN circuit 100 will be described with reference to the same figure. The inference operation block EB is a operation block included in hardware such as an edge device, and is an example of a calculation environment used during inference.
[0053] "C" in the figure represents a product-sum operation in a convolution operation circuit. "AW" in the figure represents data obtained by multiplying an input vector A by a weight matrix W, and is vector data with 16-bit integers per element. "Q" in the figure represents a quantization operation. "U" in the figure represents data obtained by quantizing AW, and is vector data with 2-bit integers per element.
[0054] The calculation environment in the inference calculation block EB (i.e., the calculation environment during inference) is lower in accuracy than the calculation environment in the quantization convolution calculation block QC (i.e., the calculation environment during learning), which will be described later. The calculation environment includes not only calculation accuracy but also the data format, calculation order, etc.
[0055] 6 is a diagram showing an example of the configuration of a quantization convolution operation block in a convolutional neural network according to this embodiment. Referring to the same figure, an example of the configuration of the quantization convolution operation block QC in the NN circuit 100 will be described. The quantization convolution operation block QC is an operation block used during learning, and is an example of a computation environment used during learning. The quantization convolution operation block QC can also be said to represent a configuration on an NDK (Network Development Kit).
[0056] The quantization convolution operation block QC shown in the figure is configured as a transformable operation block, which receives an input vector A and a weight matrix W and outputs an output vector U quantized to 2 bits per element. "X1" in the figure indicates an operation (postscaler after quantization) in which an affine transformation operation using first scaling coefficients (first scaling factors) Sa1 and Sb1 in floating-point format is performed on the input vector A (Sa1 × A + Sb1) and vector data As in floating-point format is output. When the quantization convolution operation block QC is configured as a transformable operation block, the input data is limited to the input vector A having 2 bits per element. Even in this case, by performing an affine transformation on the input vector A using the first scaling coefficients Sa1 and Sb1 as coefficients, it is possible to suppress a decrease in the accuracy of the input data.
[0057] Also, "X2" in the figure indicates an operation in which an affine transformation operation using second scaling coefficients (second scaling factors) Sa2 and Sb2 in floating-point format as coefficients is performed on weight matrix W (Sa2 × W + Sb2), and matrix data Ws in floating-point format is output. When the quantization convolution operation block QC is configured as a transformable operation block, the weights are limited to a weight matrix W with one bit per element. Even in this case, by performing an affine transformation on weight matrix W using scaling coefficients Sa2 and Sb2 as coefficients, it is possible to suppress a decrease in the accuracy of the weights.
[0058] Also, "Cf" in the figure indicates a convolution operation that multiplies As and Ws to output vector data AWs in floating-point format. Also, "X3" in the figure indicates an operation (prescaler before quantization) that performs an affine transformation operation on vector data AWs using third scaling coefficients (third scaling factors) Sa3 and Sb3 in floating-point format as coefficients (Sa3 x AWs + Sb3) to output vector data AWss in floating-point format. For example, "X3" is a prescaler corresponding to "X1" (postscaler after quantization).
[0059] Also, "Qf" shown in the figure indicates a quantization operation that quantizes vector data AWss in floating-point format based on a quantization parameter qf (thf0, thf1, thf2) to output vector data U of 2-bit integers per element. The quantization parameter qf is a threshold value (thf0, thf1, thf2) in the floating-point format. When the quantization convolution operation block QC is configured as a convertible operation block, the output data is limited to vector data U of 2 bits per element.
[0060] The quantization convolution operation block QC can be treated as a convertible operation block that can be converted into an operation that can be performed by the inference operation block EB by incorporating and aggregating the scaling coefficients (Sa1, Sb1, Sa2, Sb2, Sa3, Sb3) into the quantization parameters qf (thf0, thf1, thf2) in the quantization operation Qf. For example, if Sa1 is 1.5, Sa2 is 2.0, and Sa3 is 1.1, the scaling coefficients are aggregated into the quantization parameters qf by updating the quantization parameters qf (thf0, thf1, thf2) in the quantization operation Q to values 1 / 3.3 times the original quantization parameters.
[0061] Even when another type of operation P is added, the quantization convolution operation block QC can be configured as a convertible operation block depending on the type of operation P. For example, as described above, batch normalization and activation function operations can be incorporated into the quantization parameter qf (thf0, thf1, thf2) and consolidated. Furthermore, the addition of a bias value to the convolution operation result can be incorporated into the quantization parameter qf and consolidated by subtracting the bias value from the quantization parameter qf. Therefore, even when another type of operation P, such as batch normalization, activation function, or bias value addition, is added, the quantization convolution operation block QC can be configured as a convertible operation block. If the operation P cannot be incorporated into the quantization parameter qf and consolidated, the operation unit block including the operation P becomes a non-convertible operation block.
[0062] When the operation P includes multiple floating-point operations, it is desirable to perform the multiple floating-point operations in an order that minimizes the occurrence of rounding errors. This is because if rounding errors are likely to occur, errors are more likely to occur between the operation results of the quantization convolution operation block QC and the operation results of the inference operation block EB due to variations in rounding errors, which will be described later.
[0063] FIG. 7 is a diagram illustrating the learning stage of the neural network learning device according to this embodiment. An example of the learning process performed by the learning device 1 will be described with reference to the diagram. The computational environment in the learning stage is different from the computational environment in the inference stage, and it is known that the computational environment in the learning stage typically has higher computational accuracy. Conventionally, even if learning is performed in a computational environment with higher accuracy, there has been a problem in that the inference accuracy decreases when the inference environment is incorporated into an edge device or the like with low computational accuracy. In this embodiment, inference accuracy can be improved by performing learning in the learning stage while taking into account the computational environment in the inference stage. The learning device 1 includes a learning data acquisition unit 10, a first forward propagation unit 11, a second forward propagation unit 12, a calculation unit 13, and a backpropagation unit 14.
[0064] The learning data acquisition unit 10 acquires learning data to be used for learning. When the purpose is object detection, the learning data may be, for example, information in which image data is associated with the positions, classes, etc. of objects shown in the image data.
[0065] The first forward propagation unit 11 forward propagates the information acquired by the training data acquisition unit 10 through the neural network to be trained, thereby obtaining an output of the neural network from the final layer.
[0066] The second forward propagation unit 12 forward propagates the information acquired by the training data acquisition unit 10 through the neural network to be trained, thereby obtaining an output of the neural network from the final layer.
[0067] Here, the first forward propagation unit 11 and the second forward propagation unit 12 share the parameter values updated by the parameter determination unit 15. When the parameter determination unit 15 updates the parameters, both the parameters referenced by the first forward propagation unit 11 and the second forward propagation unit 12 are updated. The parameters referenced by the second forward propagation unit 12 are the parameters referenced by the first forward propagation unit 11, with their precision reduced to the number of operation bits of the second forward propagation unit 12. In other words, the first forward propagation unit 11 and the second forward propagation unit 12 have different numbers of operation bits. Specifically, the number of operation bits of the first forward propagation unit 11 is greater than the number of operation bits of the second forward propagation unit 12. More specifically, the first forward propagation unit 11 may be fixed (fixed point), and the second forward propagation unit 12 may be float (floating point). By having the first forward propagation unit 11 and the second forward propagation unit 12 share parameters, it is possible to create two forward propagation units that are essentially the same, although they have different numbers of operation bits.
[0068] The second forward propagation unit 12 is intended to be an edge device in which a neural network is implemented. The second forward propagation unit 12 may be generated based on the first forward propagation unit 11. The second forward propagation unit 12 may have its calculation order and calculation formulas determined with an edge device in which a neural network is implemented in mind. In other words, the first forward propagation unit 11 and the second forward propagation unit 12 may have different calculation orders and calculation formulas. In the following description, the model used by the first forward propagation unit 11 may be referred to as a first model, and the model used by the second forward propagation unit 12 may be referred to as a second model.
[0069] The calculation unit 13 calculates a predetermined threshold from the output obtained by the first forward propagation unit 11 and the output obtained by the second forward propagation unit 12. As an example of the threshold, the calculation unit 13 calculates the difference between the output obtained by the first forward propagation unit 11 and the output obtained by the second forward propagation unit 12. This difference may also be referred to as a "Threshold." This difference may also be referred to as a "constant term." Note that this difference is calculated in order to apply this embodiment to frameworks such as PyTorch and TensorFlow. Therefore, in a configuration including the second forward propagation unit 12 as shown in FIG. 7 , the difference does not necessarily need to be calculated. In this case, the backpropagation unit 14 may perform calculations based on the loss calculated by the second forward propagation unit 12.
[0070] The backpropagation unit 14 adds the difference (Threshold) calculated by the calculation unit 13 to the first forward propagation unit 11. As a result, the value of the first forward propagation unit 11 becomes equal to the value of the second forward propagation unit 12. This means that the inference result is equal to the value of the second forward propagation unit 12. In frameworks such as PyTorch and TensorFlow, this addition allows parameter updates to be performed based on the first forward propagation unit 11. By adopting such a configuration, learning in which the inference result is based on the second forward propagation unit 12 and parameter updates are based on the first forward propagation unit 11 can also be achieved in frameworks such as PyTorch and TensorFlow. After the difference calculated by the calculation unit 13 is added, the neural network is backpropagated. Note that this neural network is defined by the first forward propagation unit 11. The backpropagation unit 14 may be Float (floating point). Adding the difference (threshold) may be, for example, adding the coordinate values obtained as the difference as they are (shifting the coordinates). In addition, in the case of a model that performs object detection, adding the difference (threshold) may be a calculation that changes the likelihood value (weights the likelihood).
[0071] The parameter determination unit 15 determines the parameters of the neural network so that the loss is reduced as a result of the backpropagation performed by the backpropagation unit 14. Based on the parameters determined by the parameter determination unit 15, the parameters of both the first forward propagation unit 11 and the second forward propagation unit 12 are updated.
[0072] If there are multiple accelerators that may be implemented, an arithmetic expression corresponding to each accelerator may be prepared in advance. In this case, the arithmetic expression of the second forward propagation unit 12 may be selected depending on the accelerator to be implemented.
[0073] 8 is a diagram illustrating the relationship between a first model and a second model used in the learning stage by the neural network learning device according to this embodiment. The relationship between the first model and the second model will be described with reference to the diagram. The first model is a model for learning in a resource-rich computing environment. The second model is a model designed for inference on an accelerator mounted on an edge device. While the second model is actually executed in a resource-rich computing environment similar to the first model, the number of bits in the second model is intentionally reduced to use fixed (fixed point) data, since it is designed for inference on an accelerator mounted on an edge device.
[0074] Here, conventionally, when learning is performed in a computational environment with abundant resources, forward propagation is performed using a Float (floating-point) model, and backpropagation is similarly performed using a Float (floating-point) model. Learning is performed by repeating this process. On the other hand, in this embodiment, forward propagation is performed using a Float (floating-point) model and a Fixed (fixed-point) model, and a difference (threshold or constant term) is added to the computation result of the Float (floating-point) model, and backpropagation is performed using the Float (floating-point) model. By repeating this process, learning is performed assuming inference on an accelerator installed in an edge device.
[0075] The sum of the difference (threshold or constant term) may be the difference between the outputs of the last layer (in the illustrated example, the quantization convolution operation block QC1n and the quantization convolution operation block QC2n). However, the present embodiment is not limited to this example, and a difference may be calculated for each quantization convolution operation block QC, and the differences may be summed.
[0076] FIG. 9 is a block diagram showing an example of the internal configuration of the learning device 1 according to this embodiment. At least some of the functions of the learning device 1 can be implemented using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, and a bus 906. The computer itself can be implemented using existing technology. The central processing unit 901 executes instructions contained in a program read from the RAM 902 or the like. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. RAM stands for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices. The input / output devices 904 and 905 are also input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906. Furthermore, all or part of each functional unit provided in the learning device 1 may be realized using hardware (e.g., circuitry) such as an ASIC, PLD, or FPGA. Furthermore, all or part of each functional unit may be realized by a combination of software and hardware.
[0077] Summary of the Embodiment According to the embodiment described above, the learning device 1 includes a training data acquisition unit 10, which acquires training data. The learning device 1 also includes a first forward propagation unit 11, which forward propagates information acquired by the training data acquisition unit 10 through a neural network to be trained, thereby obtaining a loss. The learning device 1 also includes a second forward propagation unit 12 generated based on the first forward propagation unit 11, which forward propagates information acquired by the training data acquisition unit 10 through a neural network to be trained, thereby obtaining a loss. The number of calculation bits of the second forward propagation unit 12 is smaller than the number of calculation bits of the first forward propagation unit 11. The learning device 1 also includes a calculation unit 13, which calculates the difference between the loss obtained by the first forward propagation unit 11 and the loss obtained by the second forward propagation unit 12, and a backpropagation unit 14, which adds the difference calculated by the calculation unit 13 and backpropagates the neural network. Furthermore, the learning device 1 includes a parameter determination unit 15 that determines the parameters of the neural network so as to reduce losses as a result of backpropagation performed by the backpropagation unit 14. By employing such a configuration, the learning device 1 can reduce errors between the calculation results obtained by the functional model and the calculation results obtained by the neural network circuit.
[0078] Note that all or part of the functions of each unit of the learning device according to the above-described embodiment may be realized by recording a program for realizing these functions on a computer-readable recording medium, and then loading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.
[0079] Furthermore, "computer-readable recording media" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage units such as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over networks like the Internet or communication lines like telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within computer systems that serve as servers or clients in such cases. Furthermore, the above-mentioned programs may be programs that realize some of the aforementioned functions, or may be programs that can realize the aforementioned functions in combination with programs already stored in the computer system.
[0080] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the spirit of the present invention.
[0081] According to the present invention, it is possible to reduce the error between the calculation result based on the functional model and the calculation result based on the neural network circuit.
[0082] 300...Neural network learning device, 310...Memory unit, 320...Calculation unit, 322...Learning unit, 323...Inference unit, 325...Software generation unit, 326...Function model generation unit, NW1...Network information, NW2...Inference network information, DS...Learning data set, PM...Learned parameters, 330...Data input unit, 340...Data output unit, 350...Display unit, 360...Operation input unit, 500...Software, 200...CNN, 210...Convolution operation layer, 220...Quantization operation layer, 230...Output layer, 221...Pooling layer, 222...Batch Normalization layer, 223... activation function layer, 224... quantization layer, QC... quantization convolution operation block, EB... inference operation block, 1... learning device, 10... learning data acquisition unit, 11... first forward propagation unit, 12... second forward propagation unit, 13... calculation unit, 14... back propagation unit, 15... parameter determination unit
Claims
1. A learning device comprising: a learning data acquisition unit that acquires learning data; a first forward propagation unit that obtains a loss by forward propagating information acquired by the learning data acquisition unit through a neural network to be trained; a second forward propagation unit that is generated based on the first forward propagation unit and obtains the loss by performing an operation with a number of bits smaller than the number of operation bits of the first forward propagation unit; a calculation unit that calculates a threshold from the loss obtained by the first forward propagation unit and the loss obtained by the second forward propagation unit; a backpropagation unit that adds the threshold calculated by the calculation unit and backpropagates through the neural network; and a parameter determination unit that determines parameters of the neural network so that the loss is reduced as a result of the backpropagation performed by the backpropagation unit.
2. The learning device according to claim 1, wherein the operation order of the first forward propagation unit and the operation order of the second forward propagation unit are different from each other.
3. The learning device according to claim 2, wherein the parameters determined by the parameter determination unit are implemented in an accelerator, and an order of operations in the second forward propagation unit differs depending on the accelerator in which it is implemented.
4. The learning device according to claim 1, wherein the parameters determined by the parameter determination unit are implemented in an accelerator, and the calculation formula of the second forward propagation unit differs depending on the accelerator to be implemented.
5. The learning device according to claim 4, wherein an arithmetic expression corresponding to an accelerator to be implemented is prepared in advance, and the arithmetic expression of the second forward propagation unit is selected according to the accelerator to be implemented.
6. A learning method comprising: a learning data acquisition step of acquiring learning data; a first forward propagation step of acquiring a loss by forward propagating information acquired in the learning data acquisition step through a neural network to be trained; a second forward propagation step of acquiring a loss using a second model generated based on a first model used in the first forward propagation step, wherein the number of operation bits of the second model is smaller than the number of operation bits of the first model; a calculation step of calculating a threshold from the loss obtained in the first forward propagation step and the loss obtained in the second forward propagation step; a backpropagation step of adding the threshold calculated in the calculation step and backpropagating the neural network; and a parameter determination step of determining parameters of the neural network so that the loss is reduced as a result of the backpropagation performed in the backpropagation step.
Citation Information
Patent Citations
Information processing device, information processing method, and information processing program
WO2022009449A1
Neural network generation device, neural network computing device, edge device, neural network control method, and software generation program
WO2022230906A1