Neural network generation device

The neural network generation device and method address the challenge of operating convolutional neural networks in embedded devices with limited resources by generating a tailored neural network execution model, achieving high performance and efficient resource utilization.

JP2025072666APending Publication Date: 2025-05-09MAXELL LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025025279
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-30
Filing Date
2025-02-19
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

There is a demand for a neural network generation method that can efficiently operate convolutional neural networks in embedded devices like IoT devices, with limited hardware resources, while ensuring high performance.

Method used

A neural network generation device, method, and program that generate a neural network execution model tailored to the hardware configuration of embedded devices, incorporating an execution model generation unit that uses hardware and network information to create the model, and a learning unit that generates trained parameters for the model.

Benefits of technology

The proposed solution enables the generation of neural networks that can be operated with high performance in embedded devices, efficiently utilizing limited hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072666000001_ABST
    Figure 2025072666000001_ABST
Patent Text Reader

Abstract

To provide a neural network generation device, a neural network generation method, and a neural network generation program for generating a neural network that can be incorporated into an embedded device such as an IoT device and can be operated with high performance.SOLUTION: The neural network generation device is provided with a storage unit, an arithmetic unit, a data input unit, a data output unit, a display unit, and an operation input unit. The arithmetic unit is provided with an execution model generation unit for generating a neural network (NN) execution model based on hardware information of hardware on which a neural network execution model operates and network information of the neural network, and a learning unit for generating learned parameters of the generated NN execution model. The learning unit performs learning using the number of bits with higher accuracy than the neural network execution model to determine parameters in the NN execution model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a neural network generation device, a neural network generation method, and a neural network generation program. This application claims priority based on Japanese Patent Application No. 2020-113315 filed in Japan on June 30, 2020, the contents of which are incorporated herein by reference. [Background technology]

[0002] In recent years, convolutional neural networks (CNNs) have been used as models for image recognition and the like. Convolutional neural networks have a multi-layer structure with convolutional layers and pooling layers, and require a large number of calculations, such as convolutional calculations. Various calculation methods have been devised to speed up calculations by convolutional neural networks (Patent Document 1, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2018-077829 A Summary of the Invention [Problem to be solved by the invention]

[0004] On the other hand, image recognition using convolutional neural networks is also used in embedded devices such as IoT devices. In order to operate convolutional neural networks efficiently in embedded devices, a generation method for generating neural networks (models and circuits) that match the hardware configuration of the embedded device is desired. In the process of generating a neural network, a neural network training method that allows the neural network to operate with high performance within the limited hardware resources of the embedded device is also desired.

[0005] In consideration of the above circumstances, the present invention aims to provide a neural network generation device, a neural network generation method, and a neural network generation program that generate a neural network that can be embedded in embedded devices such as IoT devices and operate with high performance. [Means for solving the problem]

[0006] In order to solve the above problems, the present invention proposes the following means. A neural network generation device according to a first aspect of the present invention is a neural network generation device that generates a neural network execution model that calculates a neural network, and includes an execution model generation unit that generates the neural network execution model based on hardware information of hardware on which the neural network execution model operates and network information of the neural network, and a learning unit that generates learned parameters of the generated neural network execution model.

[0007] A neural network generation method according to a second aspect of the present invention is a neural network generation method for generating a neural network execution model that calculates a neural network, and includes a hardware information acquisition step of acquiring hardware information of hardware on which the neural network execution model operates, a network information acquisition step of setting network information of the neural network, an execution model generation step of generating the neural network execution model based on the hardware information and the network information, and a learning step of learning learning parameters of the generated neural network execution model.

[0008] A neural network generation program according to a third aspect of the present invention is a neural network generation program that causes a computer to generate a neural network execution model that calculates a neural network, and includes a hardware information acquisition step of causing the computer to acquire hardware information of hardware on which the neural network execution model operates, a network information acquisition step of causing the computer to set network information of the neural network, an execution model generation step of causing the computer to generate the neural network execution model based on the hardware information and the network information, and a learning step of causing the computer to learn learning parameters of the generated neural network execution model. Effect of the Invention

[0009] The neural network generating device, the neural network generating method, and the neural network generating program of the present invention can generate a neural network that can be incorporated into embedded devices such as IoT devices and can operate with high performance. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating a neural network generation device according to a first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating input and output of a calculation unit of the neural network generation device. [Diagram 3] FIG. 1 illustrates an example of a convolutional neural network. [Figure 4] FIG. 2 is a diagram for explaining a convolution operation performed by a convolution layer of the convolutional neural network. [Diagram 5] FIG. 1 illustrates an example of a neural network execution model. [Figure 6] 3 is a control flowchart of the neural network generation device. [Figure 7] 4 is a timing chart showing an example of the operation of the neural network execution model. [Figure 8] 11A and 11B are diagrams for explaining data division and data expansion of the convolution operation. [Figure 9] 10 is a timing chart showing another operation example of the neural network execution model. [Figure 10] FIG. 13 is a diagram showing partial tensors obtained by dividing output data of a convolution operation into tiles. [Figure 11] FIG. 13 is a diagram showing partial tensors obtained by dividing input data into slices. [Figure 12] FIG. 13 is a diagram showing partial tensors obtained by dividing input data into slices. [Figure 13] FIG. 13 is a diagram showing partial tensors obtained by dividing input data into slices. [Figure 14] FIG. 13 is a diagram showing other partial tensors required to output a partial tensor by a convolution operation of layer 2M+1. [Figure 15] FIG. 11 is an internal block diagram of a generated convolution operation circuit. [Figure 16] FIG. 2 is an internal block diagram of a multiplier in the convolution operation circuit. [Figure 17] FIG. 2 is an internal block diagram of a multiply-and-accumulate unit of the multiplier. [Figure 18] FIG. 2 is an internal block diagram of an accumulator circuit of the convolution operation circuit. [Figure 19] FIG. 2 is an internal block diagram of an accumulator unit of the accumulator circuit. [Figure 20] FIG. 4 is a state transition diagram of a control circuit of the convolution operation circuit. [Figure 21] FIG. 11 is an internal block diagram of a generated quantization calculation circuit. [Figure 22] FIG. 2 is an internal block diagram of a vector operation circuit and a quantization circuit of the quantization operation circuit. [Figure 23] FIG. 2 is a block diagram of an arithmetic unit of the vector arithmetic circuit. [Figure 24] FIG. 2 is an internal block diagram of a quantization unit of the quantization circuit. [Diagram 25] FIG. 1 is an internal block diagram of a generated DMAC. [Figure 26] FIG. 11 is a diagram illustrating a scaling coefficient in a quantization operation. [Figure 27] FIG. 11 is a diagram illustrating a scaling coefficient in a quantization operation. [Figure 28] FIG. 11 is a diagram illustrating a scaling coefficient in a quantization operation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] First embodiment A first embodiment of the present invention will be described with reference to FIGS. FIG. 1 is a diagram showing a neural network generation device 300 according to this embodiment.

[0012] [Neural network generating device 300] The neural network generation device 300 is a device that generates a trained neural network execution model 100 that can be embedded in an embedded device such as an IoT device. The neural network execution model 100 is a software or hardware model generated to operate a convolutional neural network 200 (hereinafter referred to as "CNN 200") in the embedded device.

[0013] The neural network generating device 300 is a device (computer) capable of executing a program, which includes a processor such as a CPU (Central Processing Unit) and hardware such as a memory. The functions of the neural network generating device 300 are realized by executing a neural network generating program in the neural network generating device 300. The neural network generating device 300 includes a storage unit 310, a calculation unit 320, a data input unit 330, a data output unit 340, a display unit 350, and an operation input unit 360.

[0014] The storage unit 310 stores hardware information HW, network information NW, a learning data set DS, a neural network execution model 100 (hereinafter referred to as "NN execution model 100"), and learned parameters PM. The hardware information HW, the learning data set DS, and the network information NW are input data input to the neural network generation device 300. The NN execution model 100 and the learned parameters PM are output data output by the neural network generation device 300. It should be noted that the "trained NN execution model 100" includes the NN execution model 100 and the learned parameters PM.

[0015] The hardware information HW is information on an embedded device (hereinafter, referred to as "target hardware") that runs the NN execution model 100. The hardware information HW includes, for example, the device type, device constraints, memory configuration, bus configuration, operating frequency, power consumption, and manufacturing process type of the target hardware. The device type includes, for example, ASIC (Application Specific Integrated Circuit) and FPGA (Field-Programmable Gate Array). The device constraint includes the upper limit of the number of arithmetic units included in the target device and the upper limit of the circuit size. The memory configuration includes the memory type, number of memories, memory capacity, and input / output data width. The bus configuration includes the bus type, bus width, bus communication standard, and devices connected on the same bus. Furthermore, when there are multiple variations of the NN execution model 100, the hardware information HW includes information on the variations of the NN execution model 100 to be used.

[0016] The network information NW is basic information of the CNN 200. The network information NW is, for example, the network configuration, input data information, output data information, quantization information, etc. of the CNN 200. The input data information is the type of input data such as image or sound, and the input data size, etc.

[0017] The training data set DS includes training data D1 used for training and test data D2 used for inference testing.

[0018] FIG. 2 is a diagram showing inputs and outputs of the calculation unit 320. As shown in FIG. The calculation unit 320 includes an execution model generation unit 321, a learning unit 322, an inference unit 323, and a hardware generation unit 324. The NN execution model 100 input to the calculation unit 320 may be generated by a device other than the neural network generation device 300.

[0019] The execution model generation unit 321 generates the NN execution model 100 based on the hardware information HW and the network information NW.

[0020] The learning unit 322 generates learned parameters PM using the NN running model 100 and the learning data D1. The inference unit 323 performs an inference test using the NN running model 100 and the test data D2.

[0021] The hardware generation unit 324 generates the neural network hardware model 400 based on the hardware information HW and the NN execution model 100. The neural network hardware model 400 is a hardware model that can be implemented in the target hardware. The neural network hardware model 400 is optimized for the target hardware based on the hardware information HW. The neural network hardware model 400 may be an RTL (Register Transfer Level), a netlist representing connections between gates and circuit modules, or a combination thereof. The neural network hardware model 400 may be a parameter list or a configuration file required to implement the NN execution model 100 in hardware. The parameter list or the configuration file is used in combination with the NN execution model 100 generated separately.

[0022] The data input unit 330 receives input of hardware information HW, network information NW, etc., required for generating the trained NN execution model 100. The hardware information HW, network information NW, etc. are input as data described in a predetermined data format, for example. The input hardware information HW, network information NW, etc. are stored in the storage unit 310. The hardware information HW, network information NW, etc. may be input or changed by the user via the operation input unit 360.

[0023] The generated trained NN execution model 100 is output to the data output unit 340. For example, the generated NN execution model 100 and the trained parameters PM are output to the data output unit 340.

[0024] The display unit 350 has a known monitor such as an LCD display. The display unit 350 can display GUI (Graphical User Interface) images generated by the calculation unit 320, a console screen for receiving commands, etc. Furthermore, when the calculation unit 320 requires information input from the user, the display unit 350 can display a message prompting the user to input information from the operation input unit 360 and a GUI image required for information input.

[0025] The operation input unit 360 is a device through which a user inputs instructions to the calculation unit 320, etc. The operation input unit 360 is a known input device such as a touch panel, a keyboard, a mouse, etc. The input of the operation input unit 360 is transmitted to the calculation unit 320.

[0026] All or part of the functions of the calculation unit 320 are realized by one or more processors, such as a central processing unit (CPU) or a graphics processing unit (GPU), executing a program stored in a program memory. However, all or part of the functions of the calculation unit 320 may be realized by hardware (e.g., a circuit unit; circuitry) such as a large scale integration (LSI), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a programmable logic device (PLD). All or part of the functions of the calculation unit 320 may be realized by a combination of software and hardware.

[0027] All or part of the functions of the calculation unit 320 may be realized using an external accelerator such as a CPU, GPU, or hardware provided in an external device such as a cloud server. The calculation speed of the calculation unit 320 can be improved by using a GPU or dedicated hardware with high calculation performance on a cloud server, for example.

[0028] The storage unit 310 is realized by a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), a ROM (Read-Only Memory), a RAM (Random Access Memory), etc. All or a part of the storage unit 310 may be provided in an external device such as a cloud server, and may be connected to the calculation unit 320, etc., via a communication line.

[0029] [Convolutional Neural Network (CNN) 200] Next, the CNN 200 will be described. Fig. 3 is a diagram showing an example of the CNN 200. The network information NW of the CNN 200 is information about the configuration of the CNN 200 described below. The CNN 200 uses low-bit weights w and quantized input data a, and is easy to incorporate into embedded devices.

[0030] The CNN 200 is a network with a multi-layer structure including a convolution layer 210 that performs a convolution operation, a quantization operation layer 220 that performs a quantization operation, and an output layer 230. In at least a part of the CNN 200, the convolution layer 210 and the quantization operation layer 220 are alternately connected. The CNN 200 is a model that is widely used for image recognition and video recognition. The CNN 200 may further include a layer having other functions, such as a fully connected layer.

[0031] FIG. 4 is a diagram illustrating the convolution operation performed by the convolution layer 210. The convolution layer 210 performs a convolution operation on the input data a using a weight w. The convolution layer 210 performs a multiply-and-accumulate operation on the input data a and the weight w.

[0032] The input data a (also called activation data or feature map) to the convolutional layer 210 is multidimensional data such as image data. In this embodiment, the input data a is a three-dimensional tensor consisting of elements (x, y, c). The convolutional layer 210 of the CNN 200 performs a convolution operation on the low-bit input data a. In this embodiment, the elements of the input data a are 2-bit unsigned integers (0, 1, 2, 3). The elements of the input data a may be, for example, 4-bit or 8-bit unsigned integers.

[0033] If the input data input to the CNN 200 has a different format from the input data a to the convolutional layer 210, such as a 32-bit floating-point type, the CNN 200 may further have an input layer before the convolutional layer 210 that performs type conversion and quantization.

[0034] The weight w (also called a filter or kernel) of the convolutional layer 210 is multidimensional data having elements that are learnable parameters. In this embodiment, the weight w is a four-dimensional tensor consisting of elements (i, j, c, d). The weight w has d three-dimensional tensors (hereinafter referred to as "weights wo") consisting of elements (i, j, c). The weight w in the trained CNN 200 is trained data. The convolutional layer 210 of the CNN 200 performs a convolution operation using a low-bit weight w. In this embodiment, the element of the weight w is a 1-bit signed integer (0, 1), where a value "0" represents +1 and a value "1" represents -1.

[0035] The convolution layer 210 performs the convolution operation shown in Equation 1 and outputs output data f. In Equation 1, s indicates the stride. The area indicated by the dotted line in Fig. 4 indicates one of the areas ao (hereinafter referred to as "application area ao") where the weight wo is applied to the input data a. The elements of the application area ao are represented by (x+i, y+j, c).

[0036]

number

[0037] The quantization operation layer 220 performs quantization and the like on the output of the convolution operation output by the convolution layer 210. The quantization operation layer 220 includes a pooling layer 221, a batch normalization layer 222, an activation function layer 223, and a quantization layer 224.

[0038] The pooling layer 221 compresses the output data f of the convolutional operation output by the convolutional layer 210 by performing calculations such as average pooling (Equation 2) and MAX pooling (Equation 3) on the output data f of the convolutional operation output by the convolutional layer 210. In Equations 2 and 3, u indicates an input tensor, v indicates an output tensor, and T indicates the size of the pooling region. In Equation 3, max is a function that outputs the maximum value of u for the combination of i and j included in T.

[0039]

number

[0040]

number

[0041] The batch normalization layer 222 normalizes the data distribution of the output data of the quantization operation layer 220 and the pooling layer 221, for example, by the operation shown in Equation 4. In Equation 4, u represents an input tensor, v represents an output tensor, α represents a scale, and β represents a bias. In the trained CNN 200, α and β are trained constant vectors.

[0042]

number

[0043] The activation function layer 223 performs an activation function operation such as ReLU (Equation 5) on the output of the quantization operation layer 220, the pooling layer 221, and the batch normalization layer 222. In Equation 5, u is an input tensor and v is an output tensor. In Equation 5, max is a function that outputs the largest numerical value among arguments.

[0044]

number

[0045] The quantization layer 224 performs quantization on the output of the pooling layer 221 and the activation function layer 223 based on the quantization parameter, for example, as shown in Equation 6. The quantization shown in Equation 6 reduces the input tensor u to 2 bits. In Equation 6, q(c) is a vector of quantization parameters. In the trained CNN 200, q(c) is a trained constant vector. The inequality sign "≦" in Equation 6 may be "<".

[0046]

number

[0047] The output layer 230 is a layer that outputs the result of the CNN 200 using an identity function, a softmax function, etc. The layer preceding the output layer 230 may be the convolution layer 210 or the quantization operation layer 220.

[0048] In the CNN 200, the quantized output data of the quantization layer 224 is input to the convolution layer 210, so the load of the convolution calculation in the convolution layer 210 is smaller than that of other convolution neural networks that do not perform quantization.

[0049] [Neural Network Execution Model 100 (NN Execution Model) 100] Next, the NN execution model 100 will be described. Fig. 5 is a diagram showing an example of the NN execution model 100. The NN execution model 100 is software and a hardware model generated to cause the CNN 200 to perform calculations on target hardware. The software includes software that controls the hardware model. The hardware model may be at a behavior level, at an RTL (Register Transfer Level), or a netlist representing connections between gates and circuit modules, or a combination thereof.

[0050] The NN execution model 100 includes a first memory 1, a second memory 2, a DMA controller 3 (hereinafter also referred to as "DMAC 3"), a convolution operation circuit 4, a quantization operation circuit 5, and a controller 6. The NN execution model 100 is characterized in that the convolution operation circuit 4 and the quantization operation circuit 5 are formed in a loop shape via the first memory 1 and the second memory 2.

[0051] The first memory 1 is a rewritable memory such as a volatile memory constituted by, for example, SRAM (Static RAM). Data is written to and read from the first memory 1 via the DMAC 3 and the controller 6. The first memory 1 is connected to an input port of the convolution operation circuit 4, and the convolution operation circuit 4 can read data from the first memory 1. The first memory 1 is also connected to an output port of the quantization operation circuit 5, and the quantization operation circuit 5 can write data to the first memory 1. The external host CPU can input and output data to and from the NN execution model 100 by writing and reading data to and from the first memory 1.

[0052] The second memory 2 is a rewritable memory such as a volatile memory constituted by, for example, an SRAM (Static RAM). Data is written to and read from the second memory 2 via the DMAC 3 and the controller 6. The second memory 2 is connected to an input port of the quantization calculation circuit 5, and the quantization calculation circuit 5 can read data from the second memory 2. The second memory 2 is also connected to an output port of the convolution calculation circuit 4, and the convolution calculation circuit 4 can write data to the second memory 2. The external host CPU can input and output data to and from the NN execution model 100 by writing and reading data to and from the second memory 2.

[0053] The DMAC 3 is connected to the external bus EB, and performs data transfer between an external memory such as a DRAM and the first memory 1. The DMAC 3 also performs data transfer between an external memory such as a DRAM and the second memory 2. The DMAC 3 also performs data transfer between an external memory such as a DRAM and the convolution operation circuit 4. The DMAC 3 also performs data transfer between an external memory such as a DRAM and the quantization operation circuit 5.

[0054] The convolution operation circuit 4 is a circuit that performs a convolution operation in the convolution layer 210 of the trained CNN 200. The convolution operation circuit 4 reads the input data a stored in the first memory 1, and performs a convolution operation on the input data a. The convolution operation circuit 4 writes output data f of the convolution operation (hereinafter, also referred to as “convolution operation output data”) to the second memory 2.

[0055] The quantization operation circuit 5 is a circuit that performs at least a part of the quantization operation in the quantization operation layer 220 of the trained CNN 200. The quantization operation circuit 5 reads the output data f of the convolution operation stored in the second memory 2, and performs a quantization operation (an operation including at least quantization among pooling, batch normalization, activation function, and quantization) on the output data f of the convolution operation. The quantization operation circuit 5 writes the output data of the quantization operation (hereinafter, also referred to as "quantization operation output data") to the first memory 1.

[0056] The controller 6 is connected to an external bus EB and operates as a slave to an external host CPU. The controller 6 has registers 61 including a parameter register and a status register. The parameter register is a register that controls the operation of the NN execution model 100. The status register is a register that indicates the status of the NN execution model 100, including a semaphore S. The external host CPU can access the registers 61 via the controller 6.

[0057] The controller 6 is connected to the first memory 1, the second memory 2, the DMAC 3, the convolution operation circuit 4, and the quantization operation circuit 5 via an internal bus IB. An external host CPU can access each block via the controller 6. For example, the external host CPU can issue commands to the DMAC 3, the convolution operation circuit 4, and the quantization operation circuit 5 via the controller 6. In addition, the DMAC 3, the convolution operation circuit 4, and the quantization operation circuit 5 can update a status register (including a semaphore S) held by the controller 6 via the internal bus IB. The status register (including a semaphore S) may be configured to be updated via a dedicated line connected to the DMAC 3, the convolution operation circuit 4, and the quantization operation circuit 5.

[0058] Since the NN execution model 100 has a first memory 1, a second memory 2, etc., it is possible to reduce the number of times that duplicate data is transferred when data is transferred from an external memory such as a DRAM by a DMAC 3. This makes it possible to significantly reduce the power consumption caused by memory access.

[0059] [Operation of the neural network generating device 300] Next, the operation of neural network generation device 300 (neural network generating method) will be described with reference to the control flowchart of neural network generation device 300 shown in Fig. 6. After performing an initialization process (step S10), neural network generation device 300 executes step S11.

[0060] <Hardware information acquisition step (S11)> In step S11, the neural network generating device 300 acquires hardware information HW of the target hardware to be operated (hardware information acquisition step). The neural network generating device 300 acquires the hardware information HW input to the data input unit 330, for example. The neural network generating device 300 may acquire the hardware information HW by displaying a GUI image required for inputting the hardware information HW on the display unit 350 and having the user input the hardware information HW from the operation input unit 360.

[0061] The hardware information HW specifically includes the memory type, memory capacity, and input / output data width of the memories to be allocated as the first memory 1 and the second memory 2.

[0062] The acquired hardware information HW is stored in the storage unit 310. Next, the neural network generation device 300 executes step S12.

[0063] <Network information acquisition step (S12)> In step S12, the neural network generation device 300 acquires the network information NW of the CNN 200 (network information acquisition step). The neural network generation device 300 acquires, for example, the network information NW input to the data input unit 330. The neural network generation device 300 may acquire the network information NW by displaying a GUI image required for inputting the network information NW on the display unit 350 and having the user input the network information NW from the operation input unit 360.

[0064] Specifically, the network information NW has a network configuration including an input layer and an output layer 230, a configuration of a convolution layer 210 including weights w and the bit width of the input data a, and a configuration of a quantization operation layer 220 including quantization information.

[0065] The acquired network information NW is stored in the storage unit 310. Next, the neural network generation device 300 executes step S13.

[0066] <Neural network execution model generation process (S13)> In step S13, the execution model generation unit 321 of the neural network generation device 300 generates the NN execution model 100 based on the hardware information HW and the network information NW (neural network execution model generation step).

[0067] The neural network execution model generating step (NN execution model generating step) includes, for example, a layer mapping step (S13-1), a convolution operation circuit generating step (S13-2), a quantization operation circuit generating step (S13-3), and a DMAC generating step (S13-4).

[0068] <Layer mapping process (S13-1)> The execution model generation unit 321 maps each layer of the CNN 200 to the convolution operation circuit 4 and the quantization operation circuit 5 formed in a loop (layer mapping step). The execution model generation unit 321 generates sequence data and software for sequentially executing each layer of the CNN 200 in the NN execution model 100. For layers including operations that cannot be performed by the NN execution model 100, such as the input layer and output layer 230, a software module that can be executed by an external operation device, such as an external host CPU, other than the NN execution model 100 is generated.

[0069] Fig. 7 is a timing chart showing an example of the operation of the NN execution model 100. The execution model generation unit 321 generates, for example, sequence data and software that enables the operation of the NN execution model 100 shown in Fig. 7 to be performed. An example of the operation of the NN execution model 100 shown in Fig. 7 will be described below.

[0070] The DMAC 3 stores input data a of the layer 1 (see FIG. 3) in the first memory 1. The DMAC 3 may divide the input data a of the layer 1 and transfer it to the first memory 1 in accordance with the order of the convolution operation performed by the convolution operation circuit 4.

[0071] The convolution operation circuit 4 reads out the input data a of layer 1 (see FIG. 3) stored in the first memory 1. The convolution operation circuit 4 performs the convolution operation of layer 1 on the input data a of layer 1. The output data f of the convolution operation of layer 1 is stored in the second memory 2.

[0072] The quantization calculation circuit 5 reads the output data f of layer 1 stored in the second memory 2. The quantization calculation circuit 5 performs a quantization calculation of layer 2 on the output data f of layer 1. The output data of the quantization calculation of layer 2 is stored in the first memory 1.

[0073] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2 stored in the first memory 1. The convolution operation circuit 4 performs a convolution operation of the layer 3 using the output data of the quantization operation of the layer 2 as input data a. The output data f of the convolution operation of the layer 3 is stored in the second memory 2.

[0074] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2M-2 (M is a natural number) stored in the first memory 1. The convolution operation circuit 4 performs the convolution operation of the layer 2M-1 using the output data of the quantization operation of the layer 2M-2 as input data a. The output data f of the convolution operation of the layer 2M-1 is stored in the second memory 2.

[0075] The quantization calculation circuit 5 reads the output data f of the layer 2M-1 stored in the second memory 2. The quantization calculation circuit 5 performs the quantization calculation of the layer 2M on the output data f of the 2M-1 layer. The output data of the quantization calculation of the layer 2M is stored in the first memory 1.

[0076] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2M stored in the first memory 1. The convolution operation circuit 4 performs a convolution operation of the layer 2M+1 using the output data of the quantization operation of the layer 2M as input data a. The output data f of the convolution operation of the layer 2M+1 is stored in the second memory 2.

[0077] The convolution operation circuit 4 and the quantization operation circuit 5 alternately perform operations to advance the operation of the CNN 200 shown in Fig. 3. In the NN execution model 100, the convolution operation circuit 4 performs the convolution operation of layers 2M-1 and 2M+1 by time sharing. Also, in the NN execution model 100, the quantization operation circuit 5 performs the quantization operation of layers 2M-2 and 2M by time sharing. Therefore, the circuit scale of the NN execution model 100 is significantly smaller than when a separate convolution operation circuit 4 and quantization operation circuit 5 are implemented for each layer.

[0078] The NN execution model 100 performs calculations of the CNN 200, which is a multilayer structure of multiple layers, using a circuit formed in a loop. The NN execution model 100 can efficiently use hardware resources due to the loop circuit configuration. Note that, since the NN execution model 100 forms a loop circuit, parameters in the convolution calculation circuit 4 and the quantization calculation circuit 5 that change in each layer are updated appropriately.

[0079] When the calculations of the CNN 200 include calculations that cannot be performed by the NN execution model 100, the NN execution model 100 transfers intermediate data to an external calculation device such as an external host CPU. After the external calculation device performs calculations on the intermediate data, the calculation results by the external calculation device are input to the first memory 1 and / or the second memory 2. The NN execution model 100 resumes calculations on the calculation results by the external calculation device.

[0080] <Convolution operation circuit generation process (S13-2)> The execution model generation unit 321 generates the convolution operation circuit 4 of the NN execution model 100 based on the hardware information HW and the network information NW (convolution operation circuit generation step). The execution model generation unit 321 divides the data of the convolution operation of the convolution layer 210 based on the memory capacities of the memories allocated as the first memory 1 and the second memory 2. The generated convolution operation circuit 4 has a configuration capable of operating the divided convolution operation data. If the size (Bc or Bd) of the blocks into which the data of the convolution operation of the convolution layer 210 is divided is reduced, the hardware scale of the convolution operation circuit 4 is reduced, but the operation efficiency of the convolution operation of the convolution layer 210 is reduced.

[0081] FIG. 8 is a diagram for explaining data division and data expansion in a convolution operation. The convolution operation circuit 4 of the NN execution model 100 divides the input data of the convolution operation (Equation 1) of the convolution layer 210 into partial tensors and performs the operation. The method of division into the partial tensors and the number of divisions are not particularly limited. The partial tensors are formed, for example, by dividing the input data a(x+i, y+j, c) into a(x+i, y+j, co). Note that the convolution operation circuit 4 of the NN execution model 100 can also perform the operation without dividing the input data of the convolution operation (Equation 1) of the convolution layer 210.

[0082] <Convolution operation circuit generation process: Data division for convolution operation> In input data division for a convolution operation, the variable c in Equation 1 is divided into blocks of size Bc, as shown in Equation 7. Also, the variable d in Equation 1 is divided into blocks of size Bd, as shown in Equation 8. In Equation 7, co is an offset, and ci is an index from 0 to (Bc-1). In Equation 8, do is an offset, and di is an index from 0 to (Bd-1). Note that size Bc and size Bd may be the same.

[0083]

number

[0084] [Number]

[0085] The input data a(x+i, y+j, c) in Equation 1 is divided in the c-axis direction by size Bc and is represented by the divided input data a(x+i, y+j, co). In the following description, the divided input data a is also referred to as "divided input data a".

[0086] The weight w(i, j, c, d) in Equation 1 is divided in the c-axis direction by size Bc and in the d-axis direction by size Bd and is represented by the divided weight w(i, j, co, do). In the following description, the divided weight w is also referred to as "divided weight w".

[0087] The output data f(x, y, do) divided by size Bd is obtained by Equation 9. By combining the divided output data f(x, y, do), the final output data f(x, y, d) can be calculated.

[0088] [Number]

[0089] <Convolution operation circuit generation step (S13-2): Data expansion> The convolution operation circuit 4 of the NN execution model 100 expands the input data a and the weight w in the convolution operation of the convolution layer 210 and performs the convolution operation.

[0090] The divided input data a(x+i, y+j, co) is expanded into vector data having Bc elements. The elements of the divided input data a are indexed by ci (0 ≦ ci < Bc). In the following description, the divided input data a expanded into vector data for each i and j is also referred to as "input vector A". The input vector A has elements from the divided input data a(x+i, y+j, co×Bc) to the divided input data a(x+i, y+j, co×Bc+(Bc-1)).

[0091] The segmentation weight w(i,j,co,do) is expanded into matrix data having Bc×Bd elements. The elements of the segmentation weight w expanded into matrix data are indexed by ci and di (0≦di<Bd). In the following description, the segmentation weight w expanded into matrix data for each i and j is also referred to as "weight matrix W". The weight matrix W has elements from the segmentation weight w(i,j,co×Bc,do×Bd) to the segmentation weight w(i,j,co×Bc+(Bc-1),do×Bd+(Bd-1)).

[0092] By multiplying the input vector A and the weight matrix W, vector data is calculated. By shaping the vector data calculated for each i, j, and co into a three-dimensional tensor, the output data f(x,y,do) can be obtained. By performing such data expansion, the convolution operation of the convolutional layer 210 can be implemented by multiplying vector data and matrix data.

[0093] The sizes (Bc and Bd) of the blocks for dividing the data of the convolution operation are set to sizes such that, for example, a predetermined number of divided input data a and a predetermined number of divided weights w can be stored in the first memory 1.

[0094] For example, assume that the size of the input data a is X×Y×C, the size of the weight w is K×K×C×D, and the size of the output data f is X×Y×D. The output data f(x,y,do) divided in the d-axis direction by the size Bd can be calculated by performing a convolution operation for each i, j, and co on the input data a(x+i,y+j,co) divided in the c-axis direction by the size Bc and the weight w(i,j,co,do) divided by the sizes Bc and Bd, and then adding them together.

[0095] If the elements of the output data f are 16 bits, the size of the output data f(x,y,do) divided by size Bd in the d-axis direction is 16·X·Y·Bd bits. On the other hand, if the elements of the input data a are 2 bits, the size of the input data a required to calculate the output data f divided by Bd is 2·X·Y·Bc bits. Also, if the elements of the weight w are 1 bit, the size of the weight w required to calculate the output data f divided by Bd is 1·K·K·Bc·Bd bits.

[0096] If the memory capacity of the second memory 2 is greater than 16·X·Y·Bd bits, the output data f(x, y, do) divided by Bd can be stored in the second memory 2. On the other hand, if the memory capacity of the first memory 1 is greater than (2·X·Y·Bc+1·K·K·Bc·Bd) bits, the input data a and weights w necessary for calculating the output data f divided by Bd can be stored in the first memory 1.

[0097] Based on the above-described relationship, when it is specified as a constraint in the hardware information HW, the size of the divided blocks (Bc or Bd) can be calculated from the upper limit of the memory capacity of the first memory 1 and the second memory 2. Also, the memory capacities of the first memory 1 and the second memory 2 can be calculated from the size of the divided blocks (Bc or Bd).

[0098] In order to enable parallel operation of the convolution operation circuit 4 and the DMAC 3, it is desirable that the memory capacity of the first memory 1 and the second memory 2 is at least twice the above-mentioned memory capacity and that double buffering can be implemented.

[0099] The above example is one example of a means for determining the size of the divided blocks (Bc or Bd) and the memory capacity of the first memory 1 and the second memory 2. The determination of the size of the divided blocks (Bc or Bd) and the memory capacity of the first memory 1 and the second memory 2 is appropriately changed depending on the memory usage mode, the number of parallel operations, etc.

[0100] <Convolution operation circuit generation process (S13-2): Division into partial tensors (1)> FIG. 9 is a timing chart showing another example of the operation of the NN execution model 100. In FIG. The NN execution model 100 may divide the input data a into partial tensors and perform operations on the partial tensors in a time-division manner.

[0101] Fig. 9 shows an example of operation when input data a is decomposed into two partial tensors. The decomposed partial tensors are called "first partial tensor a1" and "second partial tensor a2". For example, the convolution operation of layer 2M-1 is decomposed into a convolution operation corresponding to the first partial tensor a1 (in Fig. 9, indicated as "layer 2M-1(a1)") and a convolution operation corresponding to the second partial tensor a2 (in Fig. 7, indicated as "layer 2M-1(a2)").

[0102] The convolution operation and quantization operation corresponding to the first partial tensor a1 and the convolution operation and quantization operation corresponding to the second partial tensor a2 can be performed independently, as shown in FIG.

[0103] The convolution operation circuit 4 performs a convolution operation of the layer 2M-1 corresponding to the first partial tensor a1 (operation shown as layer 2M-1(a1) in FIG. 9). After that, the convolution operation circuit 4 performs a convolution operation of the layer 2M-1 corresponding to the second partial tensor a2 (operation shown as layer 2M-1(a2) in FIG. 9). In addition, the quantization operation circuit 5 performs a quantization operation of the layer 2M corresponding to the first partial tensor a1 (operation shown as layer 2M(a1) in FIG. 9). In this way, the NN execution model 100 can perform the convolution operation of the layer 2M-1 corresponding to the second partial tensor a2 and the quantization operation of the layer 2M corresponding to the first partial tensor a1 in parallel.

[0104] Next, the convolution operation circuit 4 performs a convolution operation of layer 2M+1 corresponding to the first partial tensor a1 (operation shown as layer 2M+1(a1) in FIG. 9). In addition, the quantization operation circuit 5 performs a quantization operation of layer 2M corresponding to the second partial tensor a2 (operation shown as layer 2M(a2) in FIG. 9). In this way, the NN execution model 100 can perform, in parallel, the convolution operation of layer 2M+1 corresponding to the first partial tensor a1 and the quantization operation of layer 2M corresponding to the second partial tensor a2.

[0105] By dividing the input data a into partial tensors, the NN execution model 100 can operate the convolution operation circuit 4 and the quantization operation circuit 5 in parallel. As a result, the waiting time of the convolution operation circuit 4 and the quantization operation circuit 5 is reduced, improving the efficiency of the operation processing of the NN execution model 100. In the operation example shown in FIG. 9, the number of divisions into partial tensors is 2, but even when the number of divisions is more than 2, the NN execution model 100 can similarly operate the convolution operation circuit 4 and the quantization operation circuit 5 in parallel.

[0106] <Convolution operation circuit generation process (S13-2): Division into partial tensors (2)> FIG. 10 is a diagram showing a partial tensor ft obtained by dividing output data f of a convolution operation into tiles. The input data at is obtained by dividing the input data a into tiles (blocks) of a predetermined size in the x-axis direction and the y-axis direction. The partial tensor ft is obtained by dividing the output data f into tiles (blocks) of size T in the x-axis direction and the y-axis direction. As in the above example, the size of the input data a is X×Y×C, the size of the weight w is K×K×C×D, and the size of the output data f is X×Y×D. The size of the partial tensor ft is T·T·D.

[0107] The convolution operation circuit 4 reads a portion of the input data at from the first memory 1, and performs a convolution operation of the layer 2M-1 with a partial tensor ft (referred to as a first partial tensor ft1) as an output. The first partial tensor ft1 is written to the second memory 2. Before performing a convolution operation on the remaining portion of the input data a stored in the first memory 1, the quantization operation circuit 5 performs a quantization operation of the layer 2M corresponding to the first partial tensor ft1 stored in the second memory 2. The output data of the quantization operation of the layer 2M is written to the first memory 1. As a result, the first partial tensor ft1 written to the second memory 2 becomes unnecessary.

[0108] Next, the convolution operation circuit 4 reads another part of the input data at from the first memory 1, and performs a convolution operation of the layer 2M-1 with the partial tensor ft (referred to as the second partial tensor ft2) as an output. The second partial tensor ft2 is written to the second memory 2. The second partial tensor ft2 overwrites the first partial tensor ft1. Before performing the convolution operation on the remaining part of the input data a stored in the first memory 1, the quantization operation circuit 5 performs a quantization operation of the layer 2M corresponding to the second partial tensor ft2 stored in the second memory 2. The output data of the quantization operation of the layer 2M is written to the first memory 1. As a result, the second partial tensor ft2 written to the second memory 2 becomes unnecessary.

[0109] By performing the above calculation on the remaining part of the input data a from the first memory 1, the convolution calculation of layer 2M-1 and the quantization calculation of layer 2M are completed. By dividing the output data f of the convolution calculation into tiles in this way, the size of the second memory 2 can be reduced to a memory size that can store one partial tensor ft. For example, the output data f is assumed to be 16 bits. When tiling is not performed, the second memory 2 needs to store the output data f, and the required memory size is 16·X·Y·D bits. On the other hand, when tiling is performed, the second memory 2 only needs to store one partial tensor ft, and the required memory size is 16·T 2 Reduced to D bits.

[0110] On the other hand, when tile partitioning is used, it is necessary to separately store the input data a of the convolution operation of layer 2M-1 and the output data of the quantization operation of layer 2M in the first memory 1. However, by making the size of T sufficiently smaller than X and Y, it is possible to reduce the overall memory capacity of the first memory 1 and the second memory 2.

[0111] <Convolution operation circuit generation process (S13-2): Division into partial tensors (3)> 11 to 13 are diagrams showing partial tensors as obtained by dividing input data a into slices. The partial tensors as are obtained by dividing input data a into slices (blocks) of a predetermined size in the y-axis direction.

[0112] As shown in FIG. 11, the convolution operation circuit 4 reads a part of the partial tensor as (first partial tensor as1) from the first memory 1, and performs a convolution operation of the layer 2M-1 with the partial tensor ft (referred to as the first partial tensor ft1) as an output. The first partial tensor ft1 is written to the second memory 2. Before performing a convolution operation on the remaining part of the first partial tensor as1 stored in the first memory 1, the quantization operation circuit 5 performs a quantization operation of the layer 2M corresponding to the first partial tensor ft1 stored in the second memory 2. The output data of the quantization operation of the layer 2M is written to the first memory 1. As a result, the first partial tensor ft1 written to the second memory 2 becomes unnecessary.

[0113] Next, the convolution operation circuit 4 reads another part of the first partial tensor as1 from the first memory 1, and performs a convolution operation of the layer 2M-1 with the partial tensor ft (referred to as the second partial tensor ft2) as an output. The second partial tensor ft2 is written to the second memory 2. The second partial tensor ft2 overwrites the first partial tensor ft1. Before performing the convolution operation on the remaining part of the first partial tensor as1 stored in the first memory 1, the quantization operation circuit 5 performs a quantization operation of the layer 2M corresponding to the second partial tensor ft2 stored in the second memory 2. The output data of the quantization operation of the layer 2M is written to the first memory 1. As a result, the second partial tensor ft2 written to the second memory 2 becomes unnecessary.

[0114] By performing the above operation on the remaining part of the first partial tensor as1 from the first memory 1, the convolution operation of layer 2M-1 and the quantization operation of layer 2M on the first partial tensor as1 are completed. As a result, the first partial tensor as1 written to the first memory 1 becomes unnecessary.

[0115] 12, the convolution operation circuit 4 and the quantization operation circuit 5 similarly perform a convolution operation and a quantization operation on another partial tensor as (second partial tensor as2) from the first memory 1. The first partial tensor as1 written to the first memory 1 is unnecessary and may be overwritten with the output data of the quantization operation of the layer 2M. When the convolution operation of the layer 2M-1 and the quantization operation of the layer 2M on the second partial tensor as2 are completed, the second partial tensor as2 written to the first memory 1 becomes unnecessary.

[0116] 13, the convolution operation circuit 4 and the quantization operation circuit 5 similarly perform a convolution operation and a quantization operation on another partial tensor as (third partial tensor as3) from the first memory 1. The first partial tensor as1 and the second partial tensor as2 written to the first memory 1 are unnecessary and may be overwritten with the output data of the quantization operation of the layer 2M. When the convolution operation of the layer 2M-1 and the quantization operation of the layer 2M on the third partial tensor as3 are completed, the third partial tensor as3 written to the first memory 1 becomes unnecessary.

[0117] By performing the above calculations on the remaining part of the input data a from the first memory 1, all of the convolution calculations of layer 2M-1 and the quantization calculations of layer 2M are completed. By dividing the input data a into slices in this manner, the size of the first memory 1 can be reduced compared to the example shown in FIG.

[0118] <Convolution operation circuit generation process (S13-2): Division into partial tensors (4)> FIG. 14 is a diagram showing other partial tensors required to output a partial tensor ft by a convolution operation of layer 2M+1.

[0119] In order to perform the convolution operation of layer 2M+1 and output the partial tensor ft, a partial tensor of the input of the quantization operation of layer 2M is required. Furthermore, a partial tensor of the input of the convolution operation of layer 2M-1 is required. In this way, there is a dependency relationship between the partial tensors required to output the partial tensor ft. The partial tensor ft may be calculated by sequentially calculating the partial tensors required to output the partial tensor ft based on this dependency relationship. The memory sizes of the first memory 1 and the second memory 2 need only be large enough to store the partial tensors, and the overall memory capacity of the first memory 1 and the second memory 2 can be reduced.

[0120] The sizes of the various partial tensors described above are set, for example, to sizes that allow a predetermined number of partial tensors to be stored in the first memory 1 and the second memory 2. The memory capacities of the first memory 1 and the second memory 2 may be calculated from the sizes of the partial tensors.

[0121] <Convolution operation circuit generation process (S13-2): Hardware model generation> Next, the execution model generation unit 321 generates a hardware model of the convolution operation circuit 4 from information such as the weight w and the bit width of the input data a input as the network information NW. The hardware model may be at a behavior level, at an RTL (Register Transfer Level), or a netlist representing connections between gates and circuit modules, or a combination of these. An example of the generated hardware model of the convolution operation circuit 4 will be described below.

[0122] FIG. 15 is an internal block diagram of the convolution operation circuit 4 to be generated. The convolution operation circuit 4 has a weight memory 41, a multiplier 42, an accumulator circuit 43, and a state controller 44. The convolution operation circuit 4 has a state controller 44 dedicated to the multiplier 42 and the accumulator circuit 43, and when an instruction command is input, the convolution operation can be performed without requiring an external controller.

[0123] The weight memory 41 is a memory in which the weight w used in the convolution calculation is stored, and is a rewritable memory such as a volatile memory constituted by, for example, an SRAM (Static RAM), etc. The DMAC 3 writes the weight w required for the convolution calculation into the weight memory 41 by DMA transfer.

[0124] FIG. 16 is an internal block diagram of the multiplier 42. The multiplier 42 multiplies the input vector A by the weight matrix W. As described above, the input vector A is vector data having Bc elements into which the divided input data a(x+i, y+j, co) is expanded. The weight matrix W is matrix data having Bc×Bd elements into which the divided weights w(i, j, co, do) are expanded. The multiplier 42 has Bc×Bd product-sum operation units 47 and can multiply the input vector A by the weight matrix W in parallel.

[0125] The multiplier 42 reads out the input vector A and the weight matrix W required for the multiplication from the first memory 1 and the weight memory 41, and performs the multiplication. The multiplier 42 outputs Bd product-sum operation results O(di).

[0126] FIG. 17 is an internal block diagram of the product-sum calculation unit 47. As shown in FIG. The multiply-add unit 47 multiplies an element A(ci) of an input vector A by an element W(ci,di) of a weight matrix W. The multiply-add unit 47 also adds the multiplication result to a multiplication result S(ci,di) of another multiply-add unit 47. The multiply-add unit 47 outputs an addition result S(ci+1,di). The element A(ci) is a 2-bit unsigned integer (0,1,2,3). The element W(ci,di) is a 1-bit signed integer (0,1), where a value "0" represents +1 and a value "1" represents -1.

[0127] The multiply-and-accumulate unit 47 has an inverter 47a, a selector 47b, and an adder 47c. The multiply-and-accumulate unit 47 performs multiplication using only the inverter 47a and the selector 47b, without using a multiplier. When the element W(ci,di) is "0", the selector 47b selects the input of the element A(ci). When the element W(ci,di) is "1", the selector 47b selects the complement of the element A(ci) inverted by the inverter. The element W(ci,di) is also input to the carry-in of the adder 47c. When the element W(ci,di) is "0", the adder 47c outputs a value obtained by adding the element A(ci) to S(ci,di). When W(ci,di) is "1", the adder 47c outputs a value obtained by subtracting the element A(ci) from S(ci,di).

[0128] FIG. 18 is an internal block diagram of the accumulator circuit 43. The accumulator circuit 43 accumulates the product-sum operation results O(di) of the multiplier 42 in the second memory 2. The accumulator circuit 43 has Bd accumulator units 48 and can accumulate the Bd product-sum operation results O(di) in parallel in the second memory 2.

[0129] FIG. 19 is an internal block diagram of the accumulator unit 48. The accumulator unit 48 has an adder 48a and a mask section 48b. The adder 48a adds an element O(di) of the multiply-and-accumulate result O to a partial sum that is an intermediate result of the convolution operation shown in Equation 1 and stored in the second memory 2. The addition result is 16 bits per element. The addition result is not limited to 16 bits per element, and may be, for example, 15 bits or 17 bits per element.

[0130] The adder 48a writes the addition result to the same address in the second memory 2. When the initialization signal clear is asserted, the masking unit 48b masks the output from the second memory 2 and sets the addition target for the element O(di) to zero. The initialization signal clear is asserted when no intermediate partial sums are stored in the second memory 2.

[0131] When the convolution operation by the multiplier 42 and the accumulator circuit 43 is completed, the output data f(x, y, do) is stored in the second memory 2.

[0132] The state controller 44 controls the states of the multiplier 42 and the accumulator circuit 43. The state controller 44 is also connected to the controller 6 via an internal bus IB. The state controller 44 has an instruction queue 45 and a control circuit 46.

[0133] The instruction queue 45 is a queue that stores the instruction commands C4 for the convolution operation circuit 4, and is configured by, for example, a FIFO memory. The instruction commands C4 are written to the instruction queue 45 via the internal bus IB.

[0134] The control circuit 46 is a state machine that decodes the instruction command C4 and controls the multiplier 42 and the accumulator circuit 43 based on the instruction command C4. The control circuit 46 may be implemented by a logic circuit or a CPU controlled by software.

[0135] FIG. 20 is a state transition diagram of the control circuit 46. When the instruction command C4 is input to the instruction queue 45 (Not empty), the control circuit 46 transitions from the idle state S1 to the decode state S2.

[0136] In the decode state S2, the control circuit 46 decodes the instruction command C3 output from the instruction queue 45. The control circuit 46 also reads the semaphore S stored in the register 61 of the controller 6, and determines whether the operations of the multiplier 42 and the accumulator circuit 43 instructed in the instruction command C4 are executable. If they are not executable (Not ready), the control circuit 46 waits (Wait) until they are executable. If they are executable (Ready), the control circuit 46 transitions from the decode state S2 to the execution state S3.

[0137] In the execution state S3, the control circuit 46 controls the multiplier 42 and the accumulator circuit 43 to cause the multiplier 42 and the accumulator circuit 43 to perform the operation instructed by the instruction command C4. When the operation of the multiplier 42 and the accumulator circuit 43 is completed, the control circuit 46 removes the executed instruction command C4 from the instruction queue 45 and updates the semaphore S stored in the register 61 of the controller 6. If there is an instruction in the instruction queue 45 (Not empty), the control circuit 46 transitions from the execution state S3 to the decode state S2. If there is no instruction in the instruction queue 45 (empty), the control circuit 46 transitions from the execution state S3 to the idle state S1.

[0138] 16, the execution model generation unit 321 associates the size (Bc or Bd) of the blocks into which the data of the convolution operation is divided with the number (Bc×Bd) of the multiply-add operation units 47. If the size (Bc or Bd) of the blocks into which the data of the convolution operation of the convolution layer 210 is divided is reduced, the hardware scale of the multiplier 42 is reduced, but the operation speed of the multiplier 42 is reduced.

[0139] 16, an input vector A having Bc elements and a weighting matrix W having Bc×Bd elements are input to the multiplier 42. Therefore, even if the number of product-sum calculation units 47 is made greater than Bc×Bd, the product-sum calculation units 47 cannot be used effectively.

[0140] It is desirable that the size of the input data a and weight w in the c-axis and d-axis directions and the block size (Bc and Bd) be a power of two, such as 64, 128, or 256, in order to efficiently perform division, data integration, etc.

[0141] Reducing the bit width of the weight w and the input data a input as the network information NW can reduce the hardware scale of the multiplier 42 and the accumulator circuit 43. Reducing the bit width of the weight w and the input data a can reduce the memory capacity of the first memory 1 and the second memory 2 that store them. Also, the time it takes for the DMAC 3 to transfer data to the first memory 1 and the second memory 2 can be shortened.

[0142] <Quantization operation circuit generation process (S13-3)> The execution model generation unit 321 generates the quantization operation circuit 5 of the NN execution model 100 based on the hardware information HW and the network information NW (quantization operation circuit generation process). The execution model generation unit 321 generates a hardware model of the quantization operation circuit 5 from the quantization information input as the network information NW. The hardware model may be at a behavior level, at an RTL (Register Transfer Level), or a netlist representing connections between gates and circuit modules, or a combination of these. An example of the generated hardware model of the quantization operation circuit 5 will be described below.

[0143] FIG. 21 is an internal block diagram of the quantization calculation circuit 5 to be generated. The quantization operation circuit 5 has a quantization parameter memory 51, a vector operation circuit 52, a quantization circuit 53, and a state controller 54. The quantization operation circuit 5 has a state controller 54 dedicated to the vector operation circuit 52 and the quantization circuit 53, and when an instruction command is input, the quantization operation can be performed without the need for an external controller.

[0144] The quantization parameter memory 51 is a memory in which the quantization parameter q used in the quantization operation is stored, and is a rewritable memory such as a volatile memory constituted by, for example, an SRAM (Static RAM), etc. The DMAC 3 writes the quantization parameter q required for the quantization operation into the quantization parameter memory 51 by DMA transfer.

[0145] FIG. 22 is an internal block diagram of the vector calculation circuit 52 and the quantization circuit 53. The vector operation circuit 52 performs an operation on the output data f(x, y, do) stored in the second memory 2. The vector operation circuit 52 has Bd operation units 57, and performs SIMD operations in parallel on the output data f(x, y, do).

[0146] FIG. 23 is a block diagram of the arithmetic unit 57. The arithmetic unit 57 includes, for example, an ALU 57a, a first selector 57b, a second selector 57c, a register 57d, and a shifter 57e. The arithmetic unit 57 may further include other arithmetic units included in a known general-purpose SIMD arithmetic circuit.

[0147] The vector calculation circuit 52 performs at least one of the calculations of the pooling layer 221, the batch normalization layer 222, and the activation function layer 223 in the quantization calculation layer 220 on the output data f(x, y, do) by combining the calculation units and the like contained in the calculation unit 57.

[0148] The arithmetic unit 57 can add the data stored in the register 57d and the element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The arithmetic unit 57 can store the addition result by the ALU 57a in the register 57d. The arithmetic unit 57 can initialize the addition result by inputting "0" to the ALU 57a instead of the data stored in the register 57d by the selection of the first selector 57b. For example, when the pooling area is 2×2, the shifter 57e can output the average value of the addition result by shifting the output of the ALU 57a to the right by 2 bits. The vector arithmetic circuit 52 can perform the average pooling calculation shown in Equation 2 by repeating the above calculations by the Bd arithmetic units 57.

[0149] The arithmetic unit 57 can compare the data stored in the register 57d with the element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The arithmetic unit 57 controls the second selector 57c according to the comparison result by the ALU 57a, and can select the larger of the data stored in the register 57d and the element f(di). The arithmetic unit 57 can initialize the comparison target to the minimum value by inputting the minimum value of the possible values ​​of the element f(di) to the ALU 57a by the selection of the first selector 57b. In this embodiment, the element f(di) is a 16-bit signed integer, so the minimum value of the possible values ​​of the element f(di) is "0x8000". The vector arithmetic circuit 52 can perform the MAX pooling calculation of Equation 3 by repeating the above calculations by the Bd arithmetic units 57. Note that in the MAX pooling calculation, the shifter 57e does not shift the output of the second selector 57c.

[0150] The arithmetic unit 57 can subtract data stored in the register 57d and an element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The shifter 57e can shift the output of the ALU 57a to the left (i.e., multiplication) or to the right (i.e., division). The vector arithmetic circuit 52 can perform the batch normalization calculation of Equation 4 by repeating the above calculations by the Bd arithmetic units 57.

[0151] The arithmetic unit 57 can compare the element f(di) of the output data f(x, y, do) read from the second memory 2 with "0" selected by the first selector 57b by the ALU 57a. The arithmetic unit 57 can select and output either the element f(di) or a constant value "0" previously stored in the register 57d according to the comparison result by the ALU 57a. The vector arithmetic circuit 52 can perform the ReLU arithmetic of Equation 5 by repeating the above arithmetic operations by the Bd arithmetic units 57.

[0152] The vector operation circuit 52 can perform average pooling, MAX pooling, batch normalization, activation function operations, and combinations of these operations. Since the vector operation circuit 52 can perform general-purpose SIMD operations, it may perform other operations necessary for the operations in the quantization operation layer 220. In addition, the vector operation circuit 52 may perform operations other than those in the quantization operation layer 220.

[0153] It is to be noted that the quantization calculation circuit 5 does not have to include the vector calculation circuit 52. When the quantization calculation circuit 5 does not include the vector calculation circuit 52, the output data f(x, y, do) is input to the quantization circuit 53.

[0154] The quantization circuit 53 quantizes the output data of the vector operation circuit 52. The quantization circuit 53 has Bd quantization units 58, as shown in FIG.

[0155] FIG. 24 is an internal block diagram of the quantization unit 58. The quantization unit 58 quantizes the element in(di) of the output data of the vector operation circuit 52. The quantization unit 58 has a comparator 58a and an encoder 58b. The quantization unit 58 performs the operation (Equation 6) of the quantization layer 224 in the quantization operation layer 220 on the output data (16 bits / element) of the vector operation circuit 52. The quantization unit 58 reads out the necessary quantization parameters q(th0, th1, th2) from the quantization parameter memory 51, and compares the input in(di) with the quantization parameter q by the comparator 58a. The quantization unit 58 quantizes the comparison result by the comparator 58a to 2 bits / element by the encoder 58b. Since α(c) and β(c) in Equation 4 are parameters that differ for each variable c, the quantization parameters q(th0, th1, th2) that reflect α(c) and β(c) are parameters that differ for each in(di).

[0156] The quantization unit 58 classifies the input in(di) into four regions (for example, in≦th0, th0<in≦th1, th1<in≦th2, th2<in) by comparing the input in(di) with three thresholds th0, th1, th2, encodes the classification result into 2 bits, and outputs it. The quantization unit 58 can also perform operations of Batch Normalization and activation functions together with quantization by setting quantization parameters q(th0, th1, th2).

[0157] The quantization unit 58 can perform the operation of Batch Normalization shown in Equation 4 together with quantization by setting the threshold th0 as β(c) in Equation 4 and the threshold differences (th1−th0) and (th2−th1) as α(c) in Equation 4. By increasing (th1−th0) and (th2−th1), α(c) can be decreased. By decreasing (th1−th0) and (th2−th1), α(c) can be increased.

[0158] The quantization unit 58 can perform the ReLU operation of the activation function together with the quantization of the input in(di). For example, the quantization unit 58 saturates the output value in the regions where in(di)≦th0 and th2<in(di). The quantization unit 58 can perform the operation of the activation function together with quantization by setting the quantization parameter q so that the output is non-linear.

[0159] The state controller 54 controls the states of the vector operation circuit 52 and the quantization circuit 53. Also, the state controller 54 is connected to the controller 6 via the internal bus IB. The state controller 54 has an instruction queue 55 and a control circuit 56.

[0160] The instruction queue 55 is a queue in which the instruction command C5 for the quantization operation circuit 5 is stored, and is composed of, for example, a FIFO memory. The instruction command C5 is written into the instruction queue 55 via the internal bus IB.

[0161] The control circuit 56 is a state machine that decodes the instruction command C5 and controls the vector operation circuit 52 and the quantization circuit 53 based on the instruction command C5. The control circuit 56 has a similar configuration to the control circuit 46 of the state controller 44 of the convolution operation circuit 4.

[0162] The quantization calculation circuit 5 writes the quantization calculation output data having Bd elements into the first memory 1. A suitable relationship between Bd and Bc is shown in Equation 10. In Equation 10, n is an integer.

[0163]

number

[0164] The execution model generation unit 321 determines, from the quantization information input as the network information NW, whether or not to perform pooling calculations in the quantization calculation circuit 5 and the type of the calculation, whether or not to perform batch normalization calculations and the type of the calculation, whether or not to perform activation function calculations and the type of the calculation, the quantization type, and whether or not to perform other calculations.

[0165] For example, when performing a pooling operation in the quantization operation circuit 5, the execution model generation unit 321 generates a calculation unit 57 optimized for the type of pooling (average pooling, MAX pooling, etc.) to be performed.

[0166] For example, when an activation function is calculated in the quantization calculation circuit 5, the execution model generation unit 321 generates the calculation unit 57 and the quantization unit 58 optimized for the activation function (such as ReLU calculation) to be calculated.

[0167] For example, when performing batch normalization calculation in the quantization calculation circuit 5, the execution model generation unit 321 generates the calculation unit 57 according to the batch normalization calculation. In addition, the execution model generation unit 321 adjusts the quantization parameters q (th0, th1, th2) according to the batch normalization calculation.

[0168] For example, when the quantization by the quantization operation circuit 5 is quantization of 3 bits or more, the execution model generation unit 321 generates a vector operation circuit 52 capable of performing pooling, Batch Normalization, and scaling for quantization.

[0169] For example, in order to reduce the operation load of the quantization operation circuit 5, the operation of normalizing Batch Normalization may be made more efficient. Specifically, in order to use bit shift in the normalization process of Batch Normalization, each element of the input tensor is made a power of 2. Thereby, the operation of normalizing Batch Normalization can be realized only by bit shift. Here, an additional operation circuit for converting each element of the input tensor into a power of 2 may be added to the quantization operation circuit 5 or may be added to the convolution operation circuit 4.

[0170] <DMAC Generation Step (S13-4)> The execution model generation unit 321 generates the DMAC 3 of the NN execution model 100 based on the hardware information HW and the network information NW (DMAC generation step). The execution model generation unit 321 generates a hardware model of the DMAC 3 from the information input as the network information NW. The hardware model may be at the behavioral level, may be RTL (Register Transfer Level), may be a netlist representing the connection between gates and circuit modules, or may be a combination thereof. Hereinafter, an example of the hardware model of the generated DMAC 3 will be described.

[0171] FIG. 25 is an internal block diagram of the generated DMAC 3. The DMAC 3 has a data transfer circuit 31 and a state controller 32. The DMAC 3 has a dedicated state controller 32 for the data transfer circuit 31, and when an instruction command is input, it can perform DMA data transfer without requiring an external controller.

[0172] The data transfer circuit 31 is connected to the external bus EB, and performs DMA data transfer between an external memory such as a DRAM and the first memory 1. The data transfer circuit 31 also performs DMA data transfer between an external memory such as a DRAM and the second memory 2. The data transfer circuit 31 also performs data transfer between an external memory such as a DRAM and the convolution operation circuit 4. The data transfer circuit 31 also performs data transfer between an external memory such as a DRAM and the quantization operation circuit 5. The number of DMA channels of the data transfer circuit 31 is not limited. For example, the first memory 1 and the second memory 2 may each have a dedicated DMA channel.

[0173] The state controller 32 controls the state of the data transfer circuit 31. The state controller 32 is also connected to the controller 6 via an internal bus IB. The state controller 32 has an instruction queue 33 and a control circuit .

[0174] The instruction queue 33 is a queue that stores instruction commands C3 for the DMAC 3, and is configured, for example, by a FIFO memory. One or more instruction commands C3 are written to the instruction queue 33 via the internal bus IB.

[0175] The control circuit 34 is a state machine that decodes the instruction command C3 and sequentially controls the data transfer circuit 31 based on the instruction command C3. The control circuit 34 has a similar configuration to the control circuit 46 of the state controller 44 of the convolution operation circuit 4.

[0176] The execution model generating unit 321 determines the number of DMA channels and the data bus width in the DMAC 3 from the information input as the network information NW.

[0177] For example, the execution model generation unit 321 generates a DMAC 3 with specifications (such as data bus width) that match the specifications of the external bus EB on the host side. By increasing the data bus width or the number of DMA channels, the data transmission speed between the external memory and the first memory 1 or second memory 2 can be improved.

[0178] <Learning process (S14)> In step S14, the learning unit 322 and the inference unit 323 of the neural network generation device 300 use the learning dataset DS to learn the learning parameters of the generated NN execution model 100 (learning step). The learning step (S14) includes, for example, a learned parameter generation step (S14-1) and an inference test step (S14-2).

[0179] <Learning process: Learned parameter generation process (S14-1)> The learning unit 322 generates learned parameters PM using the NN execution model 100 and the learning data D1. The learned parameters PM include learned weights w and quantization parameters q.

[0180] For example, when the NN execution model 100 is an execution model of the CNN 200 that performs image recognition, the learning data D1 is a combination of an input image and training data T. The input image is input data a that is input to the CNN 200. The training data T is the type of subject captured in the image, the presence or absence of a detection target in the image, the coordinate values ​​of the detection target in the image, and the like.

[0181] The learning unit 322 generates the learned parameters PM by supervised learning using a known technique such as backpropagation. The learning unit 322 obtains the difference E between the output of the NN execution model 100 for the input image and the teacher data T corresponding to the input image using a loss function (error function), and updates the weights w and quantization parameters q (th0, th1, th2) so that the difference E becomes smaller.

[0182] For example, when updating a weight w, the gradient of the loss function with respect to the weight w is used. The gradient is calculated, for example, by differentiating the loss function. When using backpropagation, the gradient is calculated by backpropagation.

[0183] When generating the learned parameters PM, the learning unit 322 makes operations related to convolution operations, operations related to quantization operations, and the like more accurate than the operations performed by the NN execution model 100.

[0184] The learning unit 322 improves the accuracy of calculations related to convolution calculations when calculating the gradient and updating the weights w. Specifically, 32-bit floating-point weights w, which are more accurate than the low-bit weights w (e.g., 1 bit) used by the NN execution model 100, are used for learning. Also, the accuracy of the convolution calculations performed in the convolution calculation circuit 4 of the NN execution model 100 is improved.

[0185] The learning unit 322 improves the accuracy of calculations related to the activation function when calculating the gradient and updating the weight w. Specifically, a Sigmond function, which is more accurate than an activation function such as a ReLU function implemented in the quantization calculation circuit 5 of the NN execution model 100, is used for learning.

[0186] On the other hand, when the learning unit 322 calculates output data for an input image by forward propagation, the learning unit 322 does not improve the accuracy of the convolution operation and the operation related to the activation function, but performs the operation based on the NN execution model 100. The high-accuracy weight w used when updating the weight w is converted to a low-bit value by a lookup table or the like.

[0187] When calculating the gradient and updating the weights w, the learning unit 322 improves the accuracy of the convolution operation and the operation related to the activation function, thereby preventing a decrease in accuracy of intermediate data in the operation and generating learned parameters PM that can achieve high inference accuracy.

[0188] On the other hand, when calculating output data for an input image, the learning unit 322 does not improve the accuracy of the forward propagation calculation, but performs a calculation based on the NN execution model 100. Therefore, the output data calculated by the learning unit 322 matches the output data of the NN execution model 100 that uses the generated trained parameters PM.

[0189] <Learning process: Inference test process (S14-2)> The inference unit 323 performs an inference test using the trained parameters PM generated by the learning unit 322, the NN execution model 100, and the test data D2. For example, when the NN execution model 100 is an execution model of the CNN 200 that performs image recognition, the test data D2 is a combination of an input image and teacher data T, similar to the training data D1.

[0190] The inference unit 323 displays the progress and results of the inference test on the display unit 350. The result of the inference test is, for example, the percentage of correct answers for the test data D2.

[0191] <Confirmation process (S15)> In step S15, the inference unit 323 of the neural network generation device 300 causes the display unit 350 to display a message prompting the user to input a confirmation of the result from the operation input unit 360 and a GUI image required for information input. The user inputs from the operation input unit 360 whether or not he / she accepts the result of the inference test. If an input indicating that the user accepts the result of the inference test is input from the operation input unit 360, the neural network generation device 300 then performs step S16. If an input indicating that the user does not accept the result of the inference test is input from the operation input unit 360, the neural network generation device 300 performs step S12 again. Note that the neural network generation device 300 may return to step S11 and have the user re-input the hardware information HW.

[0192] <Output process (S16)> In step S16, the hardware generation unit 324 of the neural network generation device 300 generates the neural network hardware model 400 based on the hardware information HW and the NN execution model 100. Next, the neural network generation device 300 performs step S17 and ends the process.

[0193] As described above, the neural network generation device 300, the neural network generation method, and the neural network generation program according to this embodiment make it possible to generate a neural network execution model 100 and a neural network hardware model 400 that can be incorporated into embedded devices such as IoT devices and can operate with high performance.

[0194] Although the first embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and design modifications and the like that do not depart from the gist of the present invention are also included. Also, the components shown in the above embodiment and modified examples can be appropriately combined to form a configuration.

[0195] Second Embodiment In the following description of the neural network generation device 300B according to the second embodiment of the present invention, with reference to Fig. 26 to Fig. 28, the same reference numerals are used for configurations common to those already described, and duplicated descriptions are omitted. The neural network generation device 300B differs from the neural network generation device 300 of the first embodiment only in the learning step (S14-1). The learning step (S14-1) in this embodiment will be described below.

[0196] <Learning process: Learned parameter generation process (S14-1)> The learning unit 322 generates learned parameters PM using the NN execution model 100 and the learning data D1. The learned parameters PM include learned weights w, quantization parameters q, and scaling coefficients sf.

[0197] The learning unit 322 learns the quantization parameters q (th0, th1, th2) and also learns the scaling coefficient sf (also called a scaling factor or a step size). The scaling coefficient sf is a coefficient indicating the scale of the quantized quantization operation output data, and specifically, is a coefficient by which the quantization operation output data is multiplied.

[0198] 26 to 28 are diagrams for explaining the scaling coefficient sf in the quantization operation. The quantization operation outputs quantization operation output data quantized to 2 bits (0, 1, 2, 3) based on the quantization parameter q (th0, th1, th2) as shown in Fig. 26. The scaling coefficient sf is a coefficient by which the quantization operation output data is multiplied as shown in Figs. 27 and 28. The quantization parameter q (th0, th1, th2) is a parameter suitable for the range of input data in the quantization operation, and is a multi-bit parameter of, for example, 8 bits or more.

[0199] The learning unit 322 learns the scaling factor sf so that, for example, the range of the input data in the quantization operation and the range of the data obtained by multiplying the quantization operation output data by the scaling factor sf become closer to each other. For example, when the range of the input data is narrow, the scaling factor sf becomes small. On the other hand, when the range of the input data is wide, the scaling factor sf becomes large. By multiplying the quantization operation output data by the learned scaling factor sf in this way, the learning unit 322 can reduce the decrease in accuracy associated with the quantization operation.

[0200] The scaling coefficient sf is a parameter that is learned for each layer, for example. In this case, the learning unit 322 can learn the optimal scaling coefficient sf for the range of input data in the quantization operation for each layer. Note that the scaling coefficient sf is not limited to being learned for each layer, and may be learned for each element O(di), for example.

[0201] As in the first embodiment, the learning unit 322 improves the accuracy of operations related to the convolution operation when generating the learned parameters. In this embodiment, the learning unit 322 uses data obtained by multiplying the quantization operation output data by the scaling factor sf as input data for the convolution operation. By improving the accuracy of the convolution operation and applying the scaling factor sf to the quantization operation output data, the learning unit 322 can prevent a decrease in accuracy of intermediate data in the operation and generate the learned parameters PM that can achieve higher inference accuracy.

[0202] On the other hand, in the NN execution model 100, the quantization operation output data is 2 bits, and the scaling coefficient sf is not directly multiplied to the quantization operation output data. Therefore, in the trained NN execution model 100 (including the NN execution model 100 and the trained parameters PM) used during inference rather than during learning, the trained scaling coefficient sf is incorporated into parameters of other operations. The parameters of other operations are parameters of the software that controls the trained parameters PM and the NN execution model 100, such as the quantization parameters q (th0, th1, th2), the threshold of the activation function, the parameters of batch normalization, and the weight w.

[0203] For example, the learned scaling factor sf is incorporated into the quantization parameter q(th0, th1, th2). Specifically, the quantization parameter q(th0, th1, th2) is replaced with a value obtained by dividing the quantization parameter q(th0, th1, th2) by the scaling factor sf.

[0204] For example, the learned scaling factor sf is incorporated into the parameters of batch normalization. Specifically, α(c) in Equation 4 is replaced with a value obtained by dividing α(c) by the scaling factor sf.

[0205] In the CNN 200, the type of calculation performed may differ for each layer. Therefore, the learned scaling coefficient sf is incorporated as a parameter of a calculation appropriately selected from the calculations performed by each layer.

[0206] In the present embodiment, an example has been shown in which the learning unit 332 learns a scaling coefficient sf for the quantized quantization operation output data. The learning unit 332 may use a scaling coefficient for the quantized weight w when quantizing the highly accurate weight w learned in the learning process using a lookup table or the like. The learning unit 332 can learn the scaling coefficient for the weight w by a method similar to the method for learning the scaling coefficient sf for the quantization operation output data. The scaling coefficient for the weight w is incorporated into the parameters of other operations, similar to the scaling coefficient sf for the quantization operation output data.

[0207] As described above, the neural network generation device 300B, the neural network generation method, and the neural network generation program of this embodiment make it possible to generate a neural network execution model 100 and a neural network hardware model 400 that can be incorporated into embedded devices such as IoT devices and can operate with high performance.

[0208] Although the second embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and design modifications and the like that do not depart from the gist of the present invention are also included. Also, the components shown in the above embodiment and modified examples can be appropriately combined to form a configuration.

[0209] (Variation 1) In the above embodiment, the first memory 1 and the second memory 2 are separate memories, but the aspects of the first memory 1 and the second memory 2 are not limited to this. The first memory 1 and the second memory 2 may be, for example, a first memory area and a second memory area in the same memory.

[0210] (Variation 2) For example, the data input to the NN execution model 100 described in the above embodiment is not limited to a single format, and may be composed of still images, moving images, voice, characters, numerical values, or a combination of these. The data input to the NN execution model 100 is not limited to the measurement results of physical quantity measuring instruments such as optical sensors, thermometers, Global Positioning System (GPS) measuring instruments, angular velocity measuring instruments, and anemometers that may be mounted on the edge device on which the neural network hardware model 400 is provided. Different information such as base station information received from peripheral devices via wired or wireless communication, information on vehicles and ships, weather information, and congestion information, as well as peripheral information, financial information, and personal information may be combined.

[0211] (Variation 3) The edge device in which the NN execution model 100 is installed is assumed to be a communication device such as a battery-powered mobile phone, a smart device such as a personal computer, a mobile device such as a digital camera, a game device, or a robot product, but is not limited thereto. It can also be used in products that have a high demand for peak power limit that can be supplied by Power on Ethernet (PoE), reduction of product heat generation, or long-term operation, to obtain effects not seen in other prior art examples. For example, by applying the model to an in-vehicle camera mounted on a vehicle or ship, or a surveillance camera installed in a public facility or on the road, it is possible to realize long-term shooting, and also contribute to weight reduction and high durability. In addition, the model can be applied to display devices such as televisions and displays, medical equipment such as medical cameras and surgical robots, and work robots used at manufacturing sites and construction sites to obtain similar effects.

[0212] The above-mentioned embodiment may be realized by recording the program in a computer-readable recording medium, reading the program recorded in the recording medium into a computer system, and executing the program. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into a computer system. The term "computer-readable recording medium" may also include a medium that dynamically holds a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in such a case. The above-mentioned program may be a program for realizing a part of the above-mentioned functions, or may be a program that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0213] In addition, the effects described in this specification are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that are apparent to a person skilled in the art from the description of this specification, in addition to or in place of the above effects. [Industrial Applicability]

[0214] The present invention can be applied to the generation of neural networks. [Explanation of symbols]

[0215] 300,300B Neural network generator 200 Convolutional Neural Networks (CNN) 100 Neural network execution model (NN execution model) 400 Neural Network Hardware Models 1. First Memory 2 Second Memory 3. DMA Controller (DMAC) 4. Convolution Circuit 42 Multiplier 43 Accumulator Circuit 5 Quantization operation circuit 52 Vector arithmetic circuit 53 Quantization circuit 6 Controller 61 Registers PM Learned parameters DS Training Dataset HW Hardware information NW Network information

Claims

1. A neural network generation device for generating a neural network execution model for computing a neural network, comprising: an execution model generation unit that generates the neural network execution model based on hardware information of hardware on which the neural network execution model runs and network information of the neural network; A learning unit that generates trained parameters of the generated neural network execution model; Equipped with the learning unit, when learning the parameters in the neural network execution model, performs learning using a number of bits with higher accuracy than that of the neural network execution model based on an error of an inference result of the neural network execution model configured based on hardware information of the hardware, and determines the parameters in the neural network execution model. Neural network generator.

2. The parameter in the neural network execution model is a convolution operation parameter in the neural network execution model.

2. The neural network generating device according to claim 1.

3. The parameter in the neural network execution model is a coefficient of an activation function in the neural network execution model.

2. The neural network generating device according to claim 1.

4. The parameters in the neural network execution model are determined by reducing the number of bits of the learning result by the learning unit using a lookup table.

4. The neural network generating device according to claim 2 or 3.

5. Further comprising a hardware generation unit that generates a neural network hardware model based on the hardware information and the neural network execution model.

2. The neural network generating device according to claim 1.

Citation Information

Patent Citations

  • Processing method and device, operation method and device

    EP3667569A1

  • Quantized neural network training and inference

    US20170286830A1

  • Hardware agnostic deep neural network compiler

    US20190392296A1

  • Information processing method, information processing device and program

    JP2018077829A