Neural network circuit and neural network calculation method

The neural network circuit and operation method address the challenge of implementing high-performance convolutional neural networks in embedded devices by using a convolution operation circuit with a multiplier and accumulator, achieving efficient operations despite limited resources.

JP2025079905APending Publication Date: 2025-05-23MAXELL LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023192769
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

It is challenging to implement high-performance convolutional neural networks in embedded devices like IoT devices due to limited hardware resources and the difficulty of incorporating large-scale dedicated circuits.

Method used

A neural network circuit and operation method that includes a convolution operation circuit with a multiplier and an accumulator circuit, capable of performing convolution operations on third- or higher-order tensors by dividing input data and accumulating results efficiently.

Benefits of technology

The proposed solution enables high-performance neural network operations in embedded devices, overcoming resource limitations and achieving efficient computing performance for convolutional neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079905000001_ABST
    Figure 2025079905000001_ABST
Patent Text Reader

Abstract

To provide a high-performance neural network circuit and a high-performance neural network calculation method capable of being incorporated into embedded equipment such as IoT equipment.SOLUTION: A neural network circuit comprises a convolution operation circuit to perform the convolution operation on input data and weight. The input data is the third order or higher tensor with the components in an x-axis direction, a y-axis direction, and a c-axis direction. The convolution operation circuit has a multiplier to multiply the divided input data obtained by dividing the input data into a prescribed number of elements in the c-axis direction and the corresponding weight, and has an accumulator circuit to add the multiplication results of the multiplier. The accumulator circuit has a register to record the addition result obtained by accumulating the multiplication results in the continuous divided input data in the c-axis direction, and outputs the addition result of the continuous divided input data in the c-axis direction, recorded in the register.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a neural network circuit and a neural network operation method. [Background technology]

[0002] In recent years, convolutional neural networks (CNNs) have been used as models for image recognition and the like. Convolutional neural networks have a multi-layer structure with convolutional layers and pooling layers, and require a large number of calculations, such as convolutional calculations. Various calculation methods have been devised to speed up calculations by convolutional neural networks (Patent Document 1, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2018-077829 A Summary of the Invention [Problem to be solved by the invention]

[0004] On the other hand, it is desired to realize image recognition using a convolutional neural network in embedded devices such as IoT devices. It is difficult to incorporate a large-scale dedicated circuit such as that described in Patent Document 1 into an embedded device. Also, in an embedded device with limited hardware resources such as a CPU and memory, it is difficult to realize sufficient computing performance of a convolutional neural network by software alone.

[0005] In consideration of the above circumstances, an object of the present invention is to provide a high-performance neural network circuit and neural network operation method that can be incorporated into embedded devices such as IoT devices. [Means for solving the problem]

[0006] In order to solve the above problems, the present invention proposes the following means. A neural network circuit according to a first aspect of the present invention includes a convolution operation circuit that performs a convolution operation on input data and weights, the input data being a third- or higher-order tensor having elements in the x-axis, y-axis, and c-axis directions, the convolution operation circuit including a multiplier that multiplies the weights by divided input data obtained by dividing the input data by a predetermined number of elements in the c-axis direction, and an accumulator circuit that adds up the multiplication operation results of the multipliers, the accumulator circuit including a register that records an addition operation result obtained by accumulating the multiplication operation results of the divided input data that are continuous in the c-axis direction, and outputs the addition operation result of the divided input data that are continuous in the c-axis direction recorded in the register.

[0007] A neural network computation method according to a second aspect of the present invention is a computation method for performing a convolution operation on input data, which is a third- or higher-order tensor having elements in the x-axis, y-axis, and c-axis directions, and a weight, and multiplies the input data by divided input data obtained by dividing the input data by a predetermined number of elements in the c-axis direction by the corresponding weight, accumulates the results of the multiplication operation on the divided input data that are continuous in the c-axis direction using a dedicated register, and outputs the results of the addition operation on the divided input data that are continuous in the c-axis direction recorded in the register. Effect of the Invention

[0008] The neural network circuit and neural network operation method of the present invention can be incorporated into embedded devices such as IoT devices and have high performance. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 illustrates a convolutional neural network. [Diagram 2] FIG. 2 is a diagram for explaining a convolution operation performed by a convolution layer. [Diagram 3]FIG. 13 is a diagram for explaining data expansion in a convolution operation. [Figure 4] 1 is a diagram showing an overall configuration of a neural network circuit according to a first embodiment; [Diagram 5] FIG. 2 is a diagram showing the overall configuration of an NN processing core. [Figure 6] 10 is a timing chart showing an example of the operation of the NN processing core. [Figure 7] 13 is a timing chart showing another example of the operation of the same NN processing core. [Figure 8] FIG. 1 is a diagram illustrating a NN computing multi-core. [Figure 9] FIG. 2 is an internal block diagram of the DMAC of the neural network circuit. [Figure 10] FIG. 2 is a state transition diagram of a control circuit of the DMAC. [Figure 11] FIG. 2 is an internal block diagram of a convolution operation circuit of the neural network circuit. [Figure 12] FIG. 2 is an internal block diagram of a multiplier in the convolution operation circuit. [Figure 13] FIG. 2 is an internal block diagram of a multiply-and-accumulate unit of the multiplier. [Figure 14] FIG. 2 is an internal block diagram of an accumulator circuit of the convolution operation circuit. [Figure 15] FIG. 2 is an internal block diagram of an accumulator unit of the accumulator circuit. [Figure 16] 13 is a flowchart showing an example of a convolution operation. [Figure 17] FIG. 2 is a diagram illustrating input vectors A[0] to A[n-1]. [Figure 18] FIG. 2 is an internal block diagram of a quantization calculation circuit of the neural network circuit. [Figure 19] FIG. 2 is an internal block diagram of a vector operation circuit and a quantization circuit of the quantization operation circuit. [Figure 20] FIG. 2 is a block diagram of a computing unit. [Figure 21] FIG. 2 is an internal block diagram of a vector quantization unit of the quantization circuit. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] First embodiment A first embodiment of the present invention will be described with reference to FIGS. 1 is a diagram showing a convolutional neural network 200 (hereinafter, referred to as "CNN 200"). The calculations performed by the neural network circuit 100 (hereinafter, referred to as "NN circuit 100") according to the first embodiment are at least a part of the trained CNN 200 used during inference.

[0011] [CNN200] The CNN 200 is a network with a multi-layer structure including a convolution layer 210 that performs a convolution operation, a quantization operation layer 220 that performs a quantization operation, and an output layer 230. In at least a part of the CNN 200, the convolution layer 210 and the quantization operation layer 220 are alternately connected. The CNN 200 is a model that is widely used for image recognition and video recognition. The CNN 200 may further include a layer having other functions, such as a fully connected layer.

[0012] FIG. 2 is a diagram illustrating the convolution operation performed by the convolution layer 210. The convolution layer 210 performs a convolution operation on the input data a using a weight w. The convolution layer 210 performs a multiply-and-accumulate operation on the input data a and the weight w.

[0013] The input data a (also called activation data or feature map) to the convolutional layer 210 is multidimensional data such as image data. In this embodiment, the input data a is a three-dimensional tensor consisting of elements (x, y, c). The convolutional layer 210 of the CNN 200 performs a convolution operation on the low-bit input data a. In this embodiment, the elements of the input data a are 2-bit unsigned integers (0, 1, 2, 3). The elements of the input data a may be, for example, 4-bit or 8-bit unsigned integers.

[0014] If the input data input to the CNN 200 has a different format from the input data a to the convolutional layer 210, such as a 32-bit floating-point type, the CNN 200 may further have an input layer before the convolutional layer 210 that performs type conversion and quantization.

[0015] The weight w (also called a filter or kernel) of the convolutional layer 210 is multidimensional data having elements that are learnable parameters. In this embodiment, the weight w is a four-dimensional tensor consisting of elements (i, j, c, d). The weight w has d three-dimensional tensors (hereinafter referred to as "weights wo") consisting of elements (i, j, c). The weight w in the trained CNN 200 is trained data. The convolutional layer 210 of the CNN 200 performs a convolution operation using a low-bit weight w. In this embodiment, the element of the weight w is a 1-bit signed integer (0, 1), where a value "0" represents +1 and a value "1" represents -1.

[0016] The convolution layer 210 performs the convolution operation shown in Equation 1 and outputs output data f. In Equation 1, s indicates the stride. The area indicated by the dotted line in Fig. 2 indicates one of the areas ao (hereinafter referred to as "application area ao") where the weight wo is applied to the input data a. The elements of the application area ao are represented by (x+i, y+j, c).

[0017]

number

[0018] The quantization operation layer 220 performs quantization and the like on the output of the convolution operation output by the convolution layer 210. The quantization operation layer 220 includes a pooling layer 221, a batch normalization layer 222, an activation function layer 223, and a quantization layer 224.

[0019] The pooling layer 221 compresses the output data f of the convolutional operation output by the convolutional layer 210 by performing calculations such as average pooling (Equation 2) and MAX pooling (Equation 3) on the output data f of the convolutional operation output by the convolutional layer 210. In Equations 2 and 3, u indicates an input tensor, v indicates an output tensor, and T indicates the size of the pooling region. In Equation 3, max is a function that outputs the maximum value of u for the combination of i and j included in T.

[0020]

number

[0021]

number

[0022] The batch normalization layer 222 normalizes the data distribution of the output data of the quantization operation layer 220 and the pooling layer 221, for example, by the operation shown in Equation 4. In Equation 4, u represents an input tensor, v represents an output tensor, α represents a scale, and β represents a bias. In the trained CNN 200, α and β are trained constant vectors.

[0023]

number

[0024] The activation function layer 223 performs an activation function operation such as ReLU (Equation 5) on the output of the quantization operation layer 220, the pooling layer 221, and the batch normalization layer 222. In Equation 5, u is an input tensor and v is an output tensor. In Equation 5, max is a function that outputs the largest numerical value among arguments.

[0025]

number

[0026] The quantization layer 224 performs quantization on the output of the pooling layer 221 and the activation function layer 223 based on the quantization parameter, for example, as shown in Equation 6. The quantization shown in Equation 6 reduces the input tensor u to 2 bits. In Equation 6, q(c) is a vector of quantization parameters. In the trained CNN 200, q(c) is a trained constant vector. The inequality sign "≦" in Equation 6 may be "<".

[0027]

number

[0028] The output layer 230 is a layer that outputs the result of the CNN 200 using an identity function, a softmax function, etc. The layer preceding the output layer 230 may be the convolution layer 210 or the quantization operation layer 220.

[0029] In the CNN 200, the quantized output data of the quantization layer 224 is input to the convolution layer 210, so the load of the convolution calculation in the convolution layer 210 is smaller than that of other convolution neural networks that do not perform quantization.

[0030] [Split convolution operations] The NN circuit 100 divides the input data of the convolution operation (Equation 1) of the convolution layer 210 into partial tensors and performs the operation. The method of division into the partial tensors and the number of divisions are not particularly limited. A partial tensor is formed, for example, by dividing the input data a(x+i, y+j, c) into a(x+i, y+j, co). Note that the NN circuit 100 can also perform the operation without dividing the input data of the convolution operation (Equation 1) of the convolution layer 210.

[0031] In input data division for a convolution operation, the variable c in Equation 1 is divided into blocks of size Bc, as shown in Equation 7. Also, the variable d in Equation 1 is divided into blocks of size Bd, as shown in Equation 8. In Equation 7, co is an offset, and ci is an index from 0 to (Bc-1). In Equation 8, do is an offset, and di is an index from 0 to (Bd-1). Note that size Bc and size Bd may be the same.

[0032]

number

[0033]

number

[0034] The input data a(x+i, y+j, c) in Equation 1 is divided in the c-axis direction by size Bc, and is expressed as divided input data a(x+i, y+j, co). In the following description, the divided input data a is also referred to as "divided input data a".

[0035] The weight w(i,j,c,d) in Equation 1 is divided by the size Bc in the c-axis direction and the size Bd in the d-axis direction, and is expressed as the divided weight w(i,j,co,do). In the following description, the divided weight w is also referred to as the "divided weight w".

[0036] The output data f(x, y, do) divided by the size Bd is calculated by Equation 9. The divided output data f(x, y, do) are combined to calculate the final output data f(x, y, d).

[0037]

number

[0038] [Expanding data for convolution operations] The NN circuit 100 performs a convolution operation by expanding the input data a and the weights w in the convolution operation of the convolutional layer 210.

[0039] FIG. 3 is a diagram for explaining the expansion of data in the convolution operation. The divided input data a(x+i, y+j, co) is expanded into vector data having Bc elements. The elements of the divided input data a are indexed by ci (0 ≦ ci < Bc). In the following description, the divided input data a expanded into vector data for each i and j is also referred to as "input vector A". The input vector A has elements from the divided input data a(x+i, y+j, co×Bc) to the divided input data a(x+i, y+j, co×Bc+(Bc-1)).

[0040] The divided weight w(i,j,co,do) is expanded into matrix data having Bc×Bd elements. The elements of the divided weight w expanded into matrix data are indexed by ci and di (0 ≦ di < Bd). In the following description, the divided weight w expanded into matrix data for each i and j is also referred to as "weight matrix W". The weight matrix W has elements from the divided weight w(i,j,co×Bc,do×Bd) to the divided weight w(i,j,co×Bc+(Bc-1),do×Bd+(Bd-1)).

[0041] By multiplying the input vector A and the weight matrix W, vector data is calculated. By shaping the vector data calculated for each i, j, and co into a three-dimensional tensor, the output data f(x, y, do) can be obtained. By performing such data expansion, the convolution operation of the convolutional layer 210 can be implemented by multiplying vector data and matrix data.

[0042] [NN circuit 100] FIG. 4 is a diagram showing the overall configuration of the NN circuit 100 according to the present embodiment. The NN circuit 100 includes a DMA controller 3 (hereinafter also referred to as "DMAC 3"), a controller 6, an IFU 7, and at least one neural network processing core 10 (hereinafter also referred to as "NN processing core 10").

[0043] The NN circuit 100 can implement multiple NN processing cores 10. The NN circuit 100 illustrated in FIG. 4 can implement up to four NN processing cores 10. The multiple NN processing cores 10 constitute a "neural network processing multi-core 10M (hereinafter also referred to as "NN processing multi-core 10M")" that cooperates to execute at least some of the processing of the NN 200. In this embodiment, the multiple NN processing cores 10 are daisy-chained. Note that the number of NN processing cores 10 that can be implemented in the NN circuit 100 may be five or more.

[0044] The DMAC 3 is connected to the external bus EB, and transfers data between an external memory 120 such as a DRAM and the NN processing core 10. The DMAC 3 transfers data read from the external memory 120 to any one of the multiple NN processing cores 10. Note that the DMAC 3 may be capable of transferring the same data read from the external memory 120 to the multiple NN processing cores 10, or may be capable of broadcasting the data.

[0045] The controller 6 is connected to an external bus EB, and operates as a slave to the external host CPU 110. The controller 6 includes a bus bridge 60 and a register 61.

[0046] The bus bridge 60 relays bus access from the external bus EB to the internal bus IB, and also relays write and read requests from the external host CPU 110 to the register 61.

[0047] Register 61 has parameter registers and status registers. The parameter registers are registers that control the operation of the NN circuit 100. The status register includes pointers and instruction counts of instruction sequences of each module, and is a register indicating the state of the NN circuit 100. Also, the status register may be configured to include the semaphore S. The external host CPU 110 can access the register 61 via the bus bridge 60 of the controller 6.

[0048] The controller 6 is connected to each block (DMAC 3, IFU 7, NN arithmetic core 10) of the NN circuit 100 via the internal bus IB. The external host CPU 110 can access each block of the NN circuit 100 via the controller 6. For example, the external host CPU 110 can instruct commands for the NN arithmetic core 10 via the controller 6. Also, each block can update the status register (which may include the semaphore S) that the controller 6 has via the internal bus IB. The status register may be configured to be updated via dedicated wiring connected to each block.

[0049] The IFU (Instruction Fetch Unit) 7 reads instruction commands for each block (DMAC 3, NN arithmetic core 10) of the NN circuit 100 from the external memory 120 via the external bus EB based on the instructions of the external host CPU 110. Also, the IFU 7 transfers the read instruction commands to the corresponding blocks (DMAC 3, NN arithmetic core 10) of the NN circuit 100. In the present embodiment, the instruction commands are stored in the external memory 120 in a compressed state (hereinafter, also referred to as "compressed instruction commands"). The IFU 7 reads the compressed instruction commands.

[0050] [NN arithmetic core 10] Figure 5 is a diagram showing the overall configuration of the NN arithmetic core 10. The NN calculation core 10 includes a first memory 1, a second memory 2, a convolution calculation circuit 4, and a quantization calculation circuit 5. The NN calculation core 10 is characterized in that the convolution calculation circuit 4 and the quantization calculation circuit 5 are formed in a loop shape via the first memory 1 and the second memory 2.

[0051] The first memory 1 is a rewritable memory such as a volatile memory constituted by, for example, SRAM (Static RAM). Data is written to and read from the first memory 1 via the DMAC 3 and the internal bus IB. The external host CPU 110 can input and output data to and from the NN calculation core 10 by writing and reading data to and from the first memory 1.

[0052] The first memory 1 is connected to an input port of the convolution operation circuit 4, and the convolution operation circuit 4 can read data from the first memory 1. The first memory 1 is also loop-connected (C1) to an output port of the quantization operation circuit 5, and the quantization operation circuit 5 can write data to the first memory 1. The first memory 1 is also capable of data transfer via an inter-core connection (C2) between the first memory 1 and another NN operation core 10, and the other NN operation core 10 connected to the inter-core connection (C2) can write data to the first memory 1. In this embodiment, a daisy-chain connection is used as an example of the inter-core connection (C2).

[0053] The second memory 2 is a rewritable memory such as a volatile memory constituted by, for example, SRAM (Static RAM). Data is written to and read from the second memory 2 via the DMAC 3 and the internal bus IB. The external host CPU 110 can input and output data to and from the NN calculation core 10 by writing and reading data to and from the second memory 2.

[0054] The second memory 2 is connected to the input port of the quantization operation circuit 5, and the quantization operation circuit 5 can read data from the second memory 2. Also, the second memory 2 is connected to the output port of the convolution operation circuit 4, and the convolution operation circuit 4 can write data to the second memory 2.

[0055] The convolution operation circuit 4 is a circuit that performs the convolution operation in the convolution layer 210 of the learned CNN 200. The convolution operation circuit 4 reads the input data a stored in the first memory 1, and performs a convolution operation on the input data a. The convolution operation circuit 4 writes the output data f of the convolution operation (hereinafter, also referred to as "convolution operation output data") to the second memory 2.

[0056] The quantization operation circuit 5 is a circuit that performs at least part of the quantization operation in the quantization operation layer 220 of the learned CNN 200. The quantization operation circuit 5 reads the output data f of the convolution operation stored in the second memory 2, and performs a quantization operation (an operation including at least quantization among pooling, Batch Normalization, activation function, and quantization) on the output data f of the convolution operation.

[0057] The quantization operation circuit 5 writes the output data of the quantization operation (hereinafter, also referred to as "quantization operation output data") to the first memory 1 connected in a loop (C1). Also, the quantization operation circuit 5 can transfer data via the core-to-core connection (C2) with other NN operation cores 10, and the quantization operation circuit 5 can output the quantization operation output data to other NN operation cores 10 connected by the core-to-core connection (C2).

[0058] Since the NN operation core 10 has the first memory 1, the second memory 2, etc., in the data transfer by the DMAC 3 from an external memory such as a DRAM, the number of times of data transfer of duplicate data can be reduced. Thereby, the power consumption or processing load generated by memory access can be significantly reduced.

[0059] [Operation Example 1 of NN Operation Core 10] FIG. 6 is a timing chart showing an example of the operation of the NN processing core 10. In FIG. The DMAC 3 stores the input data a of the layer 1 in the first memory 1. The DMAC 3 may divide the input data a of the layer 1 and transfer it to the first memory 1 in accordance with the order of the convolution operation performed by the convolution operation circuit 4.

[0060] The convolution operation circuit 4 reads out the input data a of layer 1 stored in the first memory 1. The convolution operation circuit 4 performs the convolution operation of layer 1 shown in FIG.

[0061] The quantization calculation circuit 5 reads the output data f of layer 1 stored in the second memory 2. The quantization calculation circuit 5 performs a quantization calculation of layer 2 on the output data f of layer 1. The output data of the quantization calculation of layer 2 is stored in the first memory 1.

[0062] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2 stored in the first memory 1. The convolution operation circuit 4 performs a convolution operation of the layer 3 using the output data of the quantization operation of the layer 2 as input data a. The output data f of the convolution operation of the layer 3 is stored in the second memory 2.

[0063] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2M-2 (M is a natural number) stored in the first memory 1. The convolution operation circuit 4 performs the convolution operation of the layer 2M-1 using the output data of the quantization operation of the layer 2M-2 as input data a. The output data f of the convolution operation of the layer 2M-1 is stored in the second memory 2.

[0064] The quantization calculation circuit 5 reads the output data f of the layer 2M-1 stored in the second memory 2. The quantization calculation circuit 5 performs the quantization calculation of the layer 2M on the output data f of the 2M-1 layer. The output data of the quantization calculation of the layer 2M is stored in the first memory 1.

[0065] The convolution operation circuit 4 reads the output data of the quantization operation of the layer 2M stored in the first memory 1. The convolution operation circuit 4 performs a convolution operation of the layer 2M+1 using the output data of the quantization operation of the layer 2M as input data a. The output data f of the convolution operation of the layer 2M+1 is stored in the second memory 2.

[0066] The convolution calculation circuit 4 and the quantization calculation circuit 5 alternately perform calculations to advance the calculations of the CNN 200 shown in Fig. 1. In the NN calculation core 10, the convolution calculation circuit 4 performs the convolution calculations of layers 2M-1 and 2M+1 by time sharing. In the NN calculation core 10, the quantization calculation circuit 5 performs the quantization calculations of layers 2M-2 and 2M by time sharing. Therefore, the circuit scale of the NN calculation core 10 is significantly smaller than when a separate convolution calculation circuit 4 and quantization calculation circuit 5 are implemented for each layer.

[0067] The NN processing core 10 performs the calculations of the CNN 200, which is a multi-layer structure of multiple layers, using a circuit formed in a loop. The NN processing core 10 can efficiently use hardware resources due to the loop circuit configuration. Since the NN processing core 10 forms a circuit in a loop, the parameters in the convolution processing circuit 4 and the quantization processing circuit 5, which change in each layer, are updated appropriately.

[0068] When the calculations of the CNN 200 include calculations that cannot be performed by the NN calculation core 10, the NN calculation core 10 transfers intermediate data to an external calculation device such as an external host CPU 110. After the external calculation device performs calculations on the intermediate data, the calculation results by the external calculation device are input to the first memory 1 and / or the second memory 2. The NN calculation core 10 resumes calculations on the calculation results by the external calculation device.

[0069] [NN calculation core 10 operation example 2] FIG. 7 is a timing chart showing another example of the operation of the NN processing core 10. In FIG. The NN processing core 10 may divide the input data a into partial tensors and perform operations on the partial tensors by time division. The method of division into the partial tensors and the number of divisions are not particularly limited.

[0070] FIG. 7 shows an example of operation when the input data a is decomposed into two partial tensors. The decomposed partial tensors are called "first partial tensor a 1 ", "The second part tensor a 2 For example, the convolution operation of layer 2M-1 is the first partial tensor a 1 The convolution operation corresponding to (in FIG. 7, “Layer 2M-1 (a 1 ) and the second partial tensor a 2 The convolution operation corresponding to (in FIG. 7, “Layer 2M-1 (a 2 )" and

[0071] First part of tensor a 1 and the second partial tensor a 2 The convolution and quantization operations corresponding to can be performed independently, as shown in FIG.

[0072] The convolution circuit 4 converts the first partial tensor a 1 The convolution operation of layer 2M-1 corresponding to (in FIG. 7, layer 2M-1(a 1 Then, the convolution circuit 4 performs the operation shown by the second partial tensor a 2 The convolution operation of layer 2M-1 corresponding to (in FIG. 7, layer 2M-1(a 2 The quantization calculation circuit 5 performs the calculation indicated by the first partial tensor a 1 quantization operation of layer 2M corresponding to (in FIG. 7, layer 2M (a 1 In this way, the NN processing core 10 performs the calculation of the second partial tensor a 2 The convolution operation of layer 2M-1 corresponding to 1 The quantization operation of layer 2M corresponding to can be performed in parallel.

[0073] Next, the convolution operation circuit 4 performs a convolution operation on the layer 2M + 1 corresponding to the first partial tensor a 1 (the operation indicated by layer 2M + 1(a 1 ) in FIG. 7). Further, the quantization operation circuit 5 performs a quantization operation on the layer 2M corresponding to the second partial tensor a 2 (the operation indicated by layer 2M(a 2 ) in FIG. 7). Thus, the NN operation core 10 can perform the convolution operation of the layer 2M + 1 corresponding to the first partial tensor a 1 and the quantization operation of the layer 2M corresponding to the second partial tensor a 2 in parallel.

[0074] The convolution operation and quantization operation corresponding to the first partial tensor a 1 and the convolution operation and quantization operation corresponding to the second partial tensor a 2 can be performed independently. Therefore, the NN operation core 10 can, for example, perform the convolution operation of the layer 2M - 1 corresponding to the first partial tensor a 1 and the quantization operation of the layer 2M + 2 corresponding to the second partial tensor a 2 in parallel. That is, the convolution operation and quantization operation that the NN operation core 10 performs in parallel are not limited to the operations of consecutive layers.

[0075] By dividing the input data a into partial tensors, the NN operation core 10 can operate the convolution operation circuit 4 and the quantization operation circuit 5 in parallel. As a result, the waiting time of the convolution operation circuit 4 and the quantization operation circuit 5 is reduced, and the operation processing efficiency of the NN operation core 10 is improved. In the operation example shown in FIG. 7, the number of divisions was 2, but even when the number of divisions is larger than 2, the NN operation core 10 can operate the convolution operation circuit 4 and the quantization operation circuit 5 in parallel.

[0076] For example, if the input data a is "the first partial tensor a 1 ", "the second partial tensor a 2 ", and "the third partial tensor a 3", the NN calculation core 10 divides the second partial tensor a 2 The convolution operation of layer 2M-1 corresponding to 3 The quantization operation of the layer 2M corresponding to the input data a may be performed in parallel with the quantization operation of the layer 2M corresponding to the input data a. The order of the operations is appropriately changed depending on the storage status of the input data a in the first memory 1 and the second memory 2.

[0077] As a method of computing partial tensors, an example (method 1) was shown in which the partial tensor in the same layer is computed by the convolution computation circuit 4 or the quantization computation circuit 5, and then the partial tensor in the next layer is computed. For example, as shown in FIG. 7, in the convolution computation circuit 4, the first partial tensor a 1 and the second part tensor a 2 The convolution operation of layer 2M-1 corresponding to (in FIG. 7, layer 2M-1(a 1 ) and Layer 2M-1(a 2 After performing the operation shown in (a), the first partial tensor a 1 and the second part tensor a 2 The convolution operation of layer 2M+1 corresponding to (in FIG. 7, layer 2M+1(a 1 ) and Layer 2M+1(a 2 ) is performed.

[0078] However, the method of computing the partial tensors is not limited to this. The method of computing the partial tensors may be a method of computing the remaining partial tensors after computing some of the partial tensors in multiple layers (Method 2). For example, in the convolution computation circuit 4, the first partial tensor a 1 Layer 2M-1 and the first partial tensor a 1 After performing the convolution operation of layer 2M+1 corresponding to 2 Layer 2M-1 and the second partial tensor a 2 2M+1 corresponding to the layer 2M+1 convolution operation may be performed.

[0079] Furthermore, the method of computing partial tensors may be a method of computing partial tensors by combining method 1 and method 2. However, when using method 2, it is necessary to perform the computation in accordance with the dependency relationship regarding the computation order of the partial tensors.

[0080] [NN calculation multi-core 10M] FIG. 8 is a diagram showing the NN calculation multi-core 10M. The NN processing multi-core 10M illustrated in Fig. 8 includes two daisy-chained NN processing cores 10. When distinguishing between the two NN processing cores 10, the two NN processing cores 10 are referred to as a "first NN processing core 10A" and a "second NN processing core 10B." In Fig. 8, the first memory 1 is abbreviated as "A," the convolution processing circuit 4 as "C," the second memory 2 as "F," and the quantization processing circuit 5 as "Q."

[0081] Specifically, the quantization calculation circuit 5 of the first NN arithmetic core 10A and the first memory 1 of the second NN arithmetic core 10B are daisy-chain connected (C2). The quantization calculation circuit 5 of the first NN arithmetic core 10A can write quantization calculation output data to the first memory 1 of the first NN arithmetic core 10A connected in a loop (C1) and / or the first memory 1 of the second NN arithmetic core 10B connected in a daisy chain (C2).

[0082] Specifically, the quantization calculation circuit 5 of the second NN calculation core 10B and the first memory 1 of the first NN calculation core 10A are daisy-chain connected (C2). The quantization calculation circuit 5 of the second NN calculation core 10B can write the quantization calculation output data to the first memory 1 of the second NN calculation core 10B connected in a loop (C1) and / or the first memory 1 of the first NN calculation core 10A connected in a daisy chain (C2).

[0083] Similarly, when the NN processing multi-core 10M has three or more NN processing cores 10, the multiple NN processing cores 10 are connected in a daisy chain. The quantization processing circuit 5 of the NN processing cores 10 other than the final stage NN processing core 10 is connected in a daisy chain (C2) with the first memory 1 of the subsequent stage NN processing core 10B. The quantization processing circuit 5 of the final stage NN processing core 10 is connected in a daisy chain (C2) with the first memory 1 of the first stage NN processing core 10. The multiple NN processing cores 10 are characterized by being formed in a daisy chain loop (linked together).

[0084] In one NN calculation core 10, the first memory (A) 1, the convolution calculation circuit (C) 4, the second memory (F) 2, and the quantization calculation circuit (Q) 5 are connected in a loop. On the other hand, in the NN calculation multi-core 10M, the first memory (A) 1, the convolution calculation circuit (C) 4, the second memory (F) 2, and the quantization calculation circuit (Q) 5 are connected in a daisy chain loop (linked together) so that the first memory (A) 1, the convolution calculation circuit (C) 4, the second memory (F) 2, and the quantization calculation circuit (Q) 5 are repeatedly arranged in the same order.

[0085] The multiple NN calculation cores 10 constituting the NN calculation multi-core 10M do not need to have the same hardware configuration. For example, the capacity and configuration of the first memory 1 of the first NN calculation core 10A may be different from the capacity and configuration of the first memory 1 of the second NN calculation core 10B. For example, the configuration of the quantization calculation circuit 5 of the first NN calculation core 10A may be different from the configuration of the quantization calculation circuit 5 of the second NN calculation core 10B.

[0086] Next, each component of the NN circuit 100 will be described in detail.

[0087] [DMAC3] FIG. 9 is an internal block diagram of the DMAC3. The DMAC 3 has a data transfer circuit 31 and a state controller 32. The DMAC 3 has a state controller 32 dedicated to the data transfer circuit 31, and when an instruction command is input, the DMAC 3 can perform DMA data transfer without requiring an external controller.

[0088] The data transfer circuit 31 is connected to the external bus EB, and performs DMA data transfer between an external memory 120 such as a DRAM and the NN processing core 10. The number of DMA channels of the data transfer circuit 31 is not limited. For example, the first NN processing core 10A and the second NN processing core 10B may each have a dedicated DMA channel.

[0089] The state controller 32 controls the state of the data transfer circuit 31. The state controller 32 is also connected to the controller 6 via an internal bus IB. The state controller 32 has an instruction queue 33 and a control circuit .

[0090] The instruction queue 33 is a queue that stores instruction commands C3 for the DMAC 3, and is configured, for example, by a FIFO memory. One or more instruction commands C3 are written to the instruction queue 33 via the IFU 7 or the internal bus IB.

[0091] The control circuit 34 is a state machine that decodes the instruction command C3 and sequentially controls the data transfer circuit 31 based on the instruction command C3. The control circuit 34 may be implemented by a logic circuit or a CPU controlled by software.

[0092] FIG. 10 is a state transition diagram of the control circuit 34. When an instruction command C3 is input to the instruction queue 33 (Not empty), the control circuit 34 transitions from the idle state ST1 to the decode state ST2.

[0093] In the decode state ST2, the control circuit 34 decodes the instruction command C3 output from the instruction queue 33. The control circuit 34 also reads the semaphore S stored in the register 61 of the controller 6, and determines whether the operation of the data transfer circuit 31 instructed in the instruction command C3 is executable. If it is not executable (Not ready), the control circuit 34 waits until it is executable (Wait). If it is executable (Ready), the control circuit 34 transitions from the decode state ST2 to the execution state ST3.

[0094] In the execution state ST3, the control circuit 34 controls the data transfer circuit 31 to cause the data transfer circuit 31 to perform the operation instructed in the instruction command C3. When the operation of the data transfer circuit 31 is completed, the control circuit 34 removes the executed instruction command C3 from the instruction queue 33 and updates the semaphore S stored in the register 61 of the controller 6. If there are instructions in the instruction queue 33 (Not empty), the control circuit 34 transitions from the execution state ST3 to the decode state ST2. If there are no instructions in the instruction queue 33 (Empty), the control circuit 34 transitions from the execution state ST3 to the idle state ST1.

[0095] [Convolution circuit 4] FIG. 11 is an internal block diagram of the convolution operation circuit 4. As shown in FIG. The convolution operation circuit 4 has a weight memory 41, a multiplier 42, an accumulator circuit 43, and a state controller 44. The convolution operation circuit 4 has a state controller 44 dedicated to the multiplier 42 and the accumulator circuit 43, and when an instruction command is input, the convolution operation can be performed without requiring an external controller.

[0096] The weight memory 41 is a memory in which the weight w used in the convolution calculation is stored, and is a rewritable memory such as a volatile memory constituted by, for example, an SRAM (Static RAM), etc. The DMAC 3 writes the weight w required for the convolution calculation into the weight memory 41 by DMA transfer.

[0097] FIG. 12 is an internal block diagram of the multiplier 42. The multiplier 42 multiplies the input vector A and the weight matrix W. The input vector A is vector data having Bc elements obtained by expanding the divided input data a(x + i, y + j, co) for each i and j as described above. The weight matrix W is matrix data having Bc × Bd elements obtained by expanding the divided weights w(i, j, co, do) for each i and j. The multiplier 42 includes Bc × Bd product-sum operation units 47 and can perform the multiplication of the input vector A and the weight matrix W in parallel.

[0098] The multiplier 42 reads the input vector A and the weight matrix W required for multiplication from the first memory 1 and the weight memory 41 and performs the multiplication. The multiplier 42 outputs Bd product-sum operation results O(di).

[0099] FIG. 13 is an internal block diagram of the product-sum operation unit 47. The product-sum operation unit 47 performs the multiplication of the element A(ci) of the input vector A and the element W(ci, di) of the weight matrix W. The product-sum operation unit 47 adds the multiplication result and the multiplication result S(ci, di) of another product-sum operation unit 47. The product-sum operation unit 47 outputs the addition result S(ci + 1, di). The element A(ci) is a 2-bit unsigned integer (0, 1, 2, 3). The element W(ci, di) is a 1-bit signed integer (0, 1), where the value "0" represents +1 and the value "1" represents -1.

[0100] The multiply-and-accumulate unit 47 has an inverter 47a, a selector 47b, and an adder 47c. The multiply-and-accumulate unit 47 performs multiplication using only the inverter 47a and the selector 47b, without using a multiplier. When the element W(ci,di) is "0", the selector 47b selects the input of the element A(ci). When the element W(ci,di) is "1", the selector 47b selects the complement of the element A(ci) inverted by the inverter. The element W(ci,di) is also input to the carry-in of the adder 47c. When the element W(ci,di) is "0", the adder 47c outputs a value obtained by adding the element A(ci) to S(ci,di). When W(ci,di) is "1", the adder 47c outputs a value obtained by subtracting the element A(ci) from S(ci,di).

[0101] FIG. 14 is an internal block diagram of the accumulator circuit 43. The accumulator circuit 43 accumulates the product-sum operation results O(di) of the multiplier 42 in the second memory 2. The accumulator circuit 43 has Bd accumulator units 48 and can accumulate the Bd product-sum operation results O(di) in parallel in the second memory 2.

[0102] FIG. 15 is an internal block diagram of the accumulator unit 48. The accumulator unit 48 has an adder 48a, a mask section 48b, a register 48c, and a selector 48d. The adder 48a adds an element O(di) of the sum-of-products operation result O and a partial sum which is an intermediate result of the convolution operation shown in Equation 1 and is stored in the register 48c or the second memory 2. The addition result by the accumulator unit 48 is 16 bits per element. The addition result is not limited to 16 bits per element, and may be, for example, 15 bits or 17 bits per element.

[0103] The adder 48a writes the addition result to the register 48c or the second memory 2. When the initialization signal clear is asserted, the masking unit 48b masks the output from the second memory 2 and sets the addition target for the element O(di) to zero.

[0104] The register 48c is a memory that temporarily stores the partial sums in progress. The register 48c records the result of an addition operation obtained by accumulating the results of the multiplication operations of the multipliers 42 that are successive in the c-axis direction. The register 48c is a small-scale recording medium formed of a register file or the like.

[0105] The selector 48d is a selector that selects data to be input to the adder 48a from data temporarily stored in the register 48c and data read from the second memory 2. When an intermediate partial sum is stored in the register 48c, the selector 48d selects and outputs the data temporarily stored in the register 48c. When an intermediate partial sum is not stored in the register 48c but is stored in the second memory 2, the selector 48d selects and outputs the data read from the second memory 2.

[0106] The state controller 44 controls the states of the multiplier 42 and the accumulator circuit 43. The state controller 44 is also connected to the controller 6 via an internal bus IB. The state controller 44 has an instruction queue 45 and a control circuit 46.

[0107] The instruction queue 45 is a queue that stores the instruction commands C4 for the convolution operation circuit 4, and is configured by, for example, a FIFO memory. The instruction commands C4 are written to the instruction queue 45 via the IFU 7 or the internal bus IB.

[0108] The control circuit 46 is a state machine that decodes the instruction command C4 and controls the multiplier 42 and the accumulator circuit 43 based on the instruction command C4. The control circuit 46 has a similar configuration to the control circuit 34 of the state controller 32 of the DMAC 3.

[0109] FIG. 16 is a flowchart showing an example of a convolution operation performed by the control circuit 46 or the like. In step S110, the DMAC 3 stores input vectors A[0] to A[n−1], which are divided input data in the c-axis direction, in the first memory 1. The variable i (0 ≦ i < n) and the register 48c are initialized to zero.

[0110] FIG. 17 is a diagram for explaining the input vectors A[0] to A[n−1]. The input vector A[i] is vector data having Bc elements in the c-axis direction. For example, the x-axis direction and the y-axis direction correspond to the screen direction in the image, and the c-axis direction corresponds to the channel direction. The number n of the input vectors A[0] to A[n−1] is desirably C / Bc from the viewpoint of calculation efficiency. However, the input vectors A[0] to A[n−1] must be storable in the first memory 1. That is, within the range storable in the first memory 1, the number n of the divided input data to be stored in the first memory 1 is selected.

[0111] For example, if the number C of elements in the c-axis direction of the input data a is 1024 and Bc is 256, the number of the input vectors A[0] to A[n−1] is “4”. That is, it is desirable that the number C of elements in the c-axis direction of the input data a held in the first memory 1 is larger than the number of elements in the c-axis direction of the input vector A[i] inputtable to the multiplier 42.

[0112] Note that the weight matrix W used for the convolution operation with A[0] to A[n−1] includes different elements in the c-axis direction. Therefore, it is desirable that the weight matrix W is also divided so as to correspond to the input vector A[i]. The product-sum operation unit 47 reads the weight matrix W[i] corresponding to the input vector A[i] from the weight memory 41 and executes a product-sum operation.

[0113] In step S120, the multiplier 42 outputs Bd product-sum operation results O(di) using the input vector A[i] as an input.

[0114] In step S130, the accumulator circuit 43 adds the product-sum operation result O(di) and the intermediate partial sum temporarily stored in the register 48c.

[0115] In step S140, the accumulator circuit 43 checks the value of the variable i. If the variable i is not n-1, the accumulator circuit 43 increments the variable i and then executes step S150. If the variable i is n-1, the accumulator circuit 43 then executes step S160.

[0116] In step S150, the accumulator circuit 43 overwrites the intermediate partial sum in the register 48c with the addition result. The accumulator circuit 43 then executes step S120 again.

[0117] The accumulator circuit 43 writes the addition result into the second memory 2 in step S160.

[0118] When the convolution operation by the multiplier 42 and the accumulator circuit 43 is completed, the output data f(x, y, do) is stored in the second memory 2. By sufficiently increasing the number of elements C in the c-axis direction of the input data a stored in the first memory 1, the output data f(x, y, do) can be calculated without reading out the partial sums of the intermediate steps from the second memory 2, and the frequency of access to the second memory 2 can be reduced. The second memory 2 outputs the output data f to the quantization operation circuit 5 while receiving multi-bit data such as 16 bits from the accumulator circuit 43. That is, the second memory 2 frequently inputs and outputs multi-bit data compared to the first memory 1. Therefore, by reducing the frequency of access to the second memory 2 as described above, the power consumption and circuit size of the second memory 2 and the peripheral circuits can be suitably reduced.

[0119] [Quantization operation circuit 5] FIG. 18 is an internal block diagram of the quantization calculation circuit 5. The quantization operation circuit 5 includes a quantization parameter memory 51, a vector operation circuit 52, a quantization circuit 53, and a state controller 54. The quantization operation circuit 5 has a dedicated state controller 54 for the vector operation circuit 52 and the quantization circuit 53, and can perform quantization operations without requiring an external controller when an instruction command is input.

[0120] The quantization parameter memory 51 is a memory that stores the quantization parameter q used for quantization operations, and is a rewritable memory such as a volatile memory composed of, for example, SRAM (Static RAM). The DMAC 3 writes the quantization parameter q required for quantization operations to the quantization parameter memory 51 by DMA transfer.

[0121] FIG. 19 is an internal block diagram of the vector operation circuit 52 and the quantization circuit 53. The vector operation circuit 52 performs operations on the output data f(x, y, do) stored in the second memory 2. The vector operation circuit 52 has Bd operation units 57 and performs SIMD operations on the output data f(x, y, do) in parallel.

[0122] FIG. 20 is a block diagram of the operation unit 57. The operation unit 57 includes, for example, an ALU 57a, a first selector 57b, a second selector 57c, a register 57d, and a shifter 57e. The operation unit 57 may further include other arithmetic units and the like that a known general-purpose SIMD operation circuit has.

[0123] The vector operation circuit 52 combines the arithmetic units and the like of the operation unit 57 to perform at least one of the operations of the pooling layer 221, the Batch Normalization layer 222, and the activation function layer 223 in the quantization operation layer 220 on the output data f(x, y, do).

[0124] The arithmetic unit 57 can add the data stored in the register 57d and the element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The arithmetic unit 57 can store the addition result by the ALU 57a in the register 57d. The arithmetic unit 57 can initialize the addition result by inputting "0" to the ALU 57a instead of the data stored in the register 57d by the selection of the first selector 57b. For example, when the pooling area is 2×2, the shifter 57e can output the average value of the addition result by shifting the output of the ALU 57a to the right by 2 bits. The vector arithmetic circuit 52 can perform the average pooling calculation shown in Equation 2 by repeating the above calculations by the Bd arithmetic units 57.

[0125] The arithmetic unit 57 can compare the data stored in the register 57d with the element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The arithmetic unit 57 controls the second selector 57c according to the comparison result by the ALU 57a, and can select the larger of the data stored in the register 57d and the element f(di). The arithmetic unit 57 can initialize the comparison target to the minimum value by inputting the minimum value of the possible values ​​of the element f(di) to the ALU 57a by the selection of the first selector 57b. In this embodiment, the element f(di) is a 16-bit signed integer, so the minimum value of the possible values ​​of the element f(di) is "0x8000". The vector arithmetic circuit 52 can perform the MAX pooling calculation of Equation 3 by repeating the above calculations by the Bd arithmetic units 57. Note that in the MAX pooling calculation, the shifter 57e does not shift the output of the second selector 57c.

[0126] The arithmetic unit 57 can subtract data stored in the register 57d and an element f(di) of the output data f(x, y, do) read from the second memory 2 by the ALU 57a. The shifter 57e can shift the output of the ALU 57a to the left (i.e., multiplication) or to the right (i.e., division). The vector arithmetic circuit 52 can perform the batch normalization calculation of Equation 4 by repeating the above calculations by the Bd arithmetic units 57.

[0127] The arithmetic unit 57 can compare the element f(di) of the output data f(x, y, do) read from the second memory 2 with "0" selected by the first selector 57b by the ALU 57a. The arithmetic unit 57 can select and output either the element f(di) or a constant value "0" previously stored in the register 57d according to the comparison result by the ALU 57a. The vector arithmetic circuit 52 can perform the ReLU arithmetic of Equation 5 by repeating the above arithmetic operations by the Bd arithmetic units 57.

[0128] The vector operation circuit 52 can perform average pooling, MAX pooling, batch normalization, activation function operations, and combinations of these operations. Since the vector operation circuit 52 can perform general-purpose SIMD operations, it may perform other operations necessary for the operations in the quantization operation layer 220. In addition, the vector operation circuit 52 may perform operations other than those in the quantization operation layer 220.

[0129] It is to be noted that the quantization calculation circuit 5 does not have to include the vector calculation circuit 52. When the quantization calculation circuit 5 does not include the vector calculation circuit 52, the output data f(x, y, do) is input to the quantization circuit 53.

[0130] The quantization circuit 53 quantizes the output data of the vector operation circuit 52. The quantization circuit 53 has Bd quantization units 58, as shown in FIG.

[0131] FIG. 21 is an internal block diagram of the quantization unit 58. The quantization unit 58 quantizes the elements in(di) of the output data of the vector operation circuit 52. The quantization unit 58 includes a comparator 58a and an encoder 58b. The quantization unit 58 performs the operation (Equation 6) of the quantization layer 224 in the quantization operation layer 220 on the output data (16 bits / element) of the vector operation circuit 52. The quantization unit 58 reads the necessary quantization parameters q(th0, th1, th2) from the quantization parameter memory 51, and the comparator 58a compares the input in(di) with the quantization parameters q. The quantization unit 58 quantizes the comparison result by the comparator 58a to 2 bits / element by the encoder 58b. Since α(c) and β(c) in Equation 4 are different parameters for each variable c, the quantization parameters q(th0, th1, th2) reflecting α(c) and β(c) are different parameters for each in(di).

[0132] The quantization unit 58 classifies the input in(di) into four regions (for example, in≦th0, th0<in≦th1, th1<in≦th2, th2<in) by comparing the input in(di) with three threshold values th0, th1, th2, and encodes and outputs the classification result to 2 bits. The quantization unit 58 can also perform operations of Batch Normalization and activation functions together with quantization by setting the quantization parameters q(th0, th1, th2).

[0133] The quantization unit 58 can perform the operation of Batch Normalization shown in Equation 4 together with quantization by setting the threshold value th0 as β(c) in Equation 4 and the threshold differences (th1−th0) and (th2−th1) as α(c) in Equation 4. By increasing (th1−th0) and (th2−th1), α(c) can be decreased. By decreasing (th1−th0) and (th2−th1), α(c) can be increased.

[0134] The quantization unit 58 can perform an activation function in conjunction with quantization of the input in(di). For example, the quantization unit 58 saturates the output value in regions where in(di) ≤ th0 and th2 < in(di). The quantization unit 58 can perform the operation of the activation function in conjunction with quantization by setting the quantization parameter q so that the output is non-linear.

[0135] The state controller 54 controls the states of the vector operation circuit 52 and the quantization circuit 53. Also, the state controller 54 is connected to the controller 6 via the internal bus IB. The state controller 54 has an instruction queue 55 and a control circuit 56.

[0136] The instruction queue 55 is a queue in which the instruction command C5 for the quantization operation circuit 5 is stored, and is configured by, for example, a FIFO memory. The instruction command C5 is written into the instruction queue 55 via the IFU7 or the internal bus IB.

[0137] The control circuit 56 is a state machine that decodes the instruction command C5 and controls the vector operation circuit 52 and the quantization circuit 53 based on the instruction command C5. The control circuit 56 has the same configuration as the control circuit 34 of the state controller 32 of the DMAC3.

[0138] The quantization operation circuit 5 writes quantization operation output data having Bd elements into the first memory 1. A preferred relationship between Bd and Bc is shown in Equation 10. In Equation 10, n is an integer.

[0139] [Number]

[0140] [Controller 6] The controller 6 transfers the instruction command transferred from the external host CPU 110 via the internal bus IB to the instruction queues of the DMAC 3, the convolution operation circuit 4, and the quantization operation circuit 5. The controller 6 may have an instruction memory that stores the instruction command for each circuit.

[0141] The controller 6 is connected to an external bus EB, and operates as a slave to the external host CPU 110. The controller 6 has registers 61 including a parameter register and a status register. The parameter register is a register that controls the operation of the NN circuit 100. The status register is a register that indicates the status of the NN circuit 100, including the semaphore S.

[0142] According to the neural network circuit 100 of this embodiment, the NN circuit 100 that can be embedded in an embedded device such as an IoT device can operate with high performance. By connecting multiple NN calculation cores 10, more neural network calculations can be performed efficiently and at high speed.

[0143] Although the first embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and design modifications and the like that do not depart from the gist of the present invention are also included. Also, the components shown in the above embodiment and modified examples can be appropriately combined to form a configuration.

[0144] (Variation 1) In the above embodiment, the input vector A[i] is 1×1 vector data in the xy plane. However, the input vector A[i] is not limited to this. The input vector A[i] may be 3×3 vector data in the xy plane.

[0145] (Variation 2) In the above embodiment, the first memory 1 and the second memory 2 are separate memories, but the aspects of the first memory 1 and the second memory 2 are not limited to this. The first memory 1 and the second memory 2 may be, for example, a first memory area and a second memory area in the same memory.

[0146] (Variation 3) For example, the data input to the NN circuit 100 described in the above embodiment is not limited to a single format, and can be composed of still images, moving images, sounds, characters, numbers, or combinations of these. The data input to the NN circuit 100 is not limited to the measurement results of physical quantity measuring instruments such as optical sensors, thermometers, Global Positioning System (GPS) measuring instruments, angular velocity measuring instruments, and anemometers that may be mounted on the edge device in which the NN circuit 100 is provided. Different information such as base station information received from peripheral devices via wired or wireless communication, information on vehicles and ships, weather information, and congestion information, as well as peripheral information, financial information, and personal information may also be combined.

[0147] (Variation 4) The edge device in which the NN circuit 100 is provided is assumed to be a communication device such as a battery-powered mobile phone, a smart device such as a personal computer, a digital camera, a game device, a robot product, and other mobile devices, but is not limited thereto. By using the NN circuit 100 in products that have a high demand for limiting the peak power that can be supplied by Power on Ethernet (PoE), reducing heat generation of the product, or long-term operation, it is possible to obtain effects not seen in other prior art. For example, by applying the NN circuit 100 to an in-vehicle camera mounted on a vehicle or ship, or a surveillance camera installed in a public facility or on the road, it is possible to realize long-term shooting, and it also contributes to weight reduction and high durability. In addition, the same effect can be obtained by applying the NN circuit 100 to a display device such as a television or a display, medical equipment such as a medical camera or a surgical robot, and a work robot used in a manufacturing site or a construction site.

[0148] (Variation 5) The NN circuit 100 may be implemented using one or more processors for part or all of the NN circuit 100. For example, the NN circuit 100 may implement part or all of the input layer or the output layer by software processing by a processor. Part of the input layer or the output layer implemented by software processing is, for example, normalization or conversion of data. Thereby, it is possible to support various forms of input formats or output formats. Note that the software executed by the processor may be configured to be rewritable using communication means or an external medium.

[0149] (Modification Example 6) The NN circuit 100 may be implemented by combining a part of the processing in the CNN 200 with a Graphics Processing Unit (GPU) or the like on the cloud. In addition to the processing performed by the edge device in which the NN circuit 100 is provided, the NN circuit 100 can perform further processing on the cloud or perform processing on the edge device in addition to the processing on the cloud, thereby realizing more complex processing with fewer resources. According to such a configuration, the NN circuit 100 can reduce the communication volume between the edge device and the cloud by processing distribution.

[0150] Also, the effects described in this specification are illustrative or exemplary only and not limiting. That is, the technology according to the present disclosure may exhibit other effects apparent to those skilled in the art from the description of this specification together with or instead of the above effects.

Industrial Applicability

[0151] The present invention can be applied to neural network operations.

Explanation of Signs

[0152] 200 Convolutional Neural Network 100 Neural Network Circuit (NN Circuit) 10 Neural Network Operation Core (NN Operation Core) 10A First Neural Network Calculation Core (First NN Calculation Core) 10B Second Neural Network Calculation Core (Second NN Calculation Core) 10M Neural network calculation multi-core (NN calculation multi-core) 1. First Memory 2 Second Memory 3. DMA Controller (DMAC) 4. Convolution Circuit 42 Multiplier 43 Accumulator Circuit 5 Quantization operation circuit 52 Vector arithmetic circuit 53 Quantization circuit 6 Controller 61 Registers 7 IFU

Claims

1. A convolution calculation circuit is provided for performing a convolution calculation on input data and weights, The input data is a third- or higher-order tensor having elements in the x-axis, y-axis, and c-axis directions; The convolution operation circuit includes: a multiplier that multiplies the weights corresponding to divided input data obtained by dividing the input data by a predetermined number of elements in the c-axis direction; an accumulator circuit for adding up the multiplication results of the multipliers; having the accumulator circuit has a register for recording an addition operation result obtained by accumulating the multiplication operation results of the divided input data that are continuous in the c-axis direction, and outputs the addition operation result of the divided input data that are continuous in the c-axis direction recorded in the register. Neural network circuit.

2. a first memory for storing the input data; the first memory stores two or more pieces of divided input data consecutive in the c-axis direction; 2. The neural network circuit of claim 1.

3. the number of the divided input data consecutive in the c-axis direction stored in the first memory is a number (C / Bc) obtained by dividing the number of elements in the c-axis direction of the input data by the predetermined number of elements.

3. The neural network circuit of claim 2.

4. a second memory for storing output data output by the accumulator circuit; a quantization operation circuit that performs a quantization operation on the output data stored in the second memory; Equipped with 2. The neural network circuit of claim 1.

5. At least one of the input data or the weights has a bit precision of 8 bits or less per element.

2. The neural network circuit of claim 1.

6. the number of elements in the c-axis direction of the input data held in the first memory is greater than the number of elements in the c-axis direction of the divided input data that can be input to the multiplier; 4. The neural network circuit according to claim 2 or 3.

7. A method for performing a convolution operation on input data, which is a third- or higher-order tensor having elements in the x-axis direction, the y-axis direction, and the c-axis direction, and a weight, comprising the steps of: multiplying the weights by divided input data obtained by dividing the input data into a predetermined number of elements in the c-axis direction; cumulatively adding the multiplication results of the divided input data successive in the c-axis direction using a dedicated register; outputting an addition operation result of the divided input data consecutive in the c-axis direction recorded in the register; Neural network calculation method.

8. storing the divided input data continuous in the c-axis direction in a first memory connected to a convolution operation circuit that performs the convolution operation; 8. The neural network computing method according to claim 7.

9. the number of the divided input data consecutive in the c-axis direction stored in the first memory is a number (C / Bc) obtained by dividing the number of elements in the c-axis direction of the input data by the predetermined number of elements.

9. The neural network computing method according to claim 8.

Citation Information

Patent Citations

  • Information processing method, information processing device and program

    JP2018077829A