Neural network computing device, method, and computing device
By storing the convolution kernel and input data in memory and using the first computing circuit to perform convolutional operations, the memory wall problem in convolutional neural network computing is solved, efficient memory calculation is achieved, and power consumption is reduced.
Patent Information
- Application Number
- CN201910441581.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-05-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-05-24
AI Technical Summary
In convolutional neural network computing, due to the huge amount of data and complex calculations, the data bus needs to frequently transmit data DRAM and neural network computing nodes, resulting in memory wall problems. Especially when the processor's operating frequency requirements are high, the data bus bandwidth is not enough to support efficient computing.
By implementing neural network computing in memory, the convolution kernel and input data are stored using multi-row and multi-column DRAM units, and the convolution operation is completed in memory through the first computing circuit, reducing the transmission of data between memory and computing nodes.
Complete neural network computing in memory reduces the need for data transmission, solves the memory wall problem, and reduces the power consumption when implementing convolutional computing.
Smart Images

Figure CN111985602B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and particularly to a neural network computing device, method, and computing device. Background Art
[0002] With the development of convolutional neural networks, image recognition, speech recognition, etc. can be performed through convolutional neural network computing. Since convolutional neural network computing is a high-parallelism and data-intensive computing, when the processor of a computing device uses the von Neumann computing architecture for convolutional neural network computing, the processor first stores the data of the convolutional neural network in the memory. Generally, the memory is a dynamic random access memory (DRAM). Then, the data bus transfers the data stored in the DRAM to the neural network computing node, and the convolutional neural network computing is completed in the neural network computing node. When the computing is completed, the data bus outputs the computing result to the DRAM. Due to the huge amount of data and complex computing of the convolutional neural network, the data bus needs to transfer the data back and forth between the DRAM and the neural network computing node. When the working frequency requirement of the processor is relatively high, the data bus bandwidth connecting the processor and the DRAM will not be sufficient to transfer the data required by the processor, and the problem of memory wall will easily occur. Summary of the Invention
[0003] Embodiments of the present invention provide a neural network computing device, method, and computing device, which can complete neural network computing in the memory and solve the problem of memory wall. The technical solution is as follows:
[0004] In a first aspect, a neural network computing device is provided. The neural network computing device includes:
[0005] A first storage array, including multiple rows and columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in the storage array is used to store convolution kernels, and at least two rows of DRAM cells in the storage array are used to store input data to be subjected to convolution calculation;
[0006] A first computing circuit, connected to the storage array, for performing multiple convolution operations on the input data and the convolution kernels. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of the at least two rows of DRAM cells.
[0007] This neural network computing device stores the convolution kernel and the input data to be subjected to convolution calculation in a first memory array including multiple rows and columns of DRAM cells. The first computing circuit can read the convolution kernel and the input data to be subjected to convolution calculation from within the first memory array and complete the convolution calculation. Thus, neural network computing is implemented in the memory, reducing the transmission of neural network data between the memory and the neural network computing nodes and solving the problem of the memory wall.
[0008] In a possible implementation manner, the first computing circuit includes:
[0009] Multiple latches, the input end of each latch is connected to a column of DRAM cells, the output end of each latch is connected to one of multiple multipliers, and each latch is used to cache a one-bit value of an element in the convolution kernel stored in the corresponding column of DRAM cells. Among them, each DRAM cell in the first row of DRAM cells is respectively used to store a one-bit value of an element in the convolution kernel;
[0010] Multiple multipliers, the first input end of each multiplier is respectively connected to a column of DRAM cells, and the second input end of each multiplier is respectively connected to the output end of a latch;
[0011] Multiple adders, the input end of each adder is respectively connected to at least two multipliers;
[0012] Among them, in one convolution calculation process, the multiple multipliers are respectively used to perform multiplication operations on the convolution kernel cached by the multiple latches and the input data stored in one row of DRAM cells in the multiple rows of DRAM cells, the multiple adders are respectively used to perform addition operations on the output results of the multiple multipliers, and the output of the multiple adders is the result of one convolution calculation.
[0013] In a possible implementation manner, this neural network computing device further includes:
[0014] A second computing circuit, the second computing circuit is connected to the first computing circuit, and the second computing circuit includes a summing circuit, which is used to sum the results of the one convolution calculation output by the multiple adders in the first computing circuit during one summing calculation process.
[0015] In a possible implementation manner, the second computing circuit is used to perform multiple summing calculations, and the second computing circuit further includes:
[0016] A buffer, connected to the summing circuit, and used to cache the calculation results of the summing circuit during multiple summing calculation processes;
[0017] Pooling circuit, the pooling circuit is connected to the buffer, and is used to perform a pooling operation on the calculation results in the multiple summation calculation processes cached in the buffer.
[0018] In a possible implementation manner, the second calculation circuit further includes:
[0019] Selection circuit, the first input end of the selection circuit is used to connect the output end of the pooling circuit, the second input end of the selection circuit is used to connect the output end of the summation circuit, and the selection circuit is used to select and output the output result of the pooling circuit or the output result of the summation circuit;
[0020] Activation circuit, connected to the output end of the selection circuit, is used to perform an activation operation on the output result of the selection circuit, and input the output result to a second storage array, and the second storage array includes multiple rows and multiple columns of DRAM cells.
[0021] In a possible implementation manner, each multiplier includes:
[0022] At least one exclusive-OR unit, the first input end of each exclusive-OR unit is connected to a latch, the second input end of each exclusive-OR unit is connected to a column of DRAM cells, and each exclusive-OR unit is used to perform an exclusive-OR operation on a bit value of an element in the convolution kernel cached in a latch and a bit value of an element in the input data stored in one of the at least two rows of DRAM cells, wherein each of the at least two rows of DRAM cells is respectively used to store a bit value of an element in the input data to be subjected to the convolution calculation;
[0023] A combination unit, the first input end of the combination unit is connected to the output end of the at least one exclusive-OR unit, the second input end of the combination unit is connected to at least one latch, and the combination unit is used to combine the data output by the at least one exclusive-OR unit and the data output by at least one latch.
[0024] Based on the above possible implementation manners, not only is the neural network calculation implemented in the memory, reducing the transmission of neural network data between the memory and the neural network calculation nodes and solving the memory wall problem, but also when implementing the convolution calculation, the convolution calculation is split into exclusive-OR operations, and then the result of the convolution calculation can be obtained through the exclusive-OR unit, thus avoiding splitting the process of the convolution calculation into many basic calculation logics and further reducing the power consumption when implementing the convolution calculation.
[0025] In a second aspect, a computing device is provided, and the computing device includes: a controller, configured to receive a convolution kernel and input data to be subjected to a convolution calculation, and send the convolution kernel and the input data to a memory connected to the controller;
[0026] The memory includes:
[0027] A first storage array, including multiple rows and columns of dynamic random access memory (DRAM) cells, for storing the convolution kernel and input data sent by the controller. Among them, the first row of DRAM cells in the storage array is used to store the convolution kernel, and at least two rows of DRAM cells in the storage array are used to store the input data to be subjected to convolution calculation; and
[0028] A first calculation circuit, connected to the storage array, for performing multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of DRAM cells among the at least two rows of DRAM cells.
[0029] In a possible implementation, the first calculation circuit includes:
[0030] Multiple latches, the input end of each latch is connected to a column of DRAM cells, the output end of each latch is connected to one of multiple multipliers, and each latch is used to cache a one-bit value of an element in the convolution kernel stored in the corresponding column of DRAM cells. Among them, each DRAM cell in the first row of DRAM cells is respectively used to store a one-bit value of an element in the convolution kernel;
[0031] Multiple multipliers, the first input end of each multiplier is respectively connected to a column of DRAM cells, and the second input end of each multiplier is respectively connected to the output end of a latch;
[0032] Multiple adders, the input end of each adder is respectively connected to at least two multipliers;
[0033] Among them, in one convolution calculation process, the multiple multipliers are respectively used to perform multiplication operations on the convolution kernel cached by the multiple latches and the input data stored in one row of the multiple rows of DRAM cells, the multiple adders are respectively used to perform addition operations on the output results of the multiple multipliers, and the output of the multiple adders is the result of one convolution calculation.
[0034] In a possible implementation, the memory further includes:
[0035] A second calculation circuit, the second calculation circuit is connected to the first calculation circuit, and the second calculation circuit includes a summation circuit, which is used to sum the results of the one convolution calculation output by the multiple adders in the first calculation circuit during one summation calculation process.
[0036] In a possible implementation, the second computing circuit is used to perform multiple summation calculations, and the second computing circuit further includes:
[0037] A buffer, connected to the summation circuit, for buffering the calculation results of the summation circuit during multiple summation calculations;
[0038] A pooling circuit, the pooling circuit is connected to the buffer, and is used to perform a pooling operation on the calculation results of the multiple summation calculations cached in the buffer.
[0039] In a possible implementation, the second computing circuit further includes:
[0040] A selection circuit, a first input end of the selection circuit is used to connect to an output end of the pooling circuit, a second input end of the selection circuit is used to connect to an output end of the summation circuit, and the selection circuit is used to select and output the output result of the pooling circuit or the output result of the summation circuit;
[0041] An activation circuit, connected to an output end of the selection circuit, for performing an activation operation on the output result of the selection circuit and inputting the output result to a second storage array, and the second storage array includes multiple rows and multiple columns of DRAM cells.
[0042] In a possible implementation, each multiplier includes:
[0043] At least one exclusive OR unit, a first input end of each exclusive OR unit is connected to a latch, a second input end of each exclusive OR unit is connected to a column of DRAM cells, and each exclusive OR unit is used to perform an exclusive OR operation on a one-bit value of an element in a convolution kernel cached in a latch and a one-bit value of an element in the input data stored in one of the at least two rows of DRAM cells, wherein each of the at least two rows of DRAM cells is respectively used to store a one-bit value of an element in the input data to be subjected to convolution calculation;
[0044] A combination unit, a first input end of the combination unit is connected to an output end of the at least one exclusive OR unit, a second input end of the combination unit is connected to at least one latch, and the combination unit is used to combine the data output by the at least one exclusive OR unit and the data output by at least one latch.
[0045] In a third aspect, a neural network calculation method is provided, and the method is executed by a neural network calculation device, and includes:
[0046] Receiving a convolution kernel of any network layer of a convolutional neural network and input data to be subjected to convolution calculation;
[0047] Store the convolution kernel in the first-row dynamic random access memory (DRAM) cells of the first memory array, where the first memory array includes multiple rows and columns of DRAM cells;
[0048] Store the input data in at least two rows of DRAM cells in the first memory array other than the first-row DRAM cells;
[0049] Based on a first computing circuit connected to the first memory array, perform multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one of the DRAM cells in at least two rows of DRAM cells.
[0050] In a possible implementation, the performing multiple convolution operations on the input data and the convolution kernel based on a first computing circuit connected to the first memory array includes:
[0051] Store the convolution kernel stored in the memory array in multiple latches within the first computing circuit. Each latch is used to cache one element of the convolution kernel stored in the DRAM cells of the corresponding column, and each DRAM cell in the first row of DRAM cells is respectively used to store one element of the convolution kernel;
[0052] Based on multiple multipliers within the first computing circuit, perform a multiplication operation on the convolution kernel cached by the multiple latches and the input data stored in one of the multiple rows of DRAM cells;
[0053] Based on multiple adders within the first computing circuit, perform an addition operation on the output results of the multiple multipliers. The output of the multiple adders is the result of one convolution calculation.
[0054] In a possible implementation, the performing a multiplication operation on the convolution kernel cached by the multiple latches and the input data stored in one of the multiple rows of DRAM cells based on multiple multipliers within the first computing circuit includes:
[0055] Based on at least one exclusive-OR unit within the multiplier, perform an exclusive-OR operation on one-bit value of one element of the convolution kernel cached by one latch and one-bit value of one element of the input data stored in one of the at least two rows of DRAM cells. Each of the at least two rows of DRAM cells is respectively used to store one-bit value of one element of the input data to be subjected to the convolution calculation;
[0056] Based on the combinational units within the multiplier, the combinational units are used to combine the data output by the at least one exclusive - OR unit and the data output by at least one latch.
[0057] In a possible implementation, the method further includes:
[0058] During a summation calculation process, based on the summation circuit within the second calculation circuit in the neural network computing device, sum the results of the one - time convolution calculation output by the multiple adders.
[0059] In a possible implementation, the method further includes:
[0060] Based on the buffer in the neural network computing device, cache the calculation results of the summation circuit during multiple summation calculation processes;
[0061] Based on the pooling circuit in the neural network computing device, perform a pooling operation on the calculation results cached in the buffer during the multiple summation calculation processes.
[0062] In a possible implementation, the method further includes:
[0063] Based on the selection circuit in the neural network computing device, select and output the output result of the pooling circuit or the output result of the summation circuit;
[0064] Based on the activation circuit in the neural network computing device, perform an activation operation on the output result of the selection circuit. Description of the Drawings
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0066] Figure 1 is a schematic diagram of performing convolution processing on a feature map based on a convolution kernel provided by an embodiment of the present invention;
[0067] Figure 2 is a schematic diagram of layer calculation of a convolutional neural network provided by an embodiment of the present invention;
[0068] Figure 3 is a schematic diagram of a neural network computing device provided by an embodiment of the present invention;
[0069] Figure 4 is a schematic diagram of the decomposition of the calculation process of a neural network provided by an embodiment of the present invention;
[0070] Figure 5 It is a schematic diagram of a circuit for implementing exclusive - or calculation provided by an embodiment of the present invention;
[0071] Figure 6 It is a schematic diagram of an adder provided by an embodiment of the present invention;
[0072] Figure 7 It is a block diagram of the logic circuit near PSA in a neural network computing device provided by an embodiment of the present invention;
[0073] Figure 8 It is a schematic diagram of a second computing circuit provided by an embodiment of the present invention;
[0074] Figure 9 It is a schematic diagram of a pooling calculation process provided by an embodiment of the present invention;
[0075] Figure 10 It is a block diagram of the logic circuit near SSA in a neural network computing device provided by an embodiment of the present invention;
[0076] Figure 11 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present invention;
[0077] Figure 12 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0078] Figure 13 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention;
[0079] Figure 14 It is a flowchart of a neural network computing method provided by an embodiment of the present invention;
[0080] Figure 15 It is a schematic diagram of a convolution calculation process provided by an embodiment of the present invention;
[0081] Figure 16 It is a schematic diagram of deploying computing data provided by an embodiment of the present invention. Detailed implementation manners
[0082] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0083] To facilitate the understanding of the technical process of the present invention, the basic principles of the convolutional neural network are first elaborated:
[0084] Overall structure of the convolutional neural network: The convolutional neural network is constructed by N layers (N is a positive integer greater than 1). The output of any layer can be used as the input of the next layer. For example, the feature map output by the first layer can be used as the feature map input of the second layer, the feature map output by the second layer can be used as the feature map input of the third layer, and so on. Until the Nth layer, the feature map is directly output.
[0085] Layers of the convolutional neural network: From the perspective of the layer position, the layers of the convolutional neural network can be divided into the input layer, hidden layers, and output layer. Among them, the input layer (input layer) refers to the first layer, the output layer (output layer) refers to the last layer, and the hidden layers (hidden layer) refer to each layer between the input layer and the output layer. From the perspective of the layer operation method, the layers of the convolutional neural network can be divided into convolutional layers, pooling layers, activation function layers, and fully connected layers, etc. Among them, the convolutional layer is the core layer in the convolutional neural network, and other layers can be regarded as variants of the convolutional layer or further process the feature map output by the convolutional layer.
[0086] Feature map: It can also be called feature mapping, activation map (activation map), activation mapping, convolved feature, etc. The feature map is the data input and output of each layer in the convolutional neural network. The feature map can be image data or speech data. The embodiments of the present invention do not specifically limit the data type of the feature map. Among them, the feature maps output by the first layer to the (N - 1)th layer include multiple feature points. Each feature point in the feature map represents a feature element on the feature map. When the feature map is image data, the feature point represents the pixel value. When the feature map is speech data, the feature point represents the speech value. The feature map output by the Nth layer may only include one feature point, and this feature point represents the category to which the feature map recognized by the convolutional neural network belongs. Or the feature map output by the Nth layer may include multiple feature points, and each feature point represents the category to which the feature map recognized by the convolutional neural network may belong and the probability of belonging to the corresponding category.
[0087] Convolution Kernel: It can also be called a filter, a feature detector, a kernel function, a kernel, a weight kernel, etc. A convolution kernel can be regarded as a weight matrix with a scanning window. The weight matrix includes multiple weights, and each weight can also be called a parameter. Different convolution kernels have different weights, so as to be able to detect different features in the feature map. The size of the convolution kernel (filter size, that is, the size of the weight matrix) can be determined according to actual business needs. Among them, the size of each convolution kernel in the convolutional layer is L×M (L and M are positive integers greater than 1), such as 2×2, and the size of each convolution kernel in the fully connected layer is 1×1. Each convolution kernel is used to perform convolution on the feature map to obtain an output feature map. Each layer can include at least one convolution kernel, and at least one feature map can be output through at least one convolution kernel.
[0088] Scanning Window: It is used to determine the area targeted for each convolution process in the feature map. The size of the scanning window is equal to the size of the convolution kernel (filter size, that is, the size of the weight matrix). The stride of the scanning window is used to determine the distance that the scanning window slides each time. The stride of the scanning window is determined according to actual needs and can be 1.
[0089] Calculation Principle of Convolution Kernel: Convolution processing can be understood as a process of corresponding multiplication and then accumulation. During the convolution process of the convolution kernel, the scanning window of the convolution kernel slides on the feature map, and by scanning each area of the feature map, each area of the feature map is calculated in turn. When the scanning window of the convolution kernel slides to any area of the feature map, the feature elements in that area are read, and a dot multiplication operation is performed on the feature elements in that area and the convolution kernel to obtain a feature point. Then, the scanning window slides a stride, so as to be in the next area of the feature map, and another feature point is output. When the sliding process ends, all the output feature points can be combined into a feature map and output to the next layer.
[0090] Exemplarily, see Figure 1 , which shows a schematic diagram of performing convolution processing on a feature map based on a convolution kernel. The size of this convolution kernel is 3×3, and the convolution kernel can be represented as a matrix of [1, 0, 1; 0, 1, 0; 1, 0, 1]. The stride of the scanning window is 1. Assuming that the scanning window of the convolution kernel is in the area shown in the figure, the feature matrix within the scanning window of the feature map is [1, 1, 1; 0, 1, 1; 0, 0, 1]. A dot multiplication operation is performed on the convolution kernel and the feature matrix to obtain a feature point. Specifically, the feature point = 1*1 + 0*1 + 1*1 + 0*0 + 1*1 + 0*1 + 1*0 + 0*0 + 1*1 = 4.
[0091] After that, the scanning window slides once, and then the dot multiplication operation is performed on the feature matrix and the convolution kernel within the scanning window at this time to obtain another feature point. By analogy, when the sliding of the scanning window ends, all the obtained feature points can form the output feature map.
[0092] For the convenience of understanding each calculation process in the convolutional neural network, refer to Figure 2 the schematic diagram of the layer calculation of a convolutional neural network provided by the embodiment of the present invention shown in Figure 2 The large cuboid in it is the feature map to be recognized in a certain layer of the convolutional neural network, and the small cube is the convolution kernel. The length, width, and height of the large cuboid are 3, 3, and 2 respectively. This feature map can be expressed as 3x3x2. Among them, the height of the large cuboid can be regarded as the depth of the feature map, that is, the number of layers of the feature map. Figure 2 This feature map in it consists of two layers of data. The length, width, and height of the small cube are all 2. The convolution kernel can be expressed as 2x2x2. Among them, the height of the small cube can be regarded as the depth of the convolution kernel, and it can also be regarded as the number of convolution kernels. That is, this feature map requires 2 convolution kernels for convolution. Each convolution kernel is used to perform convolution calculation on 1 layer of this feature map. When the convolution calculation of each convolution kernel is completed, a 2x2 feature map can be obtained, and then the 2x2 feature map is subjected to pooling processing, and then the pooled feature map is output.
[0093] In some embodiments, the above calculation process can be implemented within a neural network computing device. To further illustrate the hardware structure of this neural network computing device, refer to Figure 3 , Figure 3 is a schematic diagram of a neural network computing device provided by the embodiment of the present invention. The neural network computing device includes: a first storage array, including multiple rows and multiple columns of dynamic random access memory (DRAM) units. The first row of DRAM units in this storage array is used to store convolution kernels, and at least two rows of DRAM units in this storage array are used to store input data to be subjected to convolution calculation; a first calculation circuit, connected to this storage array, for performing multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform multiplication and addition calculations on the convolution kernel and the input data stored in one row of DRAM units among the at least two rows of DRAM units.
[0094] The first storage array can be a subarray in a DRAM, and a DRAM cell can be a cell in the subarray. Multiple DRAM cells can be arranged in a row of DRAM cells, and each row of DRAM cells is connected to a word line. Any word line is used to select a row of DRAM cells connected to that word line. Multiple DRAM cells can also be arranged in a column of DRAM cells, and each column of DRAM cells is connected to a bit line. Any bit line is used to select a column of DRAM cells connected to that bit line. When the neural network computing device selects any bit line and any word line, the any bit line and the any word line intersect at a DRAM cell. Furthermore, the neural network computing device can read the data stored in the DRAM cell or write data into the DRAM cell.
[0095] Each DRAM cell can include a capacitor and a transistor. The capacitor can store 1 bit of data volume. The amount of charge (potential level) after charging and discharging corresponds to binary data 0 and 1 respectively. A DRAM cell can be used to store 0 or 1. Since a weight can be represented by a numerical value of at least one binary bit, therefore, a weight can be stored by at least one DRAM cell, that is, a 1-bit binary value after quantization of a weight can be stored in one DRAM cell. Similarly, one DRAM cell can also store a 1-bit binary value of an input data to be subjected to convolution calculation.
[0096] In order to store data in the first storage array or read data from the first storage array, the neural network computing device can further include a row decoder and a column decoder. The row decoder is connected to each word line in the first storage array, so that each row of DRAM cells in the first storage array can be connected through the word line; the column decoder is connected to each bit line in the first storage array, so that each column of DRAM cells in the first storage array can be connected through the bit line. The neural network computing device can activate the DRAM cells connected to any word line and the DRAM cells connected to any bit line through the column decoder and the row decoder, so that read operations and write operations can be performed on the intersecting DRAM cells. The read operation is used to read the data stored in the intersecting DRAM cell, and the write operation is used to store data in the intersecting DRAM cell. Furthermore, reading data and storing data in the first storage array can be realized.
[0097] Specifically, the row decoder can select a corresponding row address (one row address corresponds to one word line) through the row address in the row address buffer to activate a specific row address, thereby achieving the addressing of the row address. After the specific row address is activated, the column decoder selects a corresponding column address (one column address corresponds to one bit line) through the column address in the column address buffer to activate a specific column address, thereby achieving the addressing of the column address. After the specific column address is also activated, the activated row address and the activated column address intersect at the position where a DRAM cell is located, so that the addressing of the DRAM cell can be achieved. Data can be stored in the DRAM cell through a write operation, or the data stored in the DRAM cell can be read through a read operation.
[0098] In some embodiments, the first memory array may further include at least one primary sense amplifier (PSA) and at least one secondary sense amplifier (SSA). The input end of a PSA is connected to a column of DRAM cells to amplify the amount of charge stored in the storage element in the DRAM. The amplified amount of charge is used to indicate whether the data stored in the DRAM cell is 0 or 1. Specifically, when the amplified amount of charge is greater than a preset value, it indicates that the data stored in the storage element is 1; otherwise, it indicates that the data stored in the storage element is 0. The present invention does not specifically limit this preset value.
[0099] The input end of an SSA is connected to multiple first memory arrays. The SSA is used to select and read out the output result of the PSA, and then write the output result to the read-write bus of the neural network computing device.
[0100] It should be noted that the above-mentioned first row of DRAM cells is used to refer to any row of DRAM cells in the first memory array. Since there can be multiple convolution kernels in the convolution layer of the convolutional neural network, each convolution kernel can be stored in any row of DRAM cells in the first memory array. For example, convolution kernel 1 is stored in the 9th row of the memory array, and convolution kernel 2 is stored in the 10th row of the memory array.
[0101] It should be noted that the neural network computing device may include multiple first storage arrays. In some embodiments, since the depth of the input feature map (the input data for performing convolution calculation) may be relatively large, the input feature map can be split into multiple layers, and the convolution kernel is used to perform convolution on each layer of the feature map. Then, there may be a relatively large amount of input data for the convolution calculation corresponding to the convolution kernel. And when the number of weights in a convolution kernel is too large, the neural network computing device can also split the convolution kernel with too many weights into multiple sub-convolution kernels, and each sub-convolution kernel can correspond to multiple groups of input data for the convolution calculation to be performed. Eventually, the number of convolution kernels and the input data for the convolution calculation corresponding to the convolution kernels in any network layer is relatively large. Considering the actual area of the first storage array, then, a single first storage array may not be able to store a large number of convolution kernels and the input data for the convolution calculation to be performed. Therefore, multiple first storage arrays can be used to jointly store the convolution kernels and the input data for the convolution calculation to be performed in any network layer.
[0102] To further illustrate the principle of splitting the input feature map and convolution kernel, refer to FIG. 4. Figure 4 FIG. 4 is a schematic diagram of the decomposition of the computing process of a neural network provided by an embodiment of the present invention. As can be seen from Figure 4 FIG. 4, the input feature map is a 3D feature map. The input feature map is split into multiple two-dimensional sub-feature maps. The scanning window of the 4×4 convolution kernel, and the convolution sliding sub-interval calculation process is further decomposed into the 3×3 convolution kernel sliding 4 times on the 4×4 scanning window. The area covered by the scanning window after each sliding is regarded as a sub-feature map, and each convolution kernel will perform convolution on each sub-feature map. The data on each sub-feature map can form a group of input data for the convolution calculation to be performed. Each group of input data for the convolution calculation to be performed can be stored in any row of DRAM cells other than the first row of DRAM cells in the storage array. Therefore, at least two rows of DRAM cells in the storage array can be used to store the input data for the convolution calculation to be performed.
[0103] From the above convolution calculation principle, it can be seen that the convolution calculation includes multiple calculation processes. For example, multiplying the weights by the corresponding input data, and adding the products of the weights in the convolution kernel and the corresponding input data to obtain feature points. Since the convolution kernel may correspond to multiple groups of input data for the convolution calculation to be performed, multiple feature point outputs corresponding to the convolution kernel will be obtained. These multiple feature points can form the feature map after one convolution of the convolution kernel, that is, the result of one convolution calculation. Therefore, the first calculation circuit is connected to the first storage array so that the first calculation circuit can read the convolution kernel and the input data for the convolution calculation to be performed from the first storage array, and implement the convolution calculation based on the read data.
[0104] To obtain the result of a convolution calculation for any convolution kernel, in one possible implementation, the first calculation circuit includes: a plurality of latches, the input end of each latch is connected to a column of DRAM cells, the output end of each latch is connected to one of a plurality of multipliers, and each latch is used to cache a single-bit value of an element in the convolution kernel stored in the corresponding column of DRAM cells, where each DRAM cell in the first row of DRAM cells is respectively used to store a single-bit value of an element in the convolution kernel; a plurality of multipliers, the first input end of each multiplier is respectively connected to a column of DRAM cells, and the second input end of each multiplier is respectively connected to the output end of a latch; a plurality of adders, the input end of each adder is respectively connected to at least two multipliers; where, during a convolution calculation process, the plurality of multipliers are respectively used to perform a dot multiplication operation on the convolution kernel cached by the plurality of latches and the input data stored in one row of the plurality of rows of DRAM cells, the plurality of adders are respectively used to perform an addition operation on the output results of the plurality of multipliers, and the output of the plurality of adders is the result of a convolution calculation.
[0105] The weights in the convolution kernel can be regarded as elements in the convolution kernel, that is, one weight is one element. Correspondingly, one data in the input data to be subjected to convolution calculation can also be regarded as one element in the input data to be subjected to convolution calculation. Since the quantized convolution kernel and the quantized input data are stored in the first storage array, and the quantized data is represented by binary values, then an element in the convolution kernel can be represented by data of multiple bits. For example, the binary bits 11 after quantization of the weight value 3, where the first 1 in 11 is the most significant bit of the weight value 3, and the second 1 is the second most significant bit of the weight value 3. Thus, two DRAM cells in one row can be used to store the weight 3. Correspondingly, multiple DRAM cells in one row can also be used to store one input data in the input data to be convolved. The embodiments of the present invention do not make specific limitations on the number of bits of the quantized weights and the quantized input data.
[0106] For any convolution kernel, when the neural network computing device performs convolution calculation, it calculates row by row in the first storage array, and the process is as follows:
[0107] First, the neural network computing device selects the row of DRAM cells storing the convolution kernel, opens multiple latches, and outputs the elements in the convolution kernel to the multiple latches for caching. Then, the connection between the multiple latches and the first memory array is disconnected, and a row of DRAM cells (e.g., the first row of DRAM cells storing the input data of the first window) is opened. Thus, the multiplier performs the first multiplication calculation on the input data of the first row of DRAM cells and the elements in the convolution kernel in the latches. An adder is connected to at least two multipliers to perform the first addition operation on the calculation results of at least two multipliers. The calculation results of multiple adders constitute the result of the first convolution calculation. The result of one convolution calculation is also a feature point. During the second calculation, the second row of DRAM cells (storing the input data of the second window) is opened, and the multiplier performs the second multiplication calculation on the input data stored in the second row of DRAM cells and the convolution kernel in the latches. The adder respectively performs the second addition operation on the calculation results of at least two multipliers. The calculation results of multiple adders constitute the result of the second convolution calculation, and so on until the input data to be convolved is completely calculated. Among them, the first window can be the window after the scanning window of the convolution kernel slides for the first time on the input feature map, and the second window is the window after the scanning window of the convolution kernel slides for the second time on the input feature map.
[0108] When the neural network computing device implements the calculation in any network layer of the convolutional neural network, multiple convolution kernels may be required. For each convolution kernel, the neural network computing device performs the convolution calculation according to the rows of the first memory array as described above, so as to obtain the convolution result of each convolution kernel.
[0109] To implement the dot multiplication operation in the convolution process, in a possible implementation, each multiplier includes: at least one exclusive-OR unit, the first input terminal of each exclusive-OR unit is connected to a latch, the second input terminal of each exclusive-OR unit is connected to a column of DRAM cells, and each exclusive-OR unit is used to perform an exclusive-OR operation on one-bit value of an element in the convolution kernel cached in a latch and one-bit value of an element in the input data stored in one of the at least two rows of DRAM cells, where each of the at least two rows of DRAM cells is respectively used to store one-bit value of an element in the input data to be convolved; a combination unit, the first input terminal of the combination unit is connected to the output terminal of the at least one exclusive-OR unit, and the second input terminal of the combination unit is connected to at least one latch, and the combination unit is used to combine the data output by the at least one exclusive-OR unit and the data output by at least one latch. For the convenience of description, the output data of the combination unit is called the first data, and the first data is also the product of an element in the convolution kernel and an element in the input data to be convolved.
[0110] An exclusive - OR unit can output the second - highest bit of the first data, and then the combinational unit can combine the weights in the convolutional kernel with the second - highest bit of the first data to obtain the first data.
[0111] To illustrate the reasons for using the exclusive - OR unit and the combinational unit, take 1 - bit weight W0 and 2 - bit input data D0 as an example. D0 can be represented by 1 - bit D0_0 and 1 - bit D0_1. Specifically, D0 = D0_0D0_1, and the product of D0 and W0 can be expressed as:
[0112]
[0113] From the above formula, when W0 = 0, the highest bit of the product of D0 and W0 is W0, and the second - highest bit of the product of D0 and W0 is equal to the exclusive - OR of D0 and W0. When W0 = 1, the highest bit of the product of D0 and W0 is the 1 added in i.e., the highest bit of the product of D0 and W0 is W0, and the second - highest bit of the product of D0 and W0 is the second - highest bit, and the second - highest bit
[0114] Since when quantizing weights, if any weight is less than or equal to the first preset value, the first weight is quantized to 0, and when any weight is greater than the first preset value, the first weight is quantized to 1. Then, in a possible implementation, for any weight stored in the at least one latch, when the any weight is less than or equal to the first preset value, the any weight is used as the highest bit of the first data, and the exclusive - OR of the any weight and the corresponding input data is used as the second - highest bit of the first data, and thus a first data can be obtained; when the any weight is greater than the first preset value, the any weight is used as the highest bit of the second data, the first data corresponding to the any weight is used as the second - highest bit of the second data, and the second data is added to the second preset value to obtain the first data. The embodiments of the present invention do not specifically limit the first preset value and the second preset value.
[0115] To further illustrate the input and output of the exclusive - OR unit, based on the example in Figure 4 see Figure 5 Figure 5 This is a schematic diagram of a circuit for implementing exclusive - OR calculation provided by an embodiment of the present invention. A latch is connected to a switch, and this switch is connected to a PSA. When performing calculations, first, the neural network computing device closes the switch and selects the row decoder through a read operation to open the row where the weight W0 is located. The weight W0 stored in the opened row is written into the latch through the PSA. Then, the switch is disconnected to disconnect the data bus from the latch. The row where the input data D0 is located is selected through a read operation. After being amplified by the PSA, the input data D0_1 is input to one input terminal of the exclusive - OR unit 1, and the weight W0 stored in the latch is input to the other input terminal of the exclusive - OR unit 1. Thus, the exclusive - OR unit 1 can output the exclusive - OR of W0 and D0_1. Similarly, the exclusive - OR unit 2 can output the exclusive - OR of W0 and D0_0. The W0, the exclusive - OR of W0 and D0_1, and the exclusive - OR of W0 and D0_0 in the latch are input to the input terminal of the combination unit so that the combination unit can combine them into the product of D0 and W0. For each element in a convolution kernel, the above - mentioned process is executed, and the products of each element and each input data can be obtained. Furthermore, each product can be added by multiple adders to obtain a feature point, which is also the result of one convolution calculation.
[0116] Since the convolution kernels stored in the first storage array are quantized convolution kernels, and the stored input data is also quantized data. Each weight in each convolution kernel may be 1 - bit data or multi - bit data after quantization. Similarly, each input data may be 1 - bit data or multi - bit data after quantization. Therefore, the quantity stored in the first array is relatively large, and multiple adders can be used to implement step - by - step summation to obtain the result of one convolution calculation.
[0117] To further illustrate the principle of step - by - step summation, refer to Figure 6 , Figure 6 This is a schematic diagram of an adder provided by an embodiment of the present invention. Figure 6 One of them can process at most 32 3 - bit data. Thus, this adder (which can also be called an adder tree) can add 32 3 - bit data to obtain an 8 - bit sum. Then, when the convolution kernel size is greater than 32, the result of one convolution calculation is the sum result of more than 32 3 - bit data. The 8 - bit sum obtained by this one adder tree cannot be used as the result of one convolution calculation. At least the partial sum results output by multiple adder trees need to be summed again to obtain the result of one convolution calculation.
[0118] To further illustrate the process of the neural network computing device implementing convolution operation, based on the example in Figure 5 , refer to Figure 7 , Figure 7It is a logic circuit block diagram near the PSA in a neural network computing device provided by an embodiment of the present invention. Inside the dashed box is the newly added logic circuit diagram in the neural network computing device. It can be seen that each PSA is connected to a column of the first storage array. The weights W0 and W1 in the convolution kernel are stored in the first row of the first storage array, and a set of input data D0 (D0_0D0_1) and D1 (D1_0D1_1) are stored in the second row of the first storage array. Through the PSA, the weight W0 can be stored in the latch, and the W0 stored in the latch and D0_0 in the first storage array are input to an exclusive OR unit. The exclusive OR unit outputs the exclusive OR of W0 and D0_0. The exclusive OR of W0 and D0_0, the exclusive OR of W0 and D1_0, and W0 are input to the combination unit. The combination unit outputs the product D0*W0 of D0 and W0. Similarly, another combination unit can obtain the product D1*W1 of D1 and W1. The sum of D0*W0 and D1*W1 is obtained through the adder tree (adder), and the dot product result of the convolution kernel [W0, W1] and a set of input data [D0, D1] is obtained, which is also the result of one convolution calculation of the convolution kernel [W0, W1].
[0119] When the convolution kernel is greater than 32, the summation process needs to be split. Therefore, the result of one convolution calculation is not the final convolution result. Thus, it is also necessary to perform an accumulation process on the results of multiple convolution calculations to obtain the final convolution result. In a possible implementation manner, the neural network computing device further includes: a second computing circuit, which is connected to the first computing circuit. The second computing circuit includes a summation circuit for summing the results of the one convolution calculation output by multiple adders in the first computing circuit during one summation calculation process.
[0120] The summation circuit can obtain the final convolution result through cyclic accumulation. In one possible implementation, the first input terminal of the summation circuit is connected to the output terminal of the summation circuit, and the second input terminal of the summation circuit is connected to multiple adders. After any adder inputs the result of a convolution calculation to the summation circuit once, the summation circuit adds the result of a convolution calculation input from the second input terminal to the summation result of N convolution calculations input from the first input terminal. The output terminal of the summation circuit can output the summation result of N + 1 convolution calculations, where N is a positive integer. For example, at the beginning, the data input to the first input terminal of the summation circuit is T0, where T0 = 0 is the summation result of 0 convolution calculations. The data input to the second input terminal of the summation circuit is S1, where S1 is the summation result of the first convolution calculation output by a first calculation circuit. The output terminal of the summation circuit outputs T1 = T0 + S1 to the first input terminal of the summation circuit. When the data input to the second input terminal of the summation circuit is S2 again, the output terminal of the summation circuit outputs T2 = T1 + S2 to the first input terminal of the summation circuit, where S2 is the summation result of the first convolution calculation output by the first calculation circuit, and T2 is the summation result of two convolution calculations. When the result of a convolution calculation output by the first calculation circuit is the result of the last convolution calculation, the output terminal of the summation circuit can output the final convolution result.
[0121] It should be noted that by splitting the dot multiplication operation in the convolution calculation into exclusive OR operations, the process of splitting the convolution calculation into many basic calculation logics is avoided, further reducing the power consumption when implementing the convolution calculation.
[0122] Since the summation circuit sums the results of multiple convolution calculations, each time the summation circuit performs an accumulation, it also needs to determine whether the result of a convolution calculation in this accumulation is the result of the last convolution calculation among the results of multiple convolution calculations. If the result of the convolution calculation in this accumulation is the result of the last convolution calculation, the summation circuit ends the accumulation; otherwise, the summation circuit will continue to receive the results of the convolution calculation and perform the accumulation again. For example, Figure 8 the summation circuit in the schematic diagram of a second calculation circuit provided by the embodiment of the present invention shown Figure 8 the 8-bit data input to an input port 1 of the summation circuit is the output data of a first calculation circuit, and the 32-bit data input to input port 2 is the result of the previous summation of the summation circuit. After the summation circuit finishes the summation this time, it inputs the result of this summation to input port 2 so that when input port 1 receives another 8-bit data again, another summation calculation is performed until the final 32-bit summation result is obtained. The final 32-bit summation result is also the final convolution result.
[0123] In some embodiments, the computing device may perform a pooling operation on the final convolution result to downsample the final convolution result. In one possible implementation, the second computing circuit is used to perform multiple summation calculations. The second computing circuit further includes: a buffer connected to the summation circuit for buffering the calculation results of the summation circuit during multiple summation calculations; a pooling circuit connected to the buffer for performing a pooling operation on the calculation results of the multiple summation calculations buffered in the buffer.
[0124] Before introducing the pooling circuit, taking Figure 9 a schematic diagram of a pooling calculation process provided by an embodiment of the present invention shown below as an example, the pooling operation is described as follows: For ease of description, the final convolution result is referred to as the first feature map. The purpose of the pooling operation is to downsample the first feature map. Then, some relatively large feature points can be selected from the first feature map to achieve downsampling. Specifically, for a 4*4 first feature map, 4 2*2 pooling kernels can be used for pooling. 2*2 indicates the size of the pooling kernel. The neural network computing device can split the 4*4 first feature map into 4 2*2 first sub-feature Figures 1 - 4 , and the size of each 2*2 first sub-feature map is the same as the size of a pooling kernel. Then, a largest feature point is selected from each first sub-feature map. For example, the feature points in the first sub-feature Figure 1 include 1, 1, 5, 6. Then, the feature point 6 can be selected from the first sub-feature Figure 1 . Thus, 4 feature points are selected from the first feature map, and these 4 feature points are combined into a second feature map. The second feature map is also the result of performing a pooling operation on the first feature map.
[0125] To implement the pooling operation within the pooling circuit, the pooling circuit further includes multiple comparators, and each comparator is used to compare the feature points in each first sub-feature map. In one possible implementation, a comparator includes two input interfaces and an output interface. The two input interfaces are the input interface 1 and the input interface 2 respectively. The input interface 1 is connected to the output interface of the comparator. When a feature point in a first sub-feature map is input to the input interface 2, if any feature point is received by the input interface 2, the comparator compares the feature points input to the input interface 1 and 2, and outputs a larger feature point from the output interface. When the any feature point is not the last feature point in the first sub-feature map, the output interface inputs the larger feature point to the input interface 2 for the next comparison. When the any feature point is the last feature point in the first sub-feature map, the input interface 2 outputs the larger feature point to the combination sub-unit. Furthermore, each comparator can output a feature point respectively through cyclic comparison, so that the pooling circuit can output a second feature map. Another exampleFigure 8 The pooling circuit in performs a cyclic comparison on 16 32-bit data output from the buffer, and a 32-bit data can be obtained, which is the result of the pooling operation.
[0126] In some embodiments, the neural network computing device further needs to perform an activation operation on the pooling result to obtain the feature map output by the network layer. The activation operation refers to calculating the feature map using an activation function to perform a non-linear transformation on the feature map. The activation function can be a rectified linear unit (ReLU) function. The mathematical expression of the ReLU function is f(x) = max(0, x), which will convert the feature points where the pooling result is less than 0 into 0, resulting in high sparsity of the output feature map. In some embodiments, the pooling operation may not be performed on the final convolution result, and the activation operation may be directly performed. In a possible implementation manner, the second computing circuit further includes: a selection circuit, the first input end of the selection circuit is used to connect to the output end of the pooling circuit, the second input end of the selection circuit is used to connect to the output end of the summing circuit, and the selection circuit is used to select and output the output result of the pooling circuit or the output result of the summing circuit; an activation circuit, connected to the output end of the selection circuit, for performing an activation operation on the output result of the selection circuit and inputting the output result to the second storage array, and the second storage array includes multiple rows and multiple columns of DRAM cells. For example, Figure 7 the selection circuit in has input interfaces 0 and 1, where input interface 0 is connected to the summing circuit, and input interface 1 is connected to the pooling circuit. When the output result of the summing circuit meets the preset condition, the selection circuit outputs the output result of the summing circuit. When the output result of the pooling circuit meets the preset condition, the selection circuit outputs the output result of the pooling circuit.
[0127] The second storage array can be the first storage array or other storage arrays other than the first storage array. The output result is input to the second storage array so that the neural network computing device can implement the calculation process of the next network layer through the second storage array. Thus, the calculation process of each network layer of the neural network can be implemented in the neural network computing device. When the neural network computing device completes the calculation process in the last network layer of the neural network, then, the result output by the activation circuit is also the calculation result of the entire neural network.
[0128] To further illustrate the process of the second computing circuit in the neural network computing device during calculation, based on the example in Figure 8 and referring to Figure 10 , Figure 10It is a block diagram of the logic circuit near the SSA in a neural network computing device provided by an embodiment of the present invention. The adder outputs the result of each convolution calculation of the convolution kernel to the summation circuit in the second computing circuit through the SSA. In the summation circuit, the results of multiple convolution calculations are summed to obtain the final convolution result, and then the final convolution result is input into the buffer. The convolution result in the buffer is output to the pooling circuit, and the pooling circuit performs a pooling operation on the final convolution result. Finally, the pooled result is input into the activation unit, and the activation unit performs an activation operation on the pooled result. Thus, the activation unit outputs the output result of a network layer, and then the output result of the activation unit is input into the array write-back circuit. The array write-back circuit rewrites the output result into the second storage array for calculating the next network layer.
[0129] When the scale of the first storage array is 1024x1024, there can be 16 adder trees (adders) in one first storage array, and 16 convolution processes can be supported simultaneously. The neural network computing device can include multiple partial banks. Each bank can include multiple blocks, and each block includes 16 first storage arrays. In fact, through layout and wiring, it is found that if all are configured as computing circuits, the area overhead of the neural network computing device is too large. Therefore, 8 of the first storage arrays are configured as computing arrays (8 multiply accumulate (MAC) rows), and the remaining 8 first storage arrays remain unchanged, which can achieve a better balance of area performance. In this case, each bank has 64 SSAs, and all the first storage arrays can perform parallel and independent calculations. Multiple banks have independent control and read / write buses and can also run in parallel. This architecture fully considers the repeatability and parallelism of the internal resources of the computing device and can achieve the maximum parallelization of distributed computing operations. Due to the high repeatability of the first storage array design, for the failed modules that may occur in the manufacturing process of the computing device, before storing data, the DRAM cells can be verified and marked first to ensure the yield of chip manufacturing.
[0130] It should be noted that the neural network computing device may further include a data input / output interface. Among them, the data input interface is used to receive the convolution kernel, the input data to be convolved, etc. sent by other devices, and the data output interface is used to output the data output by the last network layer of the neural network to other devices.
[0131] The neural network computing device provided by the embodiments of the present invention stores the convolution kernel and the input data to be subjected to convolution calculation in a first storage array including multiple rows and columns of DRAM cells. The first computing circuit can read the convolution kernel and the input data to be subjected to convolution calculation from the first storage array and complete the convolution calculation, thereby implementing neural network computing in the memory, reducing the transmission of neural network data between the memory and the neural network computing node, and solving the problem of the memory wall. Moreover, the neural network computing device splits the multiplication operation in the convolution calculation into exclusive OR operations, thereby avoiding splitting the process of convolution calculation into many basic computing logics, and further reducing the power consumption when implementing convolution calculation. Multiple first storage arrays can be included in the neural network computing device, so as to achieve parallel operation and improve the efficiency of completing the convolution neural network calculation process in the memory. Moreover, the computing processes such as convolution, pooling, and activation in the convolution neural network can all be completed within the neural network computing device, so that all the computing processes in the neural network can be implemented in the memory, and the transmission of neural network data between the memory and the neural network computing node can be further reduced.
[0132] The neural network computing device can be installed on a computer device, and the computer device controls the neural network computing device, so that the computer device can complete the entire computing process of the neural network in the neural network computing device. As Figure 11 shown, Figure 11 FIG. 7 is a schematic structural diagram of a computing device provided by an embodiment of the present invention. The computing device 1100 includes:
[0133] A controller 1101, configured to receive a convolution kernel and input data to be subjected to convolution calculation, and send the convolution kernel and the input data to a memory 1102 connected to the controller;
[0134] The memory 1102 includes:
[0135] A first storage array, including multiple rows and columns of dynamic random access memory (DRAM) cells, for storing the convolution kernel and input data sent by the controller. Among them, the first row of DRAM cells in the storage array is used to store the convolution kernel, and at least two rows of DRAM cells in the storage array are used to store the input data to be subjected to convolution calculation; and
[0136] A first computing circuit, connected to the storage array, for performing multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of the at least two rows of DRAM cells.
[0137] Optionally, the first computing circuit includes:
[0138] Multiple latches, the input terminal of each latch is connected to a column of DRAM cells, the output terminal of each latch is connected to one of the multiple multipliers, and each latch is used to cache a bit value of an element in the convolution kernel stored in the DRAM cells of the corresponding column. Among them, each DRAM cell in the first row of DRAM cells is respectively used to store a bit value of an element in the convolution kernel;
[0139] Multiple multipliers, the first input terminal of each multiplier is respectively connected to a column of DRAM cells, and the second input terminal of each multiplier is respectively connected to the output terminal of a latch;
[0140] Multiple adders, the input terminal of each adder is respectively connected to at least two multipliers;
[0141] Among them, during a convolution calculation process, the multiple multipliers are respectively used to perform a multiplication operation on the convolution kernel cached by the multiple latches and the input data stored in one row of DRAM cells among the multiple rows of DRAM cells, the multiple adders are respectively used to perform an addition operation on the output results of the multiple multipliers, and the output of the multiple adders is the result of one convolution calculation.
[0142] Optionally, the memory further includes:
[0143] A second calculation circuit, the second calculation circuit is connected to the first calculation circuit, and the second calculation circuit includes a summation circuit, which is used to sum the results of the one convolution calculation output by the multiple adders in the first calculation circuit during a summation calculation process.
[0144] Optionally, the second calculation circuit is used to perform multiple summation calculations, and the second calculation circuit further includes:
[0145] A buffer, connected to the summation circuit, for caching the calculation results of the summation circuit during multiple summation calculation processes;
[0146] A pooling circuit, the pooling circuit is connected to the buffer, and is used to perform a pooling operation on the calculation results of the multiple summation calculation processes cached in the buffer.
[0147] Optionally, the second calculation circuit further includes:
[0148] A selection circuit, the first input terminal of the selection circuit is used to connect to the output terminal of the pooling circuit, the second input terminal of the selection circuit is used to connect to the output terminal of the summation circuit, and the selection circuit is used to select and output the output result of the pooling circuit or the output result of the summation circuit;
[0149] An activation circuit, connected to the output end of the selection circuit, is configured to perform an activation operation on the output result of the selection circuit and input the output result to a second memory array, where the second memory array includes multiple rows and multiple columns of DRAM cells.
[0150] Optionally, each multiplier includes:
[0151] At least one exclusive-OR unit, where a first input end of each exclusive-OR unit is connected to a latch, a second input end of each exclusive-OR unit is connected to a column of DRAM cells, and each exclusive-OR unit is configured to perform an exclusive-OR operation on a bit value of an element in a convolution kernel cached by a latch and a bit value of an element in the input data stored in one of the at least two rows of DRAM cells, where each of the at least two rows of DRAM cells is respectively configured to store a bit value of an element in the input data to be subjected to convolution calculation;
[0152] A combination unit, where a first input end of the combination unit is connected to the output end of the at least one exclusive-OR unit, and a second input end of the combination unit is connected to at least one latch, and the combination unit is configured to combine the data output by the at least one exclusive-OR unit and the data output by at least one latch.
[0153] The computing device provided by the embodiment of the present invention stores a convolution kernel and input data to be subjected to convolution calculation in a first memory array including a memory. A first computing circuit can read the convolution kernel and the input data to be subjected to convolution calculation from the first memory array and complete the convolution calculation, thereby implementing neural network calculation in the memory, reducing the transmission of neural network data between the memory and a neural network calculation node, and solving the problem of the memory wall. Moreover, in the memory, the multiplication operation in the convolution calculation is split into exclusive-OR operations, thereby avoiding splitting the process of the convolution calculation into many basic calculation logics and further reducing the power consumption when implementing the convolution calculation. The memory may include multiple first memory arrays, so that parallel operation can be implemented, and the efficiency of completing the convolution neural network calculation process in the memory is improved. In addition, calculation processes such as convolution, pooling, and activation in the convolution neural network can all be completed in the memory, so that all calculation processes in the neural network can be implemented in the memory, and further reducing the transmission of neural network data between the memory and a neural network calculation node.
[0154] See Figure 12 , Figure 12 which is a schematic diagram of an implementation environment provided by the embodiment of the present invention. Figure 12The computer device includes a bus interface (peripheral component interconnect-express, PCI-e), a controller, and dual inline memory modules (DIMMs). Users can install neural network computing devices within the DIMMs. Users can train a quantized neural network according to the requirements of the application scenario to obtain a network structure and weight values that meet the accuracy requirements. Then, the computer device uses a compiler to process the original data (such as images, audio, etc.), network structure, and weights of the application scenario. After processing, they are allocated to the memory address space and an instruction set is generated. Then, the computer device writes the original data, weights, and instruction set into the neural network computing device. Finally, the neural network computing device automatically performs calculations. After the calculations are completed, it sends a message to the computer device's central processing units (CPUs) through an interrupt. The CPU reads back the result data to complete the entire neural network computing operation. When compiling the original data, the compiler needs to call functions in the function library to perform the compilation. The compiler can compile the original data, etc. based on the programming language and send the compilation result to the driver of the computing device. Then, the driver controls the runtime to run the compilation result within the computer device. It should be noted that the application programming interface (API) in the user interface can also call functions in the function library to implement the functions of the API.
[0155] Figure 13 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present invention. The computer device 1300 may vary greatly due to different configurations or performances, and may include one or more CPUs 401 and one or more memories 1302. Among them, at least one instruction is stored in the memory 1302, and the at least one instruction is loaded and executed by the processor 1301 to implement the methods provided by the following various method embodiments. Of course, the computer device 1300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The computer device 1300 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0156] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the neural network calculation method in the following embodiments. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0157] To further illustrate the process of the neural network computing device implementing the calculations within the convolutional neural network, refer to Figure 14 the flowchart of a neural network calculation method provided by an embodiment of the present invention as shown in
[0158] 1401. The computer device sends a convolutional kernel and the input data of the first network layer to be subjected to convolutional calculation to the neural network computing device.
[0159] This step can be executed by the processor in the computer device. In an embodiment of the present invention, the neural network computing device can be the memory in the computer device. The memory can include multiple rows and multiple columns of DRAM units. Among them, the first network layer can be any layer in the neural network. The input data to be subjected to convolutional calculation can be divided into multiple groups of input data. Each group of input data includes multiple input data. Each group of input data is the data in the area covered by the scanning window of a convolutional kernel after sliding once on the feature map input by any network layer. Since the area of the input feature map is larger than the area of the convolutional kernel, the scanning window of a convolutional kernel needs to slide multiple times on the input feature map. Therefore, a convolutional kernel can correspond to multiple groups of input data. Since a convolutional kernel can include multiple weights, then, one weight corresponds to one input data in a group of input data. Each group of input data can be regarded as a sub-feature map.
[0160] To further illustrate the correspondence between the convolutional kernel and the sub-feature map, refer to Figure 15 the schematic diagram of a convolutional calculation process provided by an embodiment of the present invention as shown in Figure 15 In which, the input feature map has 2 layers, there are 2 convolutional kernels. Convolutional kernel 1 slides 4 times on the first layer of the input feature map, and thus 4 sub-feature maps corresponding to convolutional kernel 1 can be obtained, which are sub-features Figures 1 - 4 , convolutional kernel 2 slides 4 times on the second layer of the feature map, and thus 4 sub-feature maps corresponding to convolutional kernel 2 can also be obtained, which are sub-feature maps A - D.
[0161] 1402. The neural network computing device receives the convolution kernels of the first network layer of the convolutional neural network and the input data to be subjected to convolution calculation.
[0162] The input data of the first network layer can be the data to be processed that is first input into the neural network computing device, or the output data of other network layers. The input data of other network layers can be the data output by the previous layer of that other network layer, or the data to be processed that is first input into the neural network computing device.
[0163] Since the convolution kernels of each layer are available in the neural network computing device, the neural network computing device can directly obtain multiple groups of input data to be subjected to convolution calculation according to the positional relationship between the convolution kernels of each layer and the output data of the previous layer. One convolution kernel can correspond to at least one group of input data to be subjected to convolution calculation.
[0164] 1403. The neural network computing device stores the convolution kernel in the first-row dynamic random access memory (DRAM) cells of the first memory array, where the first memory array includes multiple rows and multiple columns of DRAM cells.
[0165] 1404. The neural network computing device stores the input data in at least two rows of DRAM cells in the first memory array except the first-row DRAM cells.
[0166] Since each convolution kernel can correspond to at least one sub-feature map (i.e., one group of input data), there is a corresponding relationship between multiple sub-feature maps and at least one convolution kernel. Therefore, the neural network computing device can directly associate and store the multiple sub-feature maps with the at least one convolution kernel according to this corresponding relationship.
[0167] To ensure the corresponding relationship between the convolution kernel and the sub-feature map, the DRAM cells storing any convolution kernel and the sub-feature map corresponding to any convolution kernel are located in the same target column in the first memory array, so that the input data in the sub-feature map can correspond to the weights in the convolution kernel. To further illustrate the positional relationship between the DRAM cells, see Figure 16 , Figure 16 is a schematic diagram of deploying computing data provided by an embodiment of the present invention. From Figure 16 the first memory array in, it can be seen that convolution kernel 1 is stored in the 9th row and the A target column of the first memory array. The coordinates of convolution kernel 1 on the first memory array can be expressed as (9, A). The weight W0 of convolution kernel 1 is stored in the A1 sub-column of the 9th row. That is, the coordinates of weight W0 on the first memory array can be expressed as (9, A1). Convolution kernel 2 is stored in the 10th row of the first memory array. Convolution kernel 1-2, sub-feature Figures 1 - 4, Sub - feature maps A - D are simultaneously located in column A of the first storage array. By activating the data of two rows and a target column, the correspondence between the convolutional kernel and the feature map is achieved. For example, when activating the data of row 9 and row 1 in the A - th target column, convolutional kernel 1 and the sub - feature Figure 1 will correspond. It can also be regarded as the scanning window of convolutional kernel 1 sliding once and covering the sub - feature Figure 1 . Then the neural network computing device can perform convolutional calculations based on the data of row 9 and row 1 in the A - th target column.
[0168] Specifically, the neural network computing device can store any data in any memory address in the first storage array by performing a write operation on that memory address. For example, when storing weight W in the unit corresponding to memory address xxxx, the neural network computing device can charge the capacitor in the DRAM unit corresponding to memory address xxxx through a write operation. The amount of charge in the capacitor can represent weight W, and thus weight W can be stored in the unit corresponding to memory address xxxx. Since the weights in the convolutional kernel and the input data to be convolved are represented by quantized data, the quantized data may be 1 bit or multiple bits, and each DRAM unit corresponding to each memory address in the first storage array can only store 1 - bit data. Therefore, multiple adjacent DRAM units in the first storage array can be used to store a weight or an input data. For example, Figure 10 W0, W1, D0_0, D0_1, D1_0, and D1_1 in it are all 1 - bit data. Among them, W0 and W1 are two weights in the convolutional kernel, D0_0 and D0_1 can represent D0, D1_0 and D1_1 can represent D1, and D0 is the feature element in the sub - feature map corresponding to W0, D1 is the feature element in the sub - feature map corresponding to W1. Therefore, multiple DRAM units can be used to store a weight or an input data.
[0169] It should be noted that before using the first storage array to store data, it is necessary to first query the target list stored in the neural network computing device. This target list is used to store the fault conditions of each DRAM unit in the first storage array. When the target list queries that any DRAM unit is marked as faulty, then that any DRAM unit cannot be used to store data. When the target list queries that any DRAM unit is marked as fault - free, then that any DRAM unit is used to store data.
[0170] The DRAM cells in the first storage array can be verified and marked through the target list, which can avoid storing weights or input data in faulty DRAM cells, so as to avoid being unable to read weights or input data from the first storage array, and further avoid errors in subsequent calculation processes, thereby ensuring that the results output by the neural network computing device are more accurate.
[0171] 1405. The neural network computing device performs multiple convolution operations on the input data and the convolution kernel based on the first computing circuit connected to the first storage array. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of DRAM cells among at least two rows of DRAM cells.
[0172] In some embodiments, the first computing circuit can also implement one convolution calculation through a latch, a multiplier, and an adder. In a possible implementation manner, this step 1405 can be implemented through the process shown in the following steps 51-53.
[0173] Step 51. The neural network computing device stores the convolution kernel stored in the storage array in multiple latches in the first computing circuit. Each latch is used to cache one element of the convolution kernel stored in the DRAM cells of the corresponding column. Each DRAM cell in the first row of DRAM cells is respectively used to store one element of the convolution kernel.
[0174] Specifically, one row of DRAM cells storing the convolution kernel in the first storage array can be opened. Since each latch is connected to a column of DRAM cells. Therefore, by turning on the switch of the latch memory, 1-bit values of each element in the convolution kernel can be stored on the latches in each column.
[0175] Step 52. The neural network computing device performs a multiplication operation on the convolution kernel cached by the multiple latches and the input data stored in one row of DRAM cells among the multiple rows of DRAM cells based on the multiple multipliers in the first computing circuit.
[0176] Correspondingly, an exclusive OR operation can be performed on the elements in the convolution kernel and a set of input data. Based on the result of the exclusive OR operation and the weight value, the product of the input data stored in one row of DRAM cells among the multiple rows of DRAM cells is obtained.
[0177] In a possible implementation manner, this step 52 can be implemented through the process shown in the following steps 1-2.
[0178] Step 1. Based on at least one exclusive-OR unit in the multiplier, the neural network computing device performs an exclusive-OR operation on a single-bit value of an element in a convolution kernel cached in a latch and a single-bit value of an element in the input data stored in one of at least two DRAM cells, where each of the at least two DRAM cells is respectively used to store a single-bit value of an element in the input data for the convolution calculation to be performed.
[0179] Specifically, for any column in the first storage array, the 1-bit value of the target element stored in the latch on that column is input to the exclusive-OR unit of the multiplier, and the 1-bit value of the input data corresponding to the target element stored on that column is input to the exclusive-OR unit of the multiplier. The exclusive-OR unit can perform an exclusive-OR operation on the 1-bit value of the target element and the 1-bit value of the input data, and output the second-highest bit of the first data. Since there may be multiple weights in the convolution kernel, the above exclusive-OR process needs to be performed for each weight. Therefore, the neural network computing device needs to perform the above operation on each weight of the convolution kernel stored in the first storage array to obtain the second-highest bits of multiple first data.
[0180] Step 2. The neural network computing device is based on a combination unit in the multiplier, and the combination unit is used to combine the data output by the at least one exclusive-OR unit and the data output by at least one latch.
[0181] The data output by the exclusive-OR unit is also the second-highest bit of the first data, and the data output by the combination unit is also a first data. Specifically, for an element in the convolution kernel, the data output by the at least one exclusive-OR unit is input to the combination unit, and each 1-bit data of the element stored in the latch is stored and input to the combination unit. The combination unit combines these input data into a first data.
[0182] Since each convolution kernel corresponds to multiple input data, multiple first data will be obtained. Based on Figure 16 the data in, corresponding to convolution kernel 1, the multiple first data include D0*W0, D1*W1, D3*W2, and D4*W3.
[0183] By splitting the multiplication operation in the convolution calculation into exclusive-OR operations, the process of splitting the convolution calculation into many basic calculation logics is avoided, and the power consumption during the implementation of the convolution calculation is further reduced.
[0184] Step 53. The neural network computing device performs an addition operation on the output results of the multiple multipliers based on multiple adders in the first computing circuit, and the output of the multiple adders is the result of one convolution calculation.
[0185] Based on the example in step 52, input D0*W0, D1*W1, D3*W2, and D4*W3 into multiple adders. One of the multiple adders can output p0 = D0*W0 + D1*W1 + D3*W2 + D4*W3, which is the result of one convolution calculation, that is, the output of one feature point.
[0186] The summation of multiple first data by multiple adders can be regarded as multi-part summation. P0 is the result of multi-part summation. Since one convolution kernel corresponds to multiple groups of input data, step 1405 is performed on each group of data corresponding to the convolution kernel, so that multiple results of multi-part summation can be obtained. The multiple results of multi-part summation can form a convolution feature map. When there are multiple convolution kernels, the neural network computing device performs step 1405 on these multiple convolution kernels to obtain multiple convolution feature maps. For example, Figure 15 In the part of the multi-part summation, based on convolution kernel 1, a convolution feature is obtained Figure 1 and based on convolution kernel 2, a convolution feature is obtained Figure 2 .
[0187] For the process of performing step 1405 on each group of input data corresponding to one convolution kernel, multiple feature points can be obtained, and these multiple feature points are the results of multiple convolutions of one convolution kernel. When the number of convolution kernels is greater than 1, the above step 1405 can be performed on each convolution kernel to obtain the results of multiple convolution calculations for each convolution kernel.
[0188] 1406. In a single summation calculation process, the neural network computing device sums the results of the one convolution calculation output by the multiple adders based on the summation circuit in the second computing circuit of the neural network computing device.
[0189] The summation circuit can perform multiple summations, and can sum the result of this convolution calculation with the results of previous multiple convolution calculations. Specifically, the neural network computing device inputs the result of each convolution calculation into the summation circuit, and the summation circuit accumulates according to the result of the previous layer's summation in the summation circuit, so as to sum the results of multiple convolution calculations, thereby realizing multiple summations to obtain the final convolution result. Taking the Figure 15 summation part as an example, the neural network computing device sums the convolution feature Figure 1 and the convolution feature Figure 2 to obtain the matrix of the first feature map [R0, R1; R2, R3].
[0190] 1407. The neural network computing device caches the calculation results of the summation circuit during multiple summation calculation processes based on the buffer in the neural network computing device.
[0191] The calculation result in the multiple summation calculation process is also the final convolution result.
[0192] 1408. Based on the pooling circuit in the neural network computing device, the neural network computing device performs a pooling operation on the calculation result in the multiple summation calculation process cached in the buffer.
[0193] The neural network computing device inputs the final convolution result into multiple comparators of the pooling circuit, and the multiple comparators perform cyclic comparison, so as to output multiple representative feature points, and these multiple representative feature points are also the results of the pooling operation. For example, Figure 15 In the pooling part in, the neural network device inputs the first feature map [R0, R1; R2, R3] into the pooling circuit, and the pooling circuit performs a pooling operation on the first feature map [R0, R1; R2, R3] and outputs the pooling result D0. The specific comparison process of the comparator has been described above and will not be elaborated here.
[0194] 1409. Based on the selection circuit in the neural network computing device, the neural network computing device selects and outputs the output result of the pooling circuit or the output result of the summation circuit.
[0195] In some embodiments, the neural network computing device may not perform the process shown in steps 1407-1408, but may directly perform an activation operation on the final convolution result (step 1410). In some embodiments, the process shown in steps 1407-1408 needs to be performed. In order to enable the neural network computing device to meet different computing processes, the output result of the summation circuit and the output result of the pooling circuit can be input into the selection circuit, so that the selection circuit can select the input of the activation operation according to preset conditions. The selection process of the selection circuit has been described above and will not be elaborated here.
[0196] 1410. Based on the activation circuit in the neural network computing device, the neural network computing device performs an activation operation on the output result of the selection circuit, and uses the output result of the activation circuit as the input data for the next network layer to perform convolution calculation in the first network layer. When the first network layer is the last network layer of the convolutional neural network, the neural network computing device outputs the output result of the activation circuit.
[0197] It should be noted that since the convolutional neural network has been mapped into the neural network computing device at this time, when it is necessary to calculate the input feature map based on the target convolutional neural network, since the target convolutional neural network is not the same neural network as the above-mentioned convolutional neural network, it is necessary to remap the target convolutional neural network into the neural network computing device so that the computing process in the target convolutional neural network can be implemented in the neural network computing device.
[0198] In a possible implementation, when the neural network computing device implements the computing process in the convolutional neural network based on other input feature maps, for the first network layer of the convolutional neural network, it receives the input data to be subjected to convolutional calculation in the first network layer of the convolutional neural network, and performs the steps of the above-mentioned neural network calculation; when the neural network computing device implements the computing process in the target convolutional neural network based on other input feature maps, for the first network layer of the target convolutional neural network, it receives the convolutional kernels of each network layer of the target convolutional neural network and the input data to be subjected to convolutional calculation in the first network layer of the target convolutional neural network, and performs the steps of the above-mentioned neural network calculation.
[0199] The method provided by the embodiments of the present invention stores the convolutional kernels and the input data to be subjected to convolutional calculation in the first storage array of the memory. The first computing circuit can read the convolutional kernels and the input data to be subjected to convolutional calculation from the first storage array and complete the convolutional calculation, thereby implementing neural network calculation in the memory, reducing the transmission of neural network data between the memory and the neural network computing node, and solving the problem of the memory wall. Moreover, by splitting the multiplication operation in the convolutional calculation into exclusive OR operations, it is avoided to split the process of convolutional calculation into many basic calculation logics, further reducing the power consumption when implementing convolutional calculation. Furthermore, the calculation processes such as convolution, pooling, and activation in the convolutional neural network can all be completed within the neural network computing device, so that all the calculation processes in the neural network can be implemented in the memory, and the transmission of neural network data between the memory and the neural network computing node can be further reduced.
[0200] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated herein one by one.
[0201] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0202] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A neural network computing device, characterized in that, Comprising: A first storage array, including multiple rows and columns of dynamic random access memory (DRAM) cells. The first row of DRAM cells in the storage array is used to store convolution kernels, and at least two rows of DRAM cells in the storage array are used to store input data for performing convolution calculations. The input data stored in each row of the two rows of DRAM cells includes data on a sub-feature map in an input feature map. Moreover, an element in the convolution kernel and the corresponding element on the sub-feature map are stored in the same column of DRAM cells. The sub-feature map is the area covered by the scanning window of the convolution kernel on the feature map. A first calculation circuit, connected to the storage array, for performing multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of the at least two rows of DRAM cells.
2. The computing device according to claim 1, wherein The first calculation circuit includes: Multiple latches, with the input end of each latch connected to a column of DRAM cells, and the output end of each latch connected to one of multiple multipliers. Each latch is used to cache a one-bit value of an element in the convolution kernel stored in the corresponding column of DRAM cells. Among them, each DRAM cell in the first row of DRAM cells is respectively used to store a one-bit value of an element in the convolution kernel. Multiple multipliers, with the first input end of each multiplier respectively connected to a column of DRAM cells, and the second input end of each multiplier respectively connected to the output end of a latch. Multiple adders, with the input end of each adder respectively connected to at least two multipliers. Among them, during one convolution calculation process, the multiple multipliers are respectively used to perform multiplication operations on the convolution kernel cached by the multiple latches and the input data stored in one row of the multiple rows of DRAM cells. The multiple adders are respectively used to perform addition operations on the output results of the multiple multipliers. The output of the multiple adders is the result of one convolution calculation.
3. The computing device according to claim 2, wherein It further includes: A second calculation circuit, the second calculation circuit is connected to the first calculation circuit. The second calculation circuit includes a summation circuit, which is used to sum the results of the one convolution calculation output by the multiple adders in the first calculation circuit during one summation calculation process.
4. The computing device according to claim 3, wherein The second calculation circuit is used to perform multiple summation calculations. The second calculation circuit further includes: A buffer, connected to the summation circuit, for caching the calculation results of the summation circuit during multiple summation calculation processes. A pooling circuit, the pooling circuit is connected to the buffer, for performing a pooling operation on the calculation results cached in the buffer during the multiple summation calculation processes.
5. The computing device according to claim 4, wherein The second calculation circuit further includes: A selection circuit, with the first input end of the selection circuit used to connect to the output end of the pooling circuit, and the second input end of the selection circuit used to connect to the output end of the summation circuit. The selection circuit is used to selectively output the output result of the pooling circuit or the output result of the summation circuit. An activation circuit, connected to the output end of the selection circuit, is configured to perform an activation operation on the output result of the selection circuit and input the output result to a second storage array, where the second storage array includes multiple rows and multiple columns of DRAM cells.
6. The computing device according to claim 2, wherein Each of the multipliers includes: At least one exclusive OR unit, where a first input end of each exclusive OR unit is connected to a latch, a second input end of each exclusive OR unit is connected to a column of DRAM cells, and each exclusive OR unit is configured to perform an exclusive OR operation on a bit value of an element in a convolution kernel cached by a latch and a bit value of an element in the input data stored in one of the at least two rows of DRAM cells, where each of the at least two rows of DRAM cells is respectively configured to store a bit value of an element in the input data to be subjected to the convolution calculation; A combination unit, where a first input end of the combination unit is connected to the output end of the at least one exclusive OR unit, and a second input end of the combination unit is connected to at least one latch, and the combination unit is configured to combine the data output by the at least one exclusive OR unit and the data output by at least one latch.
7. A computing device, characterized in that, It includes: A controller, configured to receive a convolution kernel and input data to be subjected to the convolution calculation, and send the convolution kernel and the input data to a memory connected to the controller; The memory includes: A first storage array, including multiple rows and multiple columns of dynamic random access memory (DRAM) cells, configured to store the convolution kernel and the input data sent by the controller, where the first row of DRAM cells in the storage array is configured to store the convolution kernel, at least two rows of DRAM cells in the storage array are configured to store the input data to be subjected to the convolution calculation, the input data stored in each row of the two rows of DRAM cells includes data on a sub-feature map in the input feature map, and an element in the convolution kernel and the corresponding element on the sub-feature map are stored in the same column of DRAM cells, and the sub-feature map is the area covered by the scanning window of the convolution kernel on the feature map; and A first calculation circuit, connected to the storage array, configured to perform multiple convolution operations on the input data and the convolution kernel, and each convolution calculation is configured to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one of the at least two rows of DRAM cells.
8. The computing device according to claim 7, wherein The first calculation circuit includes: Multiple latches, where an input end of each latch is connected to a column of DRAM cells, an output end of each latch is connected to one of the multiple multipliers, and each latch is configured to cache a bit value of an element in the convolution kernel stored in the corresponding column of DRAM cells, where each of the DRAM cells in the first row is respectively configured to store a bit value of an element in the convolution kernel; Multiple multipliers, where a first input end of each multiplier is respectively connected to a column of DRAM cells, and a second input end of each multiplier is respectively connected to an output end of a latch; Multiple adders, where an input end of each adder is respectively connected to at least two multipliers; Among them, during a convolution calculation process, the multiple multipliers are respectively used to perform multiplication operations on the convolution kernels cached by the multiple latches and the input data stored in one row of DRAM units among the multiple rows of DRAM units, and the multiple adders are respectively used to perform addition operations on the output results of the multiple multipliers, and the output of the multiple adders is the result of one convolution calculation.
9. The computing device according to claim 8, wherein The memory further includes: A second calculation circuit, the second calculation circuit is connected to the first calculation circuit, and the second calculation circuit includes a summation circuit, which is used to sum the results of the one convolution calculation output by the multiple adders in the first calculation circuit during a summation calculation process.
10. The computing device according to claim 9, wherein The second calculation circuit is used to perform multiple summation calculations, and the second calculation circuit further includes: A buffer, connected to the summation circuit, for buffering the calculation results of the summation circuit during multiple summation calculation processes; A pooling circuit, the pooling circuit is connected to the buffer, and is used to perform a pooling operation on the calculation results of the multiple summation calculation processes cached in the buffer.
11. The computing device according to claim 10, wherein The second calculation circuit further includes: A selection circuit, the first input end of the selection circuit is used to connect to the output end of the pooling circuit, the second input end of the selection circuit is used to connect to the output end of the summation circuit, and the selection circuit is used to select and output the output result of the pooling circuit or the output result of the summation circuit; An activation circuit, connected to the output end of the selection circuit, for performing an activation operation on the output result of the selection circuit and inputting the output result to a second storage array, and the second storage array includes multiple rows and multiple columns of DRAM units.
12. The computing device according to claim 8, wherein Each multiplier includes: At least one exclusive-OR unit, the first input end of each exclusive-OR unit is connected to a latch, the second input end of each exclusive-OR unit is connected to a column of DRAM units, and each exclusive-OR unit is used to perform an exclusive-OR operation on one-bit value of an element in a convolution kernel cached by a latch and one-bit value of an element in the input data stored in one row of the at least two rows of DRAM units, wherein each of the at least two rows of DRAM units is respectively used to store one-bit value of an element in the input data to be subjected to the convolution calculation; A combination unit, the first input end of the combination unit is connected to the output end of the at least one exclusive-OR unit, the second input end of the combination unit is connected to at least one latch, and the combination unit is used to combine the data output by the at least one exclusive-OR unit and the data output by at least one latch.
13. A neural network computing method, characterized in that, The method is executed by a neural network computing device and includes: Receiving a convolution kernel of any network layer of a convolutional neural network and input data to be subjected to a convolution calculation; Storing the convolution kernel in the first-row dynamic random access memory (DRAM) units of a first storage array, wherein the first storage array includes multiple rows and multiple columns of DRAM units; Store the input data in at least two rows of DRAM cells in the first storage array except for the first row of DRAM cells. The input data stored in each row of the two rows of DRAM cells includes the data on a sub-feature map in the input feature map. Moreover, an element in the convolution kernel and the corresponding element on the sub-feature map are stored in the same column of DRAM cells. The sub-feature map is the area covered by the scanning window of the convolution kernel on the feature map. Based on the first computing circuit connected to the first storage array, perform multiple convolution operations on the input data and the convolution kernel. Each convolution calculation is used to perform a multiply-accumulate calculation on the convolution kernel and the input data stored in one row of the at least two rows of DRAM cells.
14. The method according to claim 13, wherein The performing multiple convolution operations on the input data and the convolution kernel based on the first computing circuit connected to the first storage array includes: Store the convolution kernel stored in the storage array in multiple latches in the first computing circuit. Each latch is used to cache an element in the convolution kernel stored in the DRAM cells of the corresponding column. Each DRAM cell in the first row of DRAM cells is respectively used to store an element in the convolution kernel. Based on multiple multipliers in the first computing circuit, perform a multiplication operation on the convolution kernel cached by the multiple latches and the input data stored in one row of the multiple rows of DRAM cells. Based on multiple adders in the first computing circuit, perform an addition operation on the output results of the multiple multipliers. The output of the multiple adders is the result of one convolution calculation.
15. The method according to claim 14, wherein The method further includes: During one summation calculation process, based on the summation circuit in the second computing circuit in the neural network computing device, sum the results of the one convolution calculation output by the multiple adders.
Citation Information
Patent Citations
Binary convolutional neural network processor and using method thereof
CN107153873A
Data processing method, device and apparatus
CN108573305A
FD-SOI process-based calculation accelerator in binary convolutional neural network memory
CN109784483A