Convolutional neural network accelerator based on lookup table and in-memory calculation
By using a method based on lookup table and in-memory computing in the convolutional neural network accelerator, serial input and compression of eigenvalue data streams is solved, and the addition tree hardware complexity is achieved to reduce power consumption and improve area benefits.
Patent Information
- Application Number
- CN202510281269.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
When existing convolutional neural network accelerators improve the calculation parallelism, the area and power consumption of the adder tree increase exponentially, resulting in excessive hardware complexity and difficulty in balancing hardware complexity and power consumption.
A convolutional neural network accelerator based on lookup tables and in-memory calculations is adopted to reduce the number of adder trees and reduce the overall hardware complexity by serial input and compression of the eigenvalue data stream by using matrix multiplication operations in the form of lookup tables.
It effectively reduces the flip probability of the accelerator calculation circuit, reduces overall power consumption, and improves area benefits by optimizing the area of the lookup table, and supports high sparseness calculation of the unstructured pruning compression model.
Smart Images

Figure CN120218147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a convolutional neural network accelerator based on a lookup table and in-memory computing. Background Art
[0002] In deep learning accelerators, data frequently interacts between the on-chip main memory (Global Buffer) and the processing element (PE) array. The power consumption and latency caused by data interaction also account for a larger proportion in the overall design of the accelerator as the scale of the deep learning network expands, thus giving rise to the problem of the memory wall. To solve the problem of reduced energy efficiency of the accelerator caused by the memory wall, deep learning accelerators based on in-memory computing (CIM) have emerged.
[0003] The types of in-memory computing structures can be divided into two major categories, volatile and non-volatile, according to the type of memory. Among them, non-volatile storage devices mainly include phase change memory (PCM), resistive random access memory (RRAM), magnetic random access memory (MRAM), and Flash, etc. Volatile memory is mainly represented by static random access memory (SRAM) and dynamic random access memory (DRAM). Thanks to mature process conditions, compared with other types of in-memory computing chips, in-memory computing chips based on SRAM have more stable and controllable performance, and the mass production difficulty is also the lowest.
[0004] Today's mainstream SRAM in-memory computing structures can be divided into two major categories, namely, analog in-memory computing structures (Analog-SRAM CIM) that use analog signals to complete multiplication and addition calculations and digital in-memory computing structures (Digital-SRAM CIM) that use digital signals to complete multiplication and addition calculations. In Analog-SRAM CIM, the final result of the multiplication and addition operation usually requires a high-resolution analog-to-digital converter (ADC) to convert the analog signal obtained by SRAM calculation into a digital signal. However, the higher the resolution of the ADC, the greater the area and power consumption overhead. Therefore, in Analog-SRAM CIM, a full-precision ADC is often not used, but a compromise is made between the overall performance and the calculation accuracy.
[0005] Digital-SRAM CIM performs multiplication and addition calculations using pure digital logic. Therefore, Digital-SRAM CIM does not introduce errors during matrix multiplication. In addition, Digital-SRAM CIM is more stable than Analog-SRAM CIM, less affected by process variations, and it is easier to adjust the overall performance of CIM by controlling voltage and frequency. Although deep learning accelerators based on Digital-SRAM CIM have the above advantages, designing high-performance Digital SRAM-CIM deep learning accelerators also faces many challenges.
[0006] Digital SRAM-CIM mainly consists of two parts: a storage module and a computing module. For deep learning algorithms such as convolutional neural networks, the computing part of CIM mainly consists of a large number of adder trees. Once the computing parallelism of CIM is increased, the area of the adder tree will increase exponentially, becoming the main overhead of the area and power consumption of CIM. How to reduce the hardware complexity of the adder tree is an important issue to be solved when designing a high-performance CIM accelerator.
[0007] Most current convolutional neural network accelerators adopt a fixed-weight-to-CIM computing structure. In this structure, the weights are pre-loaded into CIM and remain unchanged for a long time. Then, by inputting the feature values into CIM, the convolution operation of the weights and feature values is realized. Common feature value input data streams can be divided into two categories: serial input and parallel input. The serial input data stream of feature values splits multi-bit feature values into multiple single bits and inputs them into the CIM array in sequence according to time order. The advantage of this feature value data stream structure is its simplicity and the ability to support matrix multiplication calculations with variable bit widths. The disadvantage is that the calculations are relatively independent of each other, so the optimization space is limited. The parallel input data stream of feature values inputs multi-bit feature values into CIM simultaneously, so the calculation time can be effectively shortened. However, the cost is that the hardware complexity is much higher than that of serial input because multi-bit multipliers must be used to complete the multiplication operation of multi-bit data. Moreover, the larger the bit width of the feature values, the more complex the structure of the multiplier part. Therefore, how to select an appropriate data stream in combination with the characteristics of the accelerator hardware design and balance the relationship between the hardware complexity and power consumption of Digital SRAM-CIM is also an issue that needs to be carefully considered when designing the accelerator.
[0008] In addition, deep learning models are highly redundant. Therefore, when designing hardware accelerators, a large number of model compression schemes are applied to model training, resulting in a sparse model. Because of the regular structure of SRAM CIM, its basic circuit can be regarded as an M×N SRAM array, and the weights are fixed on the storage nodes. To improve the computing efficiency, the data fixed on the storage nodes remain unchanged for a long time, which means that the computing cycle is often much longer than the writing cycle. Due to the above characteristics, structured pruning is very friendly to SRAM CIM. However, it is difficult to achieve a high sparsity with structured pruning. Therefore, although it can simplify the hardware design, the performance improvement of SRAM CIM by introducing structured pruning is very limited. On the contrary, unstructured pruning can relatively easily make the model reach a high sparsity, but it is extremely unfriendly to the hardware design of SRAM CIM. So there has never been a good solution for the application of SRAM CIM in unstructured pruning models. Thus, designing an SRAM CIM that can effectively utilize the high sparsity brought by unstructured pruning without increasing the hardware complexity is also an important research direction. Summary of the Invention
[0009] The purpose of this application is to propose a convolutional neural network accelerator based on a lookup table and in-memory computing for the above-mentioned technical problems.
[0010] The present invention provides a convolutional neural network accelerator based on a lookup table and in-memory computing, which includes an on-chip main memory and a plurality of arithmetic processing units. The arithmetic processing unit includes a data preprocessing unit, an in-memory computing array unit, a first adder, and an accumulator connected in sequence. During the calculation process of the convolutional neural network, the eigenvalue data stream used is serially input based on bit eigenvalues and the input eigenvalue sequence is compressed to obtain a compressed eigenvalue sequence. The weights used in the calculation process of the convolutional neural network are compressed to obtain compressed weights. The on-chip main memory is used to buffer the compressed eigenvalue sequence and the compressed weights and transmit them to the arithmetic processing unit. The data preprocessing unit is used to decompress the compressed eigenvalue sequence and the compressed weights, and classify the bit eigenvalues corresponding to the decompressed eigenvalue sequence into sparse-type bit eigenvalues and dense-type bit eigenvalues. The sparse-type bit eigenvalues and the dense-type bit eigenvalues are interleaved and input into the in-memory computing array unit according to the pattern of dense type - sparse type - dense type and perform convolution calculation with the decompressed weights. For the sparse-type bit eigenvalues, the in-memory computing array unit will perform skip-zero calculation. For the dense-type bit eigenvalues, the in-memory computing array unit will not perform skip-zero calculation. The in-memory computing array unit includes a plurality of in-memory computing macro modules arranged in an array, and performs matrix multiplication operations in the form of a lookup table through the in-memory computing macro modules to obtain a first partial sum result. The first adder is used to add the first partial sum results output by the in-memory computing macro modules in each column of the in-memory computing array unit to obtain a second partial sum result. The accumulator is used to calculate the second partial sum result to obtain the multiply-accumulate result corresponding to the eigenvalue data stream.
[0011] Preferably, the process of serially inputting the eigenvalue data stream based on bit eigenvalues is as follows:
[0012] The first eigenvalue and the second eigenvalue with a bit width of N bits in the eigenvalue data stream are respectively divided into a first bit eigenvalue sequence and a second bit eigenvalue sequence, and both the first bit eigenvalue sequence and the second bit eigenvalue sequence contain N bit eigenvalues;
[0013] The bit eigenvalues in the first bit eigenvalue sequence and the bit eigenvalues in the second bit eigenvalue sequence are arranged in the same order and alternately input from low to high according to the order to form an eigenvalue sequence.
[0014] Preferably, the compression method for weights is binary Huffman coding; the compression method for the eigenvalue sequence is two-stage compression. The first-stage compression method is binary Huffman coding to obtain the first-stage compressed bit eigenvalues with the flag bits being 0 or 1. For the first-stage compressed bit eigenvalues with the flag bit being 1 after binary Huffman coding, the second-stage compression is performed. The second-stage compression method is to directly divide according to different channel dimensions or first divide into blocks and then group the bit eigenvalues on the same channel of all blocks.
[0015] Preferably, the data preprocessing unit includes a decompression module, an eigenvalue block division module, and a sparse engine module. The decompression module is used to decompress the compressed eigenvalue sequence and the compressed weights. The eigenvalue block division module is used to divide the decompressed eigenvalue sequence into blocks. The sparse engine module is used to classify the bit eigenvalues corresponding to the decompressed eigenvalue sequence.
[0016] Preferably, the in-memory computing macro module includes a weight preprocessing module, an encoder, a lookup table array, and an adder tree. The weight preprocessing module is used to perform addition and subtraction operations on the decompressed weights to obtain the weight operation results, and accumulate the decompressed weights to obtain the weight accumulation results, and input the weight operation results and the weight accumulation results into the lookup table array and the encoder respectively. The lookup table array includes a number of lookup table units arranged in an array. The adder tree sums up all the outputs of the lookup table units to obtain the first partial sum result.
[0017] Preferably, the in-memory computing macro module further includes an encoder. The encoder is used to encode the decompressed eigenvalue sequence from the binary data form into the signed weight coding form. The types of the lookup table units are divided into the first lookup table and the second lookup table. The first lookup table is used to determine, according to the input control signal, to perform an unsigned operation or an inversion operation in the signed operation on the first bit eigenvalue and the second bit eigenvalue input to the lookup table array in the binary data form. The second lookup table is used to perform a signed operation on the third bit eigenvalue and the fourth bit eigenvalue in the signed weight coding form output by the encoder and complete the inversion operation based on the signed weight. The encoder includes a first register, a first inverter, and a first multiplexer. The decompressed eigenvalue sequence is input to the lookup table array bit by bit in 8 cycles. In the 9th cycle, the bit eigenvalue input in the 8th cycle stored in the first register is selected through the first multiplexer, passed through the first inverter to obtain the corresponding inversion result, and then output to the lookup table array with the second lookup table.
[0018] Preferably, the first lookup table includes 2 input logic AND gates, 1 input logic XNOR gate, 1 signal input logic gate, three rows of memory cells, 1 second inverter, 1 second multiplexer, and 9 output logic gates. The three rows of memory cells store W0, W1, and W0+W1 respectively. The second multiplexer selects, according to the control signal input by the signal input logic gate, to directly output the output of one of the memory cells in the three rows of memory cells or to output the output of one of the memory cells in the three rows of memory cells after inverting it through the second inverter; the second lookup table includes 1 input logic XOR gate, 1 input logic XNOR gate, two rows of memory cells, and 1 output logic XNOR gate. The two rows of memory cells store W0-W1 and W0+W1 respectively. The input of the output logic XNOR gate is the third eigenvalue and the output of one of the memory cells in the two rows of memory cells. Whether to invert the output of one of the memory cells in the two rows of memory cells is determined according to the third eigenvalue.
[0019] Preferably, a decoder is further included. The decoder is used to decode the multiplication-accumulation result and the weight-accumulation result corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator to obtain a decoding result; the decoder includes a subtractor and a first shifter connected in sequence. The multiplication-accumulation result corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator and the weight-accumulation result are input into the subtractor to perform a subtraction operation to obtain a subtraction result. The operation of dividing the subtraction result by two is completed through the first shifter to obtain a decoding result. The decoding result or the multiplication-accumulation result corresponding to the eigenvalue data stream obtained by using the first lookup table and passing through the adder tree and the accumulator is used as the convolution calculation result.
[0020] Preferably, the accumulator includes a second adder, a sign adder, a second shifter, and a second register. The sign adder is used to accumulate the number of bit eigenvalues representing negative numbers in the decompressed eigenvalue sequence in signed operations to obtain an accumulation result, and input the accumulation result into the second adder to add it to all partial sum results corresponding to the decompressed eigenvalue sequence to obtain an addition result; the second adder, the second shifter, and the second register form an accumulation structure, and the accumulation structure is used to perform shift accumulation on all addition results corresponding to the eigenvalue data stream.
[0021] Preferably, the second register adopts a partial sum storage structure that supports eigenvalue skip-zero calculation for structured pruning compression models; the partial sum storage structure includes a sparse output cache block and a dense output cache block. The second partial sum result output after the first dense-type bit eigenvalue passes through the in-memory computing array unit is written into the dense output cache block, and the position information of the second partial sum result output after the first sparse-type bit eigenvalue passes through the in-memory computing array unit is written into the sparse output cache block; before the second partial sum result output after the next dense-type bit eigenvalue passes through the in-memory computing array unit is written into the dense output cache block, it will be added to the second partial sum results with the same position information in the sparse output cache block and the dense output cache block, and then written into the dense output cache block; the next sparse-type bit eigenvalue continues to be written into the sparse output cache block, and the above steps are repeated until the second partial sum results corresponding to all sparse-type and dense-type bit eigenvalues with the same position information have completed shift addition.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] (1) The convolutional neural network accelerator based on lookup table and in-memory computing proposed by the present invention proposes a brand-new serial data stream based on bit eigenvalues. By taking advantage of the similarity of adjacent eigenvalue values, this data stream effectively reduces the switching probability of the accelerator's computing circuit through reasonable redistribution of the eigenvalue data stream, thereby reducing the overall power consumption of the accelerator. In addition, in the convolutional neural network accelerator proposed by the present invention, the weights adopt a binary Huffman coding method. To increase the reuse of the on-chip decoding circuit and avoid designing an additional decoding circuit for eigenvalue decompression, the eigenvalues also adopt the same decoding method. Therefore, the present invention also proposes a brand-new eigenvalue compression scheme to specifically compress this data stream, effectively improving the transmission efficiency of eigenvalues and saving the hardware overhead of the decoding circuit part.
[0024] (2) The convolutional neural network accelerator based on lookup table and in-memory computing proposed by the present invention designs a brand-new in-memory computing array unit. This in-memory computing array unit adopts the form of a lookup table and uses read operations to replace the multiply-accumulate operations in partial matrix multiplications. This effectively reduces the overall hardware complexity of the CIM, significantly reducing the area and power consumption of the in-memory computing array unit. Secondly, the present invention further optimizes the area of the lookup table by introducing an encoder. Finally, through area analysis and modeling of the traditional circuit architecture and the lookup table-based circuit architecture, based on the UMC55nm process, when the weight accuracy is 8, choosing a lookup table with a parallelism of 2, where the lookup table calculates the multiply-accumulate results of two inputs and two weights at a time, has the highest area efficiency.
[0025] (3) The convolutional neural network accelerator based on look-up table and in-memory computing proposed by the present invention also supports the zero-skipping calculation of eigenvalues for unstructured pruning compressed models. Combining the characteristics of the CIM accelerator, a partial sum storage structure with low hardware complexity is designed, effectively solving the problem of increased memory hardware complexity caused by zero-skipping calculation, with high area efficiency and being very effective in processing sparse features. Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 Schematic diagram of the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0028] Figure 2 Schematic diagram of dividing the feature map into 8-bit feature maps in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0029] Figure 3 Input schematic diagrams of two serial eigenvalue data streams;
[0030] Figure 4 Schematic diagram of the two-stage compression process of eigenvalues in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0031] Figure 5 Schematic diagram of the calculation process of the eigenvalue data stream in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0032] Figure 6 Schematic diagram of the circuit structure of the weight preprocessing module in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0033] Figure 7 Schematic diagram of the circuit structure of the encoder in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0034] Figure 8 13TSRAM transistor-level schematic diagram in the convolutional neural network accelerator based on look-up table and in-memory computing for the embodiments of this application;
[0035] Figure 9Schematic diagram of the circuit structure of the first lookup table in the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application;
[0036] Figure 10 Schematic diagram of the circuit structure of the second lookup table in the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application;
[0037] Figure 11 Schematic diagram of the circuit structure of the accumulator in the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application;
[0038] Figure 12 Schematic diagram of the circuit structure of the decoder in the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application;
[0039] Figure 13 Schematic diagram of the partial sum storage structure in the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application;
[0040] Figure 14 PE working flow chart under bit eigenvalue interleaved input of the convolutional neural network accelerator based on lookup table and in-memory computing according to the embodiments of the present application. Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] Reference Figure 1, a convolutional neural network accelerator provided by an embodiment of the present application, includes an on-chip main memory and a plurality of arithmetic processing units. The arithmetic processing unit includes a data preprocessing unit, an in-memory computing array unit, a first adder, and an accumulator that are connected in sequence. In the calculation process of the convolutional neural network, the eigenvalue data stream used is serially input based on bit eigenvalues and the input eigenvalue sequence is compressed to obtain a compressed eigenvalue sequence. The weights used in the calculation process of the convolutional neural network are compressed to obtain compressed weights. The on-chip main memory is used to buffer the compressed eigenvalue sequence and the compressed weights and transmit them to the arithmetic processing unit; the data preprocessing unit is used to decompress the compressed eigenvalue sequence and the compressed weights, and classify the bit eigenvalues corresponding to the decompressed eigenvalue sequence into sparse-type bit eigenvalues and dense-type bit eigenvalues, and input the sparse-type bit eigenvalues and the dense-type bit eigenvalues into the in-memory computing array unit in an interleaved pattern of dense type - sparse type - dense type and perform convolution calculation with the decompressed weights. For the sparse-type bit eigenvalues, the in-memory computing array unit will perform zero-skipping calculation, and for the dense-type bit eigenvalues, the in-memory computing array unit will not perform zero-skipping calculation; the in-memory computing array unit includes a plurality of in-memory computing macro modules arranged in an array, and performs matrix multiplication operation in the form of a lookup table through the in-memory computing macro modules to obtain a first partial sum result. The first adder is used to add the first partial sum results output by the in-memory computing macro modules in each column of the in-memory computing array unit to obtain a second partial sum result. The accumulator is used to calculate the second partial sum result to obtain the multiplication-accumulation result corresponding to the eigenvalue data stream.
[0043] Specifically, the convolutional neural network accelerator proposed by the embodiment of the present application is constructed based on a digital in-memory computing structure (Digital-SRAM CIM) that uses digital signals to complete multiplication and addition calculations. As Figure 1As shown in the figure, the entire accelerator mainly includes one on-chip main memory (Global Buffer, GB) and eight processing elements (PEs). Among them, the GB is composed of six memory units GBcell with a size of 2KB. Therefore, the storage resource of the entire GB is 12KB, and its main function is to cache the feature values and weights input to the chip. The PE is the core computing unit on the chip. Each PE mainly includes two parts. The first part is a data preprocessing unit composed of a compression decoder, an activation reshaping module (ActReshape), and a sparse engine module (Sparse Engine), which is responsible for decompressing the feature values and weights and performing simple pre-operations. The second part is an in-memory computing array unit (CIM Array) composed of 3×3 in-memory computing macro modules (CIM Macro), which is responsible for matrix multiplication operations. The storage space of each CIM Macro is 1.5KB, and all the weights are stored in all CIM Macros. In addition, the second register in the accumulator, which is composed of a sparse output buffer (SpaOutBuf) and a dense output buffer (DenOutBuf), can support the zero skipping of feature values for unstructured pruning. Among them, the SpaOutBuf is composed of a first-in, first-out memory (FIFO) with a depth of 64. The DenOutBuf is composed of a dual-port SRAM (Two-Port SRAM) with a depth of 256 and a third adder.
[0044] In the embodiments of this application, the convolutional neural network accelerator proposed adopts a fixed weight to CIM structure. There are mainly two working states of the convolutional neural network accelerator. The first is weight loading. At this time, the GB is mainly used for buffering and storing the compressed weights. The compressed weights are loaded from off-chip into the chip and cached in the GB. Then, they are input from the GB into the CIM Macro of the PE. During this process, the compression decoder will decompress the compressed weights and restore the original zero values in the weights. The second working state is calculation. At this time, the GB is mainly used for buffering and storing the compressed feature values. After the weights are loaded into the CIM, the feature value data stream is sequentially input into the CIM Macro and completes the convolution calculation with the weights stored in the CIM Macro. The first summation result output from the CIM Macro is the result of the convolution calculation of the feature values and weights.
[0045] In a specific embodiment, the process of the serial input of the feature value data stream based on bit feature values is as follows:
[0046] The first eigenvalue and the second eigenvalue with a bit width of n bits in the eigenvalue data stream are respectively divided into a first bit eigenvalue sequence and a second bit eigenvalue sequence, and both the first bit eigenvalue sequence and the second bit eigenvalue sequence contain n bit eigenvalues;
[0047] The bit eigenvalues in the first bit eigenvalue sequence and the bit eigenvalues in the second bit eigenvalue sequence are arranged in the same bit order and alternately input from low to high in bit order to form an eigenvalue sequence.
[0048] Specifically, in this case, adjacent eigenvalues on the same feature map are similar. These similarities can be fully utilized to help optimize the performance of the circuit. Therefore, based on this characteristic of eigenvalues, the embodiments of this application design a special eigenvalue data stream. This data stream first splits the eigenvalue with a bit width of 8 bits into 8 bit feature maps, each bit feature map corresponding to a bit eigenvalue, and then serially inputs these bit eigenvalues into the CIM Macro in a specified order in units of bit eigenvalues. After completing the calculation of the bit eigenvalues in one eigenvalue, the calculation of the bit eigenvalues in the next eigenvalue is completed. This can ensure a simple calculation circuit while fully utilizing the similarity of adjacent eigenvalues and reducing the dynamic power consumption of CIM.
[0049] In the embodiments of this application, the bit widths of all eigenvalues and weights are 8 bits. The eigenvalue data stream proposed in the embodiments of this application needs to perform preprocessing on each eigenvalue outside the chip as Figure 2 shown. This preprocessing is to divide the eigenvalue with a bit width of 8 bits into 8 bit feature maps (Bit Feature Map). Each pixel point in the original feature map is 8-bit data. Each pixel point in the divided bit feature map is a 1-bit bit eigenvalue. Each bit feature map corresponds to the same bit of the pixel points in the original feature map. As Figure 2 in, BitFeatureMap7 represents the bit feature map composed of the highest bits of the eigenvalues in the original feature map. And BitFeatureMap6 represents the bit feature map composed of the seventh bit of the eigenvalues in the original feature map, and so on.
[0050] Figure 3 Two different data streams are given. Figure 3 (a) represents the traditional serially input eigenvalue data stream. This eigenvalue data stream is input into the CIM Macro for calculation in units of eigenvalues. If two 4-bit data "a" and "b" in the figure need to be transmitted, it is necessary to first transmit all of "a" through 4 cycles and then transmit "b". In the CIM Macro using this eigenvalue data stream, a multi-bit eigenvalue is calculated before the next eigenvalue is input. Figure 3(b) represents the eigenvalue data stream proposed in the embodiments of the present application. This eigenvalue data stream uses bit eigenvalues as units. As Figure 3 shown in (b), the order of data input is in accordance with the bit positions. After transmitting the least significant bit of "a", i.e., "a0", the least significant bit of "b", "b0", will start to be transmitted. After the least significant bits of "a" and "b" are both transmitted, the second least significant bits of "a" and "b" will be transmitted, and so on, alternately completing the transmission of bit eigenvalues at each position.
[0051] By adopting this eigenvalue data stream, the numerical similarity of adjacent eigenvalues in the same feature map can be utilized. In a picture, there is similarity between adjacent pixel points, but the bit values of pixel points are random. If bit-serial input is performed in units of eigenvalues, then the two consecutive bit values input to the CIM Macro are random and have no similarity, which may cause more flips in the circuit. However, if the feature map is divided in the way Figure 2 shown, the adjacent values of each bit feature map are likely to be the same. That is, there will likely be consecutive "0"s and consecutive "1"s on the bit feature map. As Figure 3 shown in (a), when bit-serial input is performed in units of eigenvalues and after transmitting the two data "a" and "b", the data is flipped a total of 4 times. However, if bit-serial input is performed in units of bit eigenvalues and after transmitting the two data "a" and "b", the data is only flipped a total of 2 times, as Figure 3 shown in (b). Therefore, adopting this serial input eigenvalue data stream based on the bit feature map can theoretically reduce the inversion probability of the calculation part circuit of the Macro.
[0052] In a specific embodiment, the compression method adopted for weights is binary Huffman coding; the compression method adopted for the eigenvalue sequence is two-level compression. The first-level compression method is binary Huffman coding to obtain the first-level compressed bit eigenvalues with the flag bits being 0 or 1; for the first-level compressed bit eigenvalues with the flag bit being 1 after binary Huffman coding, the second-level compression is performed. The second-level compression method is to directly divide according to different channel dimensions or first divide into blocks and then group the bit eigenvalues on the same channel of all blocks.
[0053] Specifically, to improve the efficiency of data transfer between the chip and the outside, the eigenvalues and weights are transmitted and stored in a compressed state before passing through the data preprocessing module of the PE. The compression method used for the weights is binary Huffman coding. The compression method used for the eigenvalues is a novel compression method proposed in the embodiments of this application, and this compression method is closely related to the eigenvalue data stream used by the convolutional neural network accelerator. Although there are differences between the compression methods of eigenvalues and weights, the underlying compression logic used by both is the same. Therefore, the on-chip decoding circuit, that is, the decompression module, can support the decompression of both eigenvalues and weights simultaneously. Since the eigenvalues and weights do not access the decompression module at the same time, each PE only needs one decompression module to complete the decoding of eigenvalues and weights. This greatly saves the hardware overhead of the decoding circuit part.
[0054] Reference Figure 4 , each cuboid of the BitFeatureMap represents a group of 64 bit eigenvalues, that is, 64 bit eigenvalues input to the Macro in each cycle. Generally speaking, in the BitFeatureMap corresponding to the high bits, a large number of zero values will appear. Therefore. The input eigenvalues will first undergo the first-level compression, and the compression scheme used is also binary Huffman coding. The specific method is as follows: If it is detected that a certain 64-bit bit eigenvalue of the input is all zeros, then this eigenvalue will be discarded, such as Figure 4 the eigenvalue represented by the gray cuboid in. At the same time, a corresponding flag bit (denoted as Flag) will be set to 0. On the contrary, if the 64-bit bit eigenvalue is non-zero, the value will be stored, and their corresponding flag bits will be set to 1.
[0055] Before performing the second-level compression, these 64-bit bit eigenvalues need to be re-divided into eigenvalue sequences with each group of 8 bits. Figure 4 Two division schemes are given in. Scheme one is to divide the 64-bit bit eigenvalue into 8 sub-eigenvalue sequences of 8 bits according to method 1. Since the 64-bit bit eigenvalue before division comes from 64 different channels, the division of scheme one is carried out in the channel dimension. The advantage of scheme one is that the division method is simple and direct and can be carried out directly after the first-level compression of the bit eigenvalues. Scheme two is the method represented by method 2. It divides the bit eigenvalues on the same channel of 8 consecutive bit eigenvalues into a group. The 8-bit sub-eigenvalue sequence obtained by dividing with scheme two represents 8 consecutive eigenvalues on the same channel. The 8-bit bit eigenvalues obtained after the division and recombination by scheme two have similarity. That is, the probability that the recombined 8-bit bit eigenvalue is all zeros will increase compared with method 1. But the price is that the division method is relatively complex, and the eigenvalues need to be converted once after decoding on the chip.
[0056] Embodiments of the present application demonstrate through experiments that for both VGG16 and AlexNet, the compression method with two - stage compression and the second - stage compression being Scheme 2 is the most efficient.
[0057] In a specific embodiment, the data pre - processing unit includes a decompression module, an eigenvalue block module, and a sparse engine module. The decompression module is used to decompress the compressed eigenvalue sequence and the compressed weights. The eigenvalue block module is used to block the decompressed eigenvalue sequence. The sparse engine module is used to classify the bit - eigenvalues corresponding to the decompressed eigenvalue sequence.
[0058] Specifically, the decompression module (Compression Decoder) is used to decompress the compressed data. Among them, the decompressed weight wei can be directly given to the CIM Macro, while the bit - eigenvalues act corresponding to the decompressed eigenvalue sequence still need to be processed subsequently.
[0059] The eigenvalue block module (Act Reshape) blocks the input decompressed eigenvalue sequence to facilitate the 3*3 convolution calculation with 9 CIM Macros. This does not contain innovative points.
[0060] The sparse engine module (Sparse Engine) classifies the bit - eigenvalues corresponding to the decompressed eigenvalue sequence into bit - eigenvalues of the sparse type (SPA) and bit - eigenvalues of the dense type (DEN). The function of the sparse engine module is to serve the partial - sum storage structure in the second register of the accumulator that supports eigenvalue skip - zero. The entire convolutional neural network accelerator will skip zero for the input of SPA - type bit - eigenvalues, while it will not skip zero for the input of DEN - type bit - eigenvalues.
[0061] In a specific embodiment, the in - memory computing macro module includes a weight pre - processing module, an encoder, a lookup - table array, and an adder tree. The weight pre - processing module is used to perform addition and subtraction operations on the decompressed weights to obtain weight operation results, and accumulate the decompressed weights to obtain weight accumulation results, and input the weight operation results and weight accumulation results into the lookup - table array and the decoder respectively. The lookup - table array includes several lookup - table units arranged in an array. The adder tree sums up all the outputs of the lookup - table units to obtain a first partial - sum result.
[0062] Specifically, the traditional all - digital in - memory computing circuit architecture is composed of a storage array, an adder tree, and an accumulator. Embodiments of the present application propose a new all - digital in - memory computing circuit architecture based on a lookup table, that is, the in - memory computing macro module, as Figure 5As shown. Different from the traditional architecture, the in-memory computing macro module proposed in the embodiments of the present application includes a weight preprocessing module (Weight Preprocess Module), an encoder (Encoder), a lookup table array (LUT Array), and an adder tree (Adder Tree). The in-memory computing array unit composed of the in-memory computing macro module is subsequently connected to a first adder, an accumulator (Accumulator), and a decoder (Decoder) in sequence. After the preprocessing module decompresses the weights and outputs the data wei, the wei is input to the weight preprocessing module of the in-memory computing macro module to perform addition and subtraction operations on the weights, and then writes the processed data into the storage unit in the lookup table unit. Therefore, the output line is connected to the write bit line (WBL) of the storage unit in the lookup table unit. The data preprocessing module decompresses the input feature values and outputs the bit feature values act corresponding to the decompressed feature value sequence. The bit feature values act corresponding to the decompressed feature value sequence are input to the encoder (Encoder) of the in-memory computing macro module for encoding. The encoded data will be connected to the RWL read word line in the memory of the lookup table unit to find the corresponding data for output, completing the lookup table calculation. The output data LO of the lookup table unit then passes through the adder tree (AdderTree) to complete the overall addition calculation to obtain the first partial addition result Psum0. Similarly, Psum1 and Psum2 in the in-memory computing array units of the other two columns are obtained. Psum0, Psum1, and Psum2 are the first partial sum results obtained by performing multiply-accumulate operations on the decompressed feature value sequence and the decompressed weights in the lookup table unit. Since the overall structure is a bit-serial structure, the first partial sum result needs to be accumulated through the accumulator (Accumulator) for multiple cycles to obtain the second addition result Acc_out. If the bit feature values corresponding to the decompressed feature value sequence input are encoded before participating in the calculation, the second addition result Acc_out needs to be decoded by the decoder (Decoder) to output the final calculation result CIM_OUT.
[0063] Reference Figure 6 , the weight preprocessing module consists of a FIFO with a depth of 2 and an arithmetic logic unit (ALU) that can perform addition and subtraction. The weight preprocessing module receives two sets of decompressed weights as input, performs addition and subtraction operations on the two sets of decompressed weights, and finally writes the obtained result into the lookup table array, that is, the WBL signal. In addition, the weight preprocessing module also accumulates the values of all weights and outputs the weight accumulation result SUM_W for the decoder to perform decoding.
[0064] In a specific embodiment, the in-memory computing macro module further includes an encoder, which is used to encode the decompressed eigenvalue sequence from binary data form into signed weight encoding form; the types of lookup table units are divided into a first lookup table and a second lookup table. The first lookup table is used to determine, according to the input control signal, an unsigned operation or an inversion operation in signed operation for the first eigenvalue and the second eigenvalue input to the lookup table array in binary data form; the second lookup table is used to perform a signed operation on the third eigenvalue and the fourth eigenvalue in the signed weight encoding form output by the encoder and complete the inversion operation based on the signed weight; the encoder includes a first register, a first inverter, and a first multiplexer. The decompressed eigenvalue sequence is input to the lookup table array bit by bit in 8 cycles. In the 9th cycle, the bit eigenvalue input in the 8th cycle stored in the first register is selected through the first multiplexer, and its inversion result is obtained through the first inverter and then output to the lookup table array with the second lookup table.
[0065] Specifically, if the lookup table unit in the lookup table array adopted in the embodiment of the present application is a second lookup table; then it must be encoded by the encoder. The data input to the second lookup table needs to be encoded by the encoding unit according to the encoding method shown in Table 1, so that the decompressed eigenvalue sequence is encoded from binary data form into signed weight encoding form. And the weight is still represented in two's complement form.
[0066] Table 1
[0067]
[0068] In the input decompressed eigenvalue sequence, each 0 of each bit is regarded as -1, that is, if a certain bit is 0, it means that the weight on this bit is negative, while 1 is the same as the original binary. It can be noted from the encoding method that the lowest bit data represents +1 or -1. Therefore, the original data for encoding needs to be odd. If this representation method is to be carried out, the original data needs to be multiplied by two and then added by one first, so that it can be ensured that it is odd for encoding.
[0069] The encoding method obtained through specific derivation is: assuming that a7a6a5a4a3a2a1a0 is an 8-bit number, if the above-mentioned representation method is to be adopted, a bit needs to be added at the highest bit, and the value of this bit is the inversion result of the original highest bit value.
[0070] In the first step, multiply by two and then add one to ensure that it is odd when encoding:
[0071] a7a6a5a4a3a2a1a0 => a7a6a5a4a3a2a1a01;
[0072] In the second step, perform the encoding formula corresponding to the input value in Table 2. This encoding formula discards the last digit and then adds the inverted value of the original most significant bit at the highest bit position.
[0073]
[0074] These two steps can be combined, that is, directly invert the number a7 at the original most significant bit and add it to the existing highest bit. The specific encoding formula can be obtained as follows:
[0075]
[0076] Based on the proposed encoding formula, the embodiments of this application design corresponding circuits. The specific circuit is as Figure 7 shown. The encoder includes three parts: a first register (reg), a first inverter, and a first multiplexer. Since an input bit-serial architecture is adopted, the bit eigenvalues act corresponding to the decompressed eigenvalue sequence are input into the in-memory computing macro module bit by bit in 8 cycles. Therefore, through a Neg signal, in the ninth cycle, the inverted result of the a7 data stored in the first register reg is selected and input into the lookup table array, thereby completing the encoding of the bit eigenvalue act corresponding to the decompressed eigenvalue sequence. The encoded data is transmitted to the IN of the lookup table unit.
[0077] In a specific embodiment, the first lookup table includes 2 input logic AND gates, 1 input logic XNOR gate, 1 signal input logic gate, three rows of storage units, 1 second inverter, 1 second multiplexer, and 9 output logic gates. The three rows of storage units store W0, W1, and W0 + W1 respectively. The second multiplexer selects to directly output the output of one of the storage units in the three rows of storage units or perform an inversion operation on the output of one of the storage units in the three rows of storage units through the second inverter according to the control signal input by the signal input logic gate and then output; the second lookup table includes 1 input logic XOR gate, 1 input logic XNOR gate, two rows of storage units, and 1 output logic XNOR gate. The two rows of storage units store W0 - W1 and W0 + W1 respectively. The input of the output logic XNOR gate is the third eigenvalue and the output of one of the storage units in the two rows of storage units, and it determines whether to perform an inversion operation on the output of one of the storage units in the two rows of storage units according to the third eigenvalue.
[0078] Specifically, the embodiments of this application propose two types of lookup table units. The bit-level unit (Bitcell) of the lookup table is as Figure 8As shown below. Briefly introduce the structure and working principle of this unit. This unit consists of 13 transistors, among which N1, N2, N3, N4, P1, and P2 are the structures of the 6T STAM unit. The AND gate circuit formed by P0 and N0 is to support row selection and column selection. CWL is the column selection signal, and WWL_B is the row selection signal. Only when CWL is 1 and WWL_B is 0 can the write operation be performed on this unit. Through such a design, the memory array composed of these cells can support the bit-interleaving structure without worrying about the semi-selection problem of data. P3 is a write assist transistor. When a write operation is performed on the unit, EN being 1 makes P3 conduct, and the sources of P1 and P2 are disconnected from VDD, losing the pull-up ability, thus avoiding the competition between pull-up and pull-down during data writing and improving the write ability under low voltage. P5, P6, N5, and N6 form a tri-state gate circuit for reading out data, and this readout circuit is the basis of the lookup table unit proposed in the embodiment of this application. Through this tri-state gate readout circuit, the RBLs of multiple rows of cells can be directly connected in the embodiment of this application, facilitating the design of the lookup table unit.
[0079] The function of the lookup table unit is to complete a0×W0 + a1×W1. Based on this, two types of lookup table units are proposed, the first lookup table and the second lookup table.
[0080] In order to support signed input, the embodiment of this application designs the truth table corresponding to the first lookup table, as shown in Table 2. This truth table has three inputs, namely Mode, IN1, and IN2, and the output is LO. Among them, Mode = 0 indicates unsigned operation. At this time, it can be seen that the truth table is the same as the truth table corresponding to the traditional lookup table, and the four outputs of LO are the four multiply-accumulate results corresponding to the two inputs IN0 and IN1 and the two weights W0 and W1. When Mode = 1, it indicates signed operation. At this time, the four outputs of LO are the inverted forms of the corresponding multiply-accumulate results. For example, when IN0 = 1 and IN1 = 0, the multiply-accumulate result IN0×W0 + IN1×W1 = W0, but since it is a signed operation at this time, the output is the inverted form of W0. It should be noted that the inversion operation is not equivalent to a complete signed operation. In multiplication, signed operations are usually implemented through two's complement operations, that is, first calculate the unsigned operation result, and then perform the "invert and add 1" operation. Therefore, the output of the first lookup table Only completes the "inversion" step, and the "add 1" operation is implemented through the design of the subsequent calculation flow, which will be introduced in detail in Section 3. Based on this, the first lookup table proposed in the embodiment of this application only needs to complete the inversion function.
[0081] Table 2 Truth table corresponding to the first lookup table under dual input
[0082]
[0083]
[0084] According to the truth table in Table 2, the circuit structure of the first lookup table designed in the embodiment of the present application is as Figure 9 shown. The lookup table unit consists of three rows of storage units, that is, only 3 data need to be stored to look up the 8 data shown in Table 2. Each row has 9 bits, indicating that the weight bit width BW at this time is 8 bits, and the three data W0, W1, and W0+W1 are stored respectively. The number of logic gates for controlling signal input on the left is 1 for each row, and the output of the logic gate is connected to the read word line signal (RWL) of all storage units in that row. When the output of the logic gate is 1, the result of the corresponding row is read. The number of logic gates for output on the right is 9, which are respectively connected to the read bit lines (RBL) of 9 columns of storage units for outputting the calculation result. When IN0 = 1 and IN1 = 0, the input logic AND gate of the first row outputs 1, thus opening the read port of the first row of storage units. If Mode = 0, the output logic selects the read result W0 as the calculation result, corresponding to the unsigned calculation result IN0×W0+IN1×W1 = W0; if Mode = 1, indicating that a signed calculation is required, the inverted form of W0 is selected as the calculation result. Similarly, when IN0 = 0 and IN1 = 1, the read port of the second row of storage units is opened, and its processing flow is the same as above; when IN0 = 1 and IN1 = 1, the third row of storage units is opened, and its processing flow also follows the same logic. When IN0 = 0 and IN1 = 0, regardless of whether it is a signed calculation or not, the result of IN0×W0+IN1×W1 is 0. Therefore, the output logic does not select the read result of the storage unit, but directly selects "0" as the output of this calculation. One design detail here is that the input logic gate of the third row is an exclusive-NOR gate, which means that in this case, although the read result of the third row of storage units is not selected, the read port is also opened. Such a design ensures that the read bit line RBL is not in a high-impedance state at this time. Because if RBL is in a high-impedance state, it will cause the input level of the subsequent first inverter and first multiplexer, which act as loads, to enter an intermediate level state, resulting in a sharp increase in leakage current and ultimately a significant increase in the leakage power consumption of the entire circuit.
[0085] Specifically, when designing the truth table of the second lookup table in the embodiments of the present application, a full-bit signed weight encoding is introduced, that is, the encoding method shown in Table 1. The weights are represented as positive and negative numbers in the form of normal two's complement. For the decompressed eigenvalue sequence of the input, the encoding method shown in Table 1 is adopted, where "0" represents "-1" and "1" still represents "1". After encoding, the second lookup table only needs to be designed based on the unified norm at the bit level calculation.
[0086] Based on this encoding method, the embodiments of the present application designed the truth table of the second lookup table, as shown in Table 3. When IN0 = 1 and IN1 = 1, the result of IN0×W0 + IN1×W1 is W0 + W1, and the output LO0 of the corresponding lookup table is W0 + W1; when IN0 = 0 and IN1 = 0, the result of IN0×W0 + IN1×W1 is -(W0 + W1), and the output S0 of the corresponding lookup table is the inverted form of W0 + W1. The cases of IN0 = 1 and IN1 = 0, IN0 = 0 and IN1 = 1 are the same. It can be seen that due to the full-bit signed weight encoding of the decompressed eigenvalue sequence of the input, the size of the truth table corresponding to the second lookup table under dual inputs is half of that of the first lookup table, and the data to be stored is greatly reduced.
[0087] Table 3 Truth table corresponding to the second lookup table under dual inputs
[0088]
[0089] Based on this truth table, the circuit structure of the second lookup table proposed by the embodiments of the present application is as Figure 10 shown. The second lookup table consists of two rows of storage units, that is, only 2 data need to be stored to look up the 4 data shown in Table 3. Compared with the first lookup table, the area of the storage unit is reduced by half. Each row is still 9 bits, indicating that the weight accuracy is 8 bits at this time, and the two data W0 - W1 and W0 + W1 are stored respectively. The number of logic gates on the left and right sides is the same as that of the first lookup table. The number of logic gates for controlling signal input is 1, and the number of logic gates for output is 9. It is worth mentioning that the output logic is simplified from the structure of the first inverter plus the first multiplexer to an exclusive NOR gate, which further reduces the area of the overall structure.
[0090] The calculation principle of the second lookup table is as follows: when IN0 = 1 and IN1 = 0, the output of the input logic XOR gate in the first row is set to 1, and the read port of the storage unit is opened. Since IN0 = 1, the output logic XNOR gate at this time is equivalent to a buffer, and the stored data W0 - W1 is directly read as the calculation result, which conforms to the calculation formula IN0×W0 + IN1×W1 = 1×W0 - 1×W1 = W0 - W1; when IN0 = 0 and IN1 = 1, the read port of the storage unit in the first row is still opened, but because IN0 = 0, the output logic XNOR gate at this time is equivalent to an inverter, and the inverted form of the stored data W0 - W1 is read as the calculation result. The cases of IN0 = 1 and IN1 = 1, IN0 = 0 and IN1 = 0 are the same. In addition, consistent with the first lookup table, for signed operations, the second lookup table only needs to complete the "inversion" operation, and the "add 1" operation is still implemented through the design of the subsequent calculation flow.
[0091] For each in-memory computing macro module, all lookup table units in its lookup table array adopt the same type of lookup table unit, that is, the first lookup table or the second lookup table; for different in-memory computing macro modules, the above two types of lookup table units can be arbitrarily selected for calculation according to requirements. After obtaining the outputs of all lookup table units in the lookup table array of each in-memory computing macro module, the adder tree is a conventional adder tree, and this adder tree is responsible for adding up all the outputs LO of the lookup table units to obtain the first partial sum result.
[0092] This in-memory computing macro module preprocesses the weights and stores a part of the sum result of the adder in the storage unit in advance. Based on this, this in-memory computing macro module can use the read operation of SRAM in the form of a lookup table to replace a part of the multiply-accumulate operations. This in-memory computing macro module can reduce nearly half of the adders in the in-memory computing macro module compared with the traditional structure, thus greatly reducing the area and power consumption of the in-memory computing macro module.
[0093] The embodiment of this application consists of 9 in-memory computing macro modules to form an in-memory computing array unit. For sparse type bit feature values, the in-memory computing array unit will perform zero-skipping calculation. For dense type bit feature values, the in-memory computing array unit will not perform zero-skipping calculation. Non-zero-skipping calculation means that for the input bit feature values, regardless of whether there is a situation where an entire column of bit feature values is 0, the in-memory computing array unit adopts a sequential calculation strategy, so as to output the second partial sum result with continuous addresses; zero-skipping calculation means that for the decompressed eigenvalue sequence input by the convolutional neural network accelerator, if there is a situation where an entire column of bit feature values is 0, it will adopt a strategy of directly skipping the calculation, so as to improve the calculation speed, but the output will be the second partial sum result with discontinuous addresses.
[0094] In a specific embodiment, the accumulator includes a second adder, a sign adder, a second shifter, and a second register. The sign adder is configured to accumulate the number of bit eigenvalues representing negative numbers in the decompressed eigenvalue sequence during signed operations to obtain an accumulation result, and input the accumulation result into the second adder to add it to all partial sum results corresponding to the decompressed eigenvalue sequence to obtain an addition result. The second adder, the second shifter, and the second register form an accumulation structure, and the accumulation structure is configured to perform shift accumulation on all addition results corresponding to the eigenvalue data stream.
[0095] Specifically, referring to Figure 11 , the accumulator is composed of four parts, namely, the second adder (Adder), the sign adder (Sign Adder), the shifter, and the second register (Reg). Among them, the second adder (Adder), the shifter, and the second register (Reg) form a traditional accumulation structure, which is used to perform shift accumulation on the first partial sum results obtained by adding up the previous adder tree. The first partial sum results are the multiply-accumulation partial sum results of 1-bit data input in parallel and 8-bit weights in the storage array. However, the bit eigenvalue act corresponding to the input decompressed eigenvalue sequence is 8 bits, and the input data after encoding is 9 bits. Therefore, 9 first partial sum results will be generated in 9 cycles. The 9 first partial sum results are successively given to the accumulator for shift accumulation operations, and finally, the multiply-accumulation result Acc_out of the 8-bit data input in parallel and the 8-bit weights in the array can be obtained. Different from the traditional accumulator, this accumulator has an additional sign adder. As mentioned above, the lookup table unit in the embodiment of the present application only completes the "inversion operation" in the two's complement operation and does not complete the "add 1" operation. Therefore, the sign adder is used to accumulate the number of bit eigenvalues representing negative numbers in the input decompressed eigenvalue sequence during signed operations, and input the accumulation result into the second adder (Adder) to add it to the first partial sum results, so as to complete the "add 1" operation. In this way, the overall architecture can support signed input operations.
[0096] In a specific embodiment, a decoder is further included. The decoder is configured to decode a decoding result based on the multiplication-accumulation result and the weight accumulation result corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator. The decoder includes a subtractor and a first shifter connected in sequence. The multiplication-accumulation result and the weight accumulation result corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator are input into the subtractor to perform a subtraction operation, obtaining a subtraction result. The first shifter completes the operation of dividing the subtraction result by two to obtain the decoding result. The decoding result or the multiplication-accumulation result corresponding to the eigenvalue data stream obtained by using the first lookup table and passing through the adder tree and the accumulator is used as the convolution calculation result.
[0097] Specifically, the function of the decoder is to correctly decode the result. In the case of using the second lookup table, the data encoded by the encoder needs to multiply the original data by two and add one first. Therefore, the multiplication-accumulation result of the bit eigenvalue corresponding to the decompressed eigenvalue sequence and the weight is not the final result, and it needs to be restored to the result corresponding to the original data through the decoder.
[0098] The formula corresponding to the circuit of the decoder is as follows:
[0099]
[0100] Among them, Acc_out represents the output of the accumulator, CIM_OUT represents the correct calculation result output by the final in-memory computing macro module, and SUM_W represents the direct addition result of all weight data of the convolution kernel.
[0101] The finally obtained circuit of the decoder is as Figure 12 shown, which is composed of a subtractor and a first shifter.
[0102] In a specific embodiment, the second register adopts a partial sum storage structure that supports eigenvalue skip-zero calculation of the structured pruning compression model. The partial sum storage structure includes a sparse output cache block and a dense output cache block. The second partial sum result output after the first dense-type bit eigenvalue passes through the in-memory computing array unit is written into the dense output cache block, and the position information of the second partial sum result output after the first sparse-type bit eigenvalue passes through the in-memory computing array unit is written into the sparse output cache block. Before the second partial sum result output after the next dense-type bit eigenvalue passes through the in-memory computing array unit is written into the dense output cache block, it will be added to the second partial sum result with the same position information in the sparse output cache block and the dense output cache block, and then written into the dense output cache block. The next sparse-type bit eigenvalue continues to be written into the sparse output cache block, and the above steps are repeated until the second partial sum results corresponding to all sparse-type and dense-type bit eigenvalues with the same position information have completed the shift addition.
[0103] Specifically, the accumulation process in the in-memory computing macro module is implemented by using the partial sum storage structure as shown in Figure 13 . Among them, the Sparse Engine module classifies the bit eigenvalues corresponding to the decompressed eigenvalue sequence of the input into bit eigenvalues of the sparse type (SPA) and bit eigenvalues of the dense type (DEN). (When the sparsity exceeds 25%, it is DEN, and when it is less than 25%, it is SPA). The in-memory computing array unit (CIM Array) will skip zeros for the input of bit eigenvalues of the SPA type. Therefore, the corresponding addresses of the output partial sum data Spa_Psum are discontinuous. The in-memory computing array unit (CIM Array) will not skip zeros for the input of bit eigenvalues of the DEN type. Therefore, the corresponding addresses of the output partial sum data Den_Psum are continuous. However, Spa_Psum and Den_Psum with the same address need to be added. Therefore, in order to successfully sum Spa_Psum with discontinuous addresses and Den_Psum with corresponding addresses, the embodiment of the present application designs a partial sum storage structure composed of a sparse output buffer block (SpaOutBuf) and a dense output buffer block (DenOutBuf). The second register in the accumulator uses this partial sum storage structure, supports eigenvalue skip-zero calculation by using the sparse output buffer block and the dense output buffer block, and adds the second partial sum result output after the bit eigenvalues of the sparse type with the same position information pass through the in-memory computing array unit and the second partial sum result output after the bit eigenvalues of the dense type pass through the in-memory computing array unit to complete the multiply-accumulate operation.
[0104] Among them, SpaOutBuf is a single-port FIFO, and DenOutBuf is a dual-port SRAM. SpaOutBuf is used to store the Spa_Psum data with discontinuous addresses, and DenOutBuf is used to store the Den_Psum with continuous addresses. The workflow will be divided into three steps.
[0105] The first step: The in-memory computing array unit (CIM Array) will calculate a bit eigenvalue of the DEN type, and the generated Den_Psum with continuous addresses will be sequentially stored in DenOutBuf.
[0106] The second step: The in-memory computing array unit (CIM Array) will calculate a bit eigenvalue of the SPA type, and the generated Spa_Psum with discontinuous addresses will be temporarily stored in SpaOutBuf.
[0107] Step 3: The in-memory computing array unit (CIM Array) will calculate another bit eigenvalue of the DEN type. However, during the process of storing Den_Psum into DenOutBuf, it will be judged whether the address of the data to be read from the SpaOutBuf FIFO at this time is the same as the current Den_Psum. If they are the same, the data and the data of the current Den_Psum will be added and then written into DenOutBuf. If they are not the same, it will wait for the next judgment.
[0108] In addition, DenOutBuf is a dual-port SRAM. Therefore, during the process of writing Den_Psum to the corresponding address of DenOutBuf, it will also be accumulated with the data at the current address of DenOutBuf, thus completing the complete multiply-accumulate operation.
[0109] Spa_wdy is the write enable signal of SpaOutBuf, Spa_rdy is the read enable signal of SpaOutBuf, and Den_wdy is the write enable signal of DenOutBuf. These signals are used to control the reading and writing of the two buffers, so as to successfully complete the entire workflow.
[0110] In the convolutional neural network accelerator proposed in the embodiment of the present application, the sparse engine (Sprase Engine) module in the data preprocessing unit classifies the bit eigenvalues corresponding to the decompressed eigenvalue sequence. The bit eigenvalues with a sparsity less than 25% are sparse bit eigenvalues (denoted as SPA), and the bit eigenvalues with a sparsity greater than 25% are dense bit eigenvalues (denoted as DEN). The bit eigenvalues are not input into the in-memory computing array unit (CIMArray) for convolutional calculation in the order from low to high (or from high to low) of the bit positions, but are interleaved according to the difference in their sparsities. The bit feature map needs to be interleaved and input into the in-memory computing array unit (CIM Array) for convolutional calculation in the pattern of "DEN-SPA-DEN". The type of the first bit eigenvalue processed by the in-memory computing array unit (CIM Array) must be "DEN".
[0111] Figure 14 Illustrated the workflow of the "DEN-SPA-DEN" skip-zero strategy.
[0112] Step1: At the beginning of the convolution calculation, the first bit eigenvalue of the "DEN" type is input into the in-memory computing array unit (CIM Array) to complete the calculation, and the values in the DenOutBuf are sequentially filled by the second part and the result calculated by the in-memory computing array unit (CIM Array). Since the in-memory computing array unit (CIM Array) in the DENSE mode does not perform any zero-skipping operations, the position information of the second part and the result filled into the DenOutBuf is continuous. The data in the brackets in the figure represents the position information of the partial sum (the position information marked in the DenOutBuf is only for explaining the working principle, and only the values of the second part and the result are actually stored in the DenOutBuf).
[0113] Step2: After processing the bit eigenvalue of the "DEN" type, the in-memory computing array unit (CIM Array) will start to process a bit eigenvalue of the "SPA" type. At this time, the second part and the result output by the in-memory computing array unit (CIM Array) and their corresponding position information will be written into the SpaOutBuf. Because the position information of the second part and the result output by the in-memory computing array unit (CIM Array) is not continuous at this time, the position information of the second part and the result output is determined by the current non-zero eigenvalue data and cannot be obtained in advance.
[0114] Step3: The in-memory computing array unit (CIM Array) will process a bit eigenvalue of the "DEN" type again. When the second part and the result output by the in-memory computing array unit (CIM Array) are written into the DenOutBuf, it will check whether its position information matches the current output of the SpaOutBuf. If they are the same, before the second part and the result are written into the DenOutBuf, they need to be added to the value popped from the SpaOutBuf and then written into the DenOutBuf. For example, Figure 13At time T0 when processing the bit feature value of the second "DEN" type. At this time, the position information of the second part and the result output by the in-memory computing array unit (CIM Array) is (0, 0), but the position information of the second part and the result popped out by SpaOutBuf at this time is (0, 1). Since the value in SpaOutBuf is not fetched, no new number will be read out from SpaOutBuf. The second part and the result with the position information of (0, 1) are blocked in SpaOutBuf. At this time, the second part and the result calculated by the in-memory computing array unit (CIM Array) will be directly input into DenOutBuf. The position information of the second part and the result written into DenOutBuf at time T1 is (0, 1), and the position information of the second part and the result popped out by SpaOutBuf at this time is also (0, 1). At this time, before the second part and the result are input into DenOutBuf, they need to be added to the value popped out by SpaOutBuf. At this time, since the value popped out by SpaOutBuf is fetched, the new value, that is, the second part and the result with the position information of (0, 2), will be popped out from SpaOutBuf. Wait to be added to the second part and the result output by the in-memory computing array unit (CIM Array).
[0115] It should be noted that the first bit feature value processed by the in-memory computing array unit (CIM Array) must be a bit feature map of the "DEN" type. Because it is necessary to fill DenOutBuf completely before calculating the bit feature value of the "SPA" type to ensure that for each partial sum popped out from "SPA" during the calculation of the bit feature value of the "SPA" type, a corresponding position can be found in DenOutBuf, thus avoiding write errors.
[0116] The above partial sum storage structure is not only simple but also very effective in processing sparse features. In the above partial sum storage structure, SpaOutBuf is a single-port FIFO, and DenOutBuf is only a dual-port SRAM. This well controls the hardware complexity of the memory used to store the second part and the result. Because in general sparse accelerators, the memory used to store the second part and the result often needs to be designed into a very complex structure, while the above solution solves the storage problem of sparse data with only two simple cache blocks.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A convolutional neural network accelerator based on lookup table and in-memory computing, characterized in that: The invention comprises an on-chip main memory and a plurality of operation processing units, wherein the operation processing unit comprises a data preprocessing unit, an in-memory calculation array unit, a first adder and an accumulator connected in sequence, the eigenvalue data stream used in the calculation process of the convolutional neural network adopts a serial input based on bit eigenvalues and compresses the input eigenvalue sequence to obtain a compressed eigenvalue sequence, the weights used in the calculation process of the convolutional neural network are compressed to obtain compressed weights, the on-chip main memory is used to buffer the compressed eigenvalue sequence and the compressed weights and transmit them to the operation processing unit; the data preprocessing unit is used to decompress the compressed eigenvalue sequence and the compressed weights, and classify the bit eigenvalues corresponding to the decompressed eigenvalue sequence into sparse bit eigenvalues and dense bit eigenvalues, and classify the sparse bit eigenvalues into dense bit eigenvalues. The eigenvalues and the bit eigenvalues of the dense type are interleavedly input into the in-memory calculation array unit in a dense type-sparse type-dense type mode and convolution calculation is performed with the decompressed weights. For the bit eigenvalues of the sparse type, the in-memory calculation array unit will perform zero-jumping calculation, and for the bit eigenvalues of the dense type, the in-memory calculation array unit will not perform zero-jumping calculation; the in-memory calculation array unit includes a plurality of in-memory calculation macro modules arranged in an array, and a matrix multiplication operation based on a lookup table is performed by the in-memory calculation macro modules to obtain a first partial sum result, the first adder is used to add the first partial sum result output by the in-memory calculation macro module of each column in the in-memory calculation array unit to obtain a second partial sum result, and the accumulator is used to calculate the second partial sum result to obtain the multiplication and accumulation result corresponding to the eigenvalue data stream.
2. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 1, characterized in that: The process of the eigenvalue data stream using serial input based on bit eigenvalues is as follows: Splitting the first eigenvalue and the second eigenvalue with a bit width of N bits in the eigenvalue data stream into a first eigenvalue sequence and a second eigenvalue sequence, respectively, wherein the first eigenvalue sequence and the second eigenvalue sequence both contain N bit eigenvalues; The bit feature values in the first bit feature value sequence and the bit feature values in the second bit feature value sequence are arranged in the same order and are input alternately from low to high according to the order to form a feature value sequence.
3. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 1, characterized in that: The compression method used for the weights is binary Huffman coding; the compression method used for the eigenvalue sequence is two-level compression, the first-level compression method is binary Huffman coding, and the bit eigenvalue after the first-level compression with the flag bit being 0 or 1 is obtained; the second-level compression is performed on the bit eigenvalue after the first-level compression with the flag bit being 1 after binary Huffman coding, and the second-level compression method is to directly divide according to different channel dimensions or to first divide into blocks and then group the bit eigenvalues on the same channel on all blocks.
4. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 1, characterized in that: The data preprocessing unit includes a decompression module, an eigenvalue blocking module and a sparse engine module. The decompression module is used to decompress the compressed eigenvalue sequence and the compressed weight, the eigenvalue blocking module is used to block the decompressed eigenvalue sequence, and the sparse engine module is used to classify the bit eigenvalues corresponding to the decompressed eigenvalue sequence.
5. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 1, characterized in that: The in-memory calculation macromodule includes a weight preprocessing module, an encoder, a lookup table array and an adder tree; the weight preprocessing module is used to perform addition and subtraction operations on the decompressed weights to obtain weight operation results, and accumulate the decompressed weights to obtain weight accumulation results, and input the weight operation results and weight accumulation results into the lookup table array and the decoder respectively; the lookup table array includes a plurality of lookup table units arranged in an array; All outputs of the lookup table unit are added through the adder tree to obtain a first partial sum result.
6. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 5, characterized in that: The in-memory calculation macromodule also includes an encoder, which is used to encode the decompressed eigenvalue sequence from binary data form into a signed weight coding form; the types of the lookup table unit are divided into a first lookup table and a second lookup table, the first lookup table is used to determine the inversion operation of the first eigenvalue and the second eigenvalue input into the lookup table array in binary data form according to the input control signal; the second lookup table is used to perform a signed operation on the third eigenvalue and the fourth eigenvalue in the signed weight coding form output by the encoder and complete the inversion operation based on the sign weight; the encoder includes a first register, a first inverter and a first multiplexer, the decompressed eigenvalue sequence is input into the lookup table array bit by bit in 8 cycles, and in the 9th cycle, the bit eigenvalue input in the 8th cycle stored in the first register is selected by the first multiplexer, so that it passes through the first inverter to obtain the corresponding inversion result, and then output to the lookup table array with a second lookup table.
7. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 6, characterized in that: The first lookup table includes 2 input logic AND gates, 1 input logic XNOR gate, 1 signal input logic gate, three rows of storage cells, 1 second inverter, 1 second multiplexer and 9 output logic gates. The three rows of storage cells store W0, W1 and W0+W1 respectively. The second multiplexer selects, according to the control signal input by the signal input logic gate, to directly output the output of one of the three rows of storage cells or to invert the output of one of the three rows of storage cells through the second inverter and then output it; the second lookup table includes 1 input logic XOR gate, 1 input logic XNOR gate, two rows of storage cells and 1 output logic XNOR gate. The two rows of storage cells store W0-W1 and W0+W1 respectively. The input of the output logic XNOR gate is the third bit characteristic value and the output of one of the two rows of storage cells. It is determined whether to invert the output of one of the two rows of storage cells according to the third bit characteristic value.
8. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 6, characterized in that: It also includes a decoder, which is used to decode and obtain a decoding result according to the multiplication and accumulation results corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator and the weight accumulation result; the decoder includes a subtractor and a first shifter connected in sequence, and the multiplication and accumulation results corresponding to the eigenvalue data stream obtained by using the second lookup table and passing through the adder tree and the accumulator and the weight accumulation result are input into the subtractor to perform a subtraction operation to obtain a subtraction result, and the subtraction result is divided by two through the first shifter to obtain the decoding result, and the decoding result or the multiplication and accumulation result corresponding to the eigenvalue data stream obtained by using the first lookup table and passing through the adder tree and the accumulator is used as the convolution calculation result.
9. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 6, characterized in that: The accumulator includes a second adder, a sign adder, a second shifter and a second register, wherein the sign adder is used to accumulate the number of bit eigenvalues representing negative numbers in the signed operation in the decompressed eigenvalue sequence to obtain an accumulated result, and the accumulated result is input into the second adder to add all parts and results corresponding to the decompressed eigenvalue sequence to obtain an added result; the second adder, the second shifter and the second register constitute an accumulation structure, and the accumulation structure is used to perform shift accumulation on all addition results corresponding to the eigenvalue data stream.
10. The convolutional neural network accelerator based on lookup table and in-memory computing according to claim 9, characterized in that: The second register adopts a partial sum storage structure that supports the zero-jump calculation of the eigenvalue of the structured pruning compression model; the partial sum storage structure includes a sparse output cache block and a dense output cache block, the second partial sum result output after the first dense type bit eigenvalue passes through the in-memory calculation array unit is written into the dense output cache block, and the position information of the second partial sum result output after the first sparse type bit eigenvalue passes through the in-memory calculation array unit is written into the sparse output cache block; The second part and result of the bit feature value of the next dense type output after passing through the in-memory calculation array unit will be added with the second part and result of the same position information in the sparse output cache block and the dense output cache block before being written into the dense output cache block; The next sparse type bit feature value continues to be written into the sparse output buffer block, and the above steps are repeated until the second parts and results corresponding to all sparse type bit feature values and dense type bit feature values with the same position information are shifted and added.