A mixed-signal binary CNN processor
By designing a mixed signal binary CNN processor, using binary temperature decoding units and SC neurons, and combining filter memory alternating between near-memory computing and ping-pong methods, the problem of high energy in existing CNN chips in high-complexity tasks is solved, and efficient image classification and edge deployment are achieved.
Patent Information
- Application Number
- CN201810321430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2038-04-11
AI Technical Summary
In high complexity tasks, existing CNN chips have difficulty in edge deployment due to the high off-chip DRAM access energy, and under the demand of low-energy deep convolutional neural networks, it is difficult for the prior art to achieve efficient image classification.
A mixed signal binary CNN processor is designed, using binary temperature decoding units and energy-efficient switching capacitor (SC) neurons to realize image classification through near memory calculations, and optimize the calculation through ping-pong filter memory alternating input and output.
The image classification accuracy of 86% in the moderate complexity task was achieved and the classification energy was reduced to 3.8μJ, a 40-fold increase over TrueNorth, significantly improving the feasibility of edge deployment.
Smart Images

Figure CN110363292B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a CNN processor, in particular to a mixed signal binary CNN processor. Background Art
[0002] Convolutional Neural Network (CNN) is a feedforward neural network whose artificial neurons can respond to surrounding units within a certain coverage area and performs well in large image processing. It includes convolutional layers and pooling layers.
[0003] The trend of pushing cloud deep learning to the edge has created a demand for low-energy deep convolutional neural networks (CNNs) due to concerns about latency, bandwidth, and privacy.
[0004] Existing single-layer classifiers achieve sub-nJ operations but are limited to moderate accuracy on low-complexity tasks (90% on MNIST). Larger CNN chips offer dataflow computation for high-complexity tasks (AlexNet) with mJ energy, but edge deployment remains a challenge due to off-chip DRAM access energy.
[0005] Therefore, there is a particular need for a mixed-signal binary CNN processor to solve the above-mentioned existing problems. Summary of the invention
[0006] The purpose of the present invention is to provide a mixed-signal binary CNN processor that addresses the deficiencies of the prior art, performs moderately complex image classification (86% in CIFAR-10), and uses near-memory computing to achieve a classification energy of 3.8 μJ, a 40-fold improvement over TrueNorth.
[0007] The technical problem solved by the present invention can be achieved by adopting the following technical solutions:
[0008] A mixed signal binary CNN processor is characterized in that it includes a neuron array unit, a binary temperature decoding unit, a control unit, an input image unit, an output image unit and a storage unit, wherein an RGB image is input through the input end of the binary temperature decoding unit, the output end of the binary temperature decoding unit is connected to the input end of the neuron array unit through the input image unit, the output end of the neuron array unit is connected to the output image unit, the control unit is connected to the neuron array unit, a control instruction is input through the input end of the control unit, and the storage unit is connected to the neuron array unit.
[0009] In one embodiment of the present invention, the storage unit comprises a local memory, a first filter memory and a second filter memory, and the local memory, the first filter memory and the second filter memory are connected to the neuron array unit respectively.
[0010] Further, the first filter memory and the second filter memory alternately input and output in a ping-pong manner.
[0011] In one embodiment of the present invention, the input image unit includes an input image memory and an input demultiplexer, and the output end of the binary temperature decoding unit is connected to the input end of the neuron array unit through the input image memory and the input demultiplexer in sequence.
[0012] In one embodiment of the present invention, the output image unit comprises an output image memory and an output demultiplexer, and the output end of the neuron array unit is outputted in sequence through the output demultiplexer and the output image memory.
[0013] Compared with the prior art, the mixed-signal binary CNN processor of the present invention completes the work through the Binary Net algorithm of the binary temperature decoding unit, whose weight and activation constraints are +1 / -1, which greatly simplifies the multiplication operation (XNOR) and allows the integration of all on-chip storage units; the neuron array unit is a high-efficiency switched capacitor (SC) neuron to solve the challenge of Binary Net wide vector summation; it performs image classification of medium complexity (86% in CIFAR-10) and adopts near-memory computing to achieve 3.8μJ of classification energy, which is 40 times higher than TrueNorth, thereby achieving the purpose of the present invention.
[0014] The features of the present invention can be clearly understood by referring to the drawings and the following detailed description of the preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of the network topology of the mixed-signal binary CNN processor of the present invention;
[0016] Figure 2 It is a structural schematic diagram of the mixed signal binary CNN processor of the present invention;
[0017] Figure 3 A schematic diagram of how the mixed-signal binary CNN processor locality of the present invention is converted into reduced load;
[0018] Figure 4 A schematic diagram of the neuron principle of the mixed signal binary CNN processor of the present invention;
[0019] Figure 5It is a schematic diagram of the measurement results of the mixed signal binary CNN processor of the present invention at room temperature. DETAILED DESCRIPTION
[0020] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below with reference to specific diagrams. Example
[0021] like Figures 1 to 5 As shown, the mixed signal binary CNN processor of the present invention includes a neuron array unit 10, a binary temperature decoding unit 20, a control unit 30, an input image unit 40, an output image unit 50 and a storage unit 60. The RGB image is input through the input end of the binary temperature decoding unit 20, the output end of the binary temperature decoding unit 20 is connected to the input end of the neuron array unit 10 through the input image unit 40, the output end of the neuron array unit 10 is connected to the output image unit 50, the control unit 30 is connected to the neuron array unit 10, the control instruction is input through the input end of the control unit 30, and the storage unit 60 is connected to the neuron array unit 10.
[0022] In this embodiment, the specific circuit diagram of the neuron array unit 10, the binary temperature decoding unit 20, the control unit 30, the input image unit 40 and the output image unit 50 is shown in the attached drawings, which will not be described in detail here; the storage unit 60 adopts Hynix's HY62LF16806B.
[0023] In this embodiment, the storage unit 60 includes a local memory 61, a first filter memory 62 and a second filter memory 63, and the local memory 61, the first filter memory 62 and the second filter memory 63 are respectively connected to the neuron array unit 10. The first filter memory 62 and the second filter memory 63 are alternately input and output in a ping-pong manner.
[0024] In this embodiment, the input image unit 40 includes an input image memory 41 and an input multiplexer 42 , and the output end of the binary temperature decoding unit 20 is connected to the input end of the neuron array unit 10 through the input image memory 41 and the input multiplexer 42 in sequence.
[0025] In this embodiment, the output image unit 50 includes an output image memory 51 and an output demultiplexer 52 , and the output end of the neuron array unit 10 is outputted through the output demultiplexer 52 and the output image memory 51 in sequence.
[0026] like Figure 1As shown, the function and network topology of the mixed-signal binary CNN processor of the present invention are explained. By enforcing structural regularization, the physical architecture is allowed to maximize the locality of the CNN algorithm. Each CNN layer performs multi-channel multi-filter convolution. The number of filters in each convolution layer is limited to 256, the filter size is 2×2, and the number of channels is 256; the circuit advantages brought by this regularity are short lines and array-type low fan-out splitters, which minimize the path load between memory and logic.
[0027] like Figure 2 As shown, the mixed signal binary CNN processor of the present invention can support up to 9 layers and has customized instruction sets for input and output operations, CNN and fully connected (FC) layers. The mixed signal binary CNN processor of the present invention reads the RGB image, converts the channels into 85-level thermometer codes through the binary temperature decoding unit 20, and superimposes them into a 256-channel image as the input of the mixed signal binary CNN processor of the present invention. At the output, the 4-bit class label is digitally calculated through the local memory 61. For the CNN layer, the first filter memory 62 and the second filter memory 63 alternate input and output roles in a ping-pong manner. These storage units 60 are 256 bits wide, and each word represents a 256-channel pixel. The calculation of the mixed signal binary CNN processor of the present invention is completed in the neuron array unit 10, eliminating partial sums. The weights are transferred from the SRAM to the local memory 61 (latch) and reused, while the first filter memory 62 and the second filter memory 63 traverse the image. 64 neurons in the form of data parallel arrays process a fragment of the input image, amortizing the read energy corresponding to each filter at the input image memory 41 by 64 times. The input demultiplexer 42 interacts between the input image memory 41 (loading pixels) and the neuron array unit 10 (receiving fragments). For the FC layers, the weights are loaded from a separate SRAM bank 64 channels at a time, and the multiply-accumulate operations are performed sequentially in the digital domain.
[0028] like Figure 3, showing how locality translates into reduced load. The input demultiplexer 42 is a set of 1 to 4 demultiplexers with an output register. Each pixel of the input image can be reused to process two overlapping fragments, amortizing the input image memory 41 read energy for each filter calculation by a factor of 2. A 2 by 2 crossbar swaps pixel pairs at the input of the neuron array unit 10. The filter weights are transferred via a 4-bit bus per neuron, split into a north and south half to reduce the load of weight transfer by a factor of 2. To minimize the wiring of the neuron array to the memory, each neuron writes to the same 4 output channels in each CNN layer (1 per filter bank), allowing the output demultiplexer 52 to be implemented as a 1 to 4 demultiplexer array. Max pooling occurs stepwise during the convolution process by first reading a bit in the output image memory 51 and then writing back its logical OR with the neuron output.
[0029] like Figure 4Figure 1 shows a schematic diagram of a neuron, each of which computes a weighted sum of filters over a fragment of an input image. As storage energy is reduced through parallel distribution and reuse, multiplication is reduced to XNOR, and high fan-in addition becomes the main bottleneck. However, in the adopted SC neuron, the energy cost of addition is reduced by small voltage swings at the charge preservation nodes. In contrast, a digital adder tree would involve rail-to-rail voltage swings and a larger amount of switching capacitance in its different stages. The main noise source of the neuron is the comparator, but its energy cost is amortized by the 1024 weights, and the CNN can tolerate some noise. Therefore, the SC neuron is suitable for low-voltage operation and uses a 0.6V digital supply / analog reference voltage and a 0.8V comparator supply voltage. Since the SC neuron performs a weighted sum of data-dependent switching (in addition to the comparator), its energy varies with activity, like static CMOS. The SC neuron uses a capacitive DAC (CDAC) divided into four sections: a 1024-bit thermometer section to implement the filter, a binary weighting section for neuron bias, a threshold section (comparator), and a common-mode (CM) setting section to compensate for parasitic effects at the charge storage nodes. The comparator offset is digitized using calibration at startup, stored in a local register, and subtracted from the offset loaded from the SRAM during weight transfer. In environments where large temperature changes may cause significant offset drift, calibration can be performed periodically (e.g., once per second) at negligible cost to average energy per classification and throughput. Behavioral Monte Carlo simulations were performed to determine the amount of comparator noise, offset, and unit capacitance mismatch that the CNN can tolerate without degrading classification accuracy, resulting in a comparator designed for 4.6mV offset and a unit capacitance of 1fF. Since the voltage representing the weighted sum is generated at the charge storage node, parasitics at the top and bottom do not affect linearity. During convolution, the CDAC is periodically cleared (sampled 0V) as required by the top plate leakage. To prevent excessive charge from being drawn from the supply, the cell capacitor bottom plate node is discharged through switch CLR before the top capacitor is discharged through CLRe. To prevent asymmetric charge injection, the top plate switch is opened before the bottom plate voltage recovers the value set by the filter weights, image input, and bias.
[0030] like Figure 5Figure 2 shows measurements at room temperature. Ten different chips were measured to evaluate the accuracy differences due to thermal noise and mismatch in the SC neurons. At nominal supply voltages (VDD = VMEM = 1.0V, VNEU = 0.6V, VCOMP = 0.8V), the chip ran at up to 380 frames per second (FPS) and achieved 5.4 μJ / classification. Reducing VDD and VMEM to 0.8V achieved 3.8 μJ / classification at 237 FPS (1.43× reduction). The average classification accuracy was 86.05% (see histogram), which is the same as observed in the perfect digital model. The histogram spread is caused only by noise and mismatch in the SC neurons (which may lead to higher classification accuracy than the perfect digital model). The 95% confidence interval for the average classification accuracy is 86.01% to 86.10%, measured over 10 chips, each passed 30 times through the CIFAR-10 test set of 10,000 images. Not included in these energy figures is the 1.8V chip I / O energy, which amounts to 0.43μJ (a small fraction of the core energy).
[0031] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected, and the scope of the present invention to be protected is defined by the attached claims and their equivalents.
Claims
1. A mixed signal binary CNN processor, characterized in that It includes a neuron array unit, a binary temperature decoding unit, a control unit, an input image unit, an output image unit and a storage unit; the mixed signal binary CNN processor reads an RGB image, converts the channel into an N-level thermometer code through the binary temperature decoding unit, and superimposes the channels into an image, the RGB image is input through the input end of the binary temperature decoding unit, and the BinaryNet algorithm of the binary temperature decoding unit converts the RGB image into an N-level thermometer code and superimposes the codes into an image; the output end of the binary temperature decoding unit is connected to the input end of the neuron array unit through the input image unit, the input image unit includes an input image memory and an input multiplexer, and the output end of the binary temperature decoding unit is connected to the input end of the neuron array unit through the input image memory and the input multiplexer in turn; The output end of the neuron array unit is connected to the output image unit, the control unit is connected to the neuron array unit, the control instruction is input through the input end of the control unit, and the storage unit is connected to the neuron array unit.
2. The mixed signal binary CNN processor of claim 1, wherein: The storage unit comprises a local memory, a first filter memory and a second filter memory, and the local memory, the first filter memory and the second filter memory are respectively connected to the neuron array unit.
3. The mixed signal binary CNN processor of claim 2, wherein: The first filter memory and the second filter memory alternate input and output in a ping-pong manner.
4. The mixed signal binary CNN processor of claim 1, wherein: The output image unit comprises an output image memory and an output multiplexer, and the output end of the neuron array unit is outputted in sequence through the output multiplexer and the output image memory.
Citation Information
Patent Citations
A mixed signal binary CNN processor
CN209980298U