A programmable in-memory computing (IMC) accelerator for low-precision deep neural network
inference, also referred to as PIMCA, is provided. Embodiments of the PIMCA integrate a large number of capacitive-
coupling-based IMC static random-access memory (SRAM) macros and demonstrate large-scale integration of IMC SRAM macros. For example, a 28 nm prototype integrates 108 capacitive-
coupling-based IMC SRAM macros of a total size of 3.4 megabytes (Mb), demonstrating one of the largest IMC hardware to date. In addition, a
custom instruction set architecture (ISA) is developed featuring IMC and single-instruction-multiple-data (
SIMD) functional units with hardware loop to support a range of deep neural network (DNN) layer types. The 28 nm prototype
chip achieves a peak
throughput of 4.9 tera operations per second (TOPS) and
system-level peak energy-efficiency of 437 TOPS per
watt (TOPS / W) at 40 megahertz (MHz) with a 1
volt (V) supply.