In-Memory SRAM Accelerator ISA for Flexible Low-Precision DNN Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing deep neural network (DNN) accelerators face challenges in integrating sufficient IMC SRAM macros, flexibility to support various DNN layers, and efficient handling of generic nested loops, leading to limited system-level throughput and energy efficiency.
Innovation Solution
A programmable in-memory computing accelerator (PIMCA) with 108 capacitive-coupling-based IMC SRAM macros and a custom ISA, featuring SIMD functional units and hardware loop support, enabling large-scale integration and efficient computation of diverse DNN layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a large number of IMC SRAM macros are integrated on-chip, then system-level throughput and energy efficiency are improved, but device complexity and manufacturing difficulty increase
Solution Approach 1:
The system is divided into multiple IMC processing elements (PEs), each containing a set of IMC macros that can operate independently and in parallel. This segmentation allows the large-scale integration of 108 IMC SRAM macros to be managed through modular units, improving system-level throughput while controlling complexity through standardized building blocks.
Solution Approach 2:
The IMC PEs are designed with universal functionality to support various DNN layer types through a custom instruction set architecture. This multi-functionality allows the same hardware structure to handle different computational tasks, reducing the need for specialized circuits and managing device complexity while maintaining high productivity.
2Productivity
If hardware loop support is omitted, then device complexity is reduced, but loss of time and productivity decrease due to large overhead in latency and instruction counts
Solution Approach 1:
Hardware loop support is pre-integrated into the IMC PE architecture, allowing loop operations to be executed efficiently without software overhead. The loop control logic is built into the hardware, enabling preliminary setup and automatic iteration through DNN layers, which reduces execution time and maintains high productivity without excessive complexity.
3Adaptability or versatility
If data flow is hard-wired, then device complexity is reduced, but adaptability and versatility are limited to support layer types other than batch normalization and activation layers
Solution Approach 1:
The data flow architecture is designed to be dynamic and reconfigurable through a custom instruction set, allowing the IMC PEs to adapt to different DNN layer types. Control signals and instruction sequences dynamically route data through appropriate computational paths, providing versatility for various layer types while managing complexity through software-controlled flexibility rather than hard-wired constraints.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The PIMCA achieves a peak throughput of 4.9 tera operations per second and system-level peak energy-efficiency of 437 TOPS/W, significantly reducing on-chip memory access overhead and supporting a wide range of DNN operations.
Implementation Method 1
a plurality of in-memory computing (IMC) processing elements (PEs), each comprising a set of capacitive-coupling-based IMC static random-access memory (SRAM) macros
Data Source
AI summary
A programmable in-memory computing (IMC) accelerator for low-precision deep neural network inference, also referred to as PIMCA, is provided. Embodiments of the PIMCA integrate a large number of capacitive-coupling-based IMC static random-access memory (SRAM) macros and demonstrate large-scale integration of IMC SRAM macros. For example, a 28 nm prototype integrates 108 capacitive-coupling-based IMC SRAM macros of a total size of 3.4 megabytes (Mb), demonstrating one of the largest IMC hardware to date. In addition, a custom instruction set architecture (ISA) is developed featuring IMC and single-instruction-multiple-data (SIMD) functional units with hardware loop to support a range of deep neural network (DNN) layer types. The 28 nm prototype chip achieves a peak throughput of 4.9 tera operations per second (TOPS) and system-level peak energy-efficiency of 437 TOPS per watt (TOPS/W) at 40 megahertz (MHz) with a 1 volt (V) supply.


