Cross-Layer Reconfigurable SRAM Compute-In-Memory Macro
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current compute-in-memory (CIM) technologies face challenges in achieving sufficient reconfigurability to adapt to the rapid iteration and diversification of algorithms, particularly for edge neural networks, with existing designs failing to effectively reconfigure computational precision, kernel size, model size, and other non-intelligent algorithms with low power consumption and minimal hardware overhead.
Innovation Solution
A cross-layer reconfigurable SRAM-based CIM macro and method that incorporates a 6T SRAM cell with separate wordlines for bitline control, column-shared reconfigurable Boolean computation cells, and additional transistors for supporting various Boolean operations, enabling reconfiguration and reducing hardware overhead through pipelined bit-serial and bit-parallel additions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If computational logic is added to the SRAM cell level, then Boolean computation capability is improved, but area overhead increases significantly
Solution Approach 1:
The patent segments the computational functionality by separating Boolean computation cells from the main SRAM array. The SRAM array maintains its original storage function while dedicated computation cells handle Boolean operations, avoiding area overhead in the main array and enabling independent optimization of both storage and computation components.
Solution Approach 2:
The patent introduces peripheral circuits as intermediary components between the SRAM array and computation logic. These peripheral circuits read data from the SRAM array and feed it to Boolean computation cells, acting as a mediator that enables computation without directly modifying the SRAM cell structure and minimizing area overhead.
2Area of moving object
If fixed peripheral computation circuits are used, then area overhead is reduced, but reconfigurability for different algorithms is lost
Solution Approach 1:
The patent implements dynamic reconfigurability in peripheral computation circuits through configurable logic units that can be programmed to perform different Boolean operations. Control signals dynamically adjust the behavior of these circuits, allowing the same hardware to adapt to different algorithms and computational requirements without physical reconfiguration.
Solution Approach 2:
The patent designs universal peripheral computation circuits that can perform multiple Boolean operations (AND, OR, XOR, NOT, etc.) using the same hardware resources. This multi-functionality is achieved through configurable logic units that can be programmed to implement different logical functions, reducing area overhead while maintaining algorithm adaptability.
3Adaptability or versatility
If 8T SRAM cells are used for cache accelerator, then computation coverage is improved, but area overhead increases
Solution Approach 1:
The patent applies local quality by using different SRAM cell configurations in different regions. The main SRAM array uses conventional 6T or 8T cells optimized for storage, while dedicated computation regions use specialized Boolean computation cells with transistor configurations optimized for logic operations. This allows each region to have the quality needed for its specific function.
Solution Approach 2:
The patent adds a computational dimension to the traditional storage-focused SRAM architecture by stacking computation functionality above or alongside the storage array. This dimensional expansion allows the system to perform both storage and computation functions without significantly increasing the footprint of the storage region, effectively utilizing vertical or lateral space for computation.
4Adaptability or versatility
If reconfigurable computation modes are added, then algorithm adaptability is improved, but device complexity increases
Solution Approach 1:
The patent implements preliminary action by pre-configuring control logic and selection circuits that prepare the computation units for different operations before actual computation begins. Configuration registers and control units are designed to receive operation parameters in advance, allowing the computation circuits to be rapidly reconfigured without complex real-time adjustments during computation.
Data Source
AI summary
The present disclosure provides a cross-layer reconfigurable static random access memory (SRAM) based compute-in-memory (CIM) macro and method for edge intelligence. The cell includes a SRAM cell and a column-shared reconfigurable Boolean computation cell. A reconfiguration computation is performed based on the SRAM cell to obtain a reconfigured structure; the column-shared reconfigurable Boolean computation cell outputs a computation result based on the reconfigured structure; and a peripheral computation circuit supporting pipelined bit-serial addition outputs an in-memory addition result on a basis of a Boolean computation. In order to meet a requirement of edge artificial intelligence (AI) for a low power consumption and a low hardware overhead, and to enable an accelerator to adapt to a fast iterative software algorithm as much as possible.


