Cross-Layer Reconfigurable SRAM Compute-In-Memory Macro

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current compute-in-memory (CIM) technologies face challenges in achieving sufficient reconfigurability to adapt to the rapid iteration and diversification of algorithms, particularly for edge neural networks, with existing designs failing to effectively reconfigure computational precision, kernel size, model size, and other non-intelligent algorithms with low power consumption and minimal hardware overhead.

Innovation Solution

A cross-layer reconfigurable SRAM-based CIM macro and method that incorporates a 6T SRAM cell with separate wordlines for bitline control, column-shared reconfigurable Boolean computation cells, and additional transistors for supporting various Boolean operations, enabling reconfiguration and reducing hardware overhead through pipelined bit-serial and bit-parallel additions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If computational logic is added to the SRAM cell level, then Boolean computation capability is improved, but area overhead increases significantly

Engineering Contradiction:
ImproveBoolean computation capabilityVSAvoidSRAM cell area
Core Design Contradiction:
Adaptability or versatilityVSArea of moving object

Solution Approach 1:

The patent segments the computational functionality by separating Boolean computation cells from the main SRAM array. The SRAM array maintains its original storage function while dedicated computation cells handle Boolean operations, avoiding area overhead in the main array and enabling independent optimization of both storage and computation components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces peripheral circuits as intermediary components between the SRAM array and computation logic. These peripheral circuits read data from the SRAM array and feed it to Boolean computation cells, acting as a mediator that enables computation without directly modifying the SRAM cell structure and minimizing area overhead.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Area of moving object

If fixed peripheral computation circuits are used, then area overhead is reduced, but reconfigurability for different algorithms is lost

Engineering Contradiction:
Improveperipheral circuit areaVSAvoidalgorithm adaptability
Core Design Contradiction:
Area of moving objectVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic reconfigurability in peripheral computation circuits through configurable logic units that can be programmed to perform different Boolean operations. Control signals dynamically adjust the behavior of these circuits, allowing the same hardware to adapt to different algorithms and computational requirements without physical reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent designs universal peripheral computation circuits that can perform multiple Boolean operations (AND, OR, XOR, NOT, etc.) using the same hardware resources. This multi-functionality is achieved through configurable logic units that can be programmed to implement different logical functions, reducing area overhead while maintaining algorithm adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If 8T SRAM cells are used for cache accelerator, then computation coverage is improved, but area overhead increases

Engineering Contradiction:
Improvecomputation coverageVSAvoidSRAM cell area
Core Design Contradiction:
Adaptability or versatilityVSArea of moving object

Solution Approach 1:

The patent applies local quality by using different SRAM cell configurations in different regions. The main SRAM array uses conventional 6T or 8T cells optimized for storage, while dedicated computation regions use specialized Boolean computation cells with transistor configurations optimized for logic operations. This allows each region to have the quality needed for its specific function.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent adds a computational dimension to the traditional storage-focused SRAM architecture by stacking computation functionality above or alongside the storage array. This dimensional expansion allows the system to perform both storage and computation functions without significantly increasing the footprint of the storage region, effectively utilizing vertical or lateral space for computation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If reconfigurable computation modes are added, then algorithm adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvecomputation mode flexibilityVSAvoidcircuit configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-configuring control logic and selection circuits that prepare the computation units for different operations before actual computation begins. Configuration registers and control units are designed to receive operation parameters in advance, allowing the computation circuits to be rapidly reconfigured without complex real-time adjustments during computation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12112797B1Cross-layer reconfigurable static random access memory (SRAM) based compute-in-memory macro and method for edge intelligence
Publication Date: 2024.10.08 SHANGHAI JIAOTONG UNIV
  • US12112797B1 patent drawing
  • US12112797B1 patent drawing
  • US12112797B1 patent drawing

AI summary

The present disclosure provides a cross-layer reconfigurable static random access memory (SRAM) based compute-in-memory (CIM) macro and method for edge intelligence. The cell includes a SRAM cell and a column-shared reconfigurable Boolean computation cell. A reconfiguration computation is performed based on the SRAM cell to obtain a reconfigured structure; the column-shared reconfigurable Boolean computation cell outputs a computation result based on the reconfigured structure; and a peripheral computation circuit supporting pipelined bit-serial addition outputs an in-memory addition result on a basis of a Boolean computation. In order to meet a requirement of edge artificial intelligence (AI) for a low power consumption and a low hardware overhead, and to enable an accelerator to adapt to a fast iterative software algorithm as much as possible.