Segmented In-Memory Computing Circuit for Multi-Bit MVM Throughput
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current analog in-memory computation circuits face inefficiencies in processing high-dimensional Matrix Vector Multiplications due to limitations in the storage and processing of multi-bit weights in SRAM arrays, which affect the precision and throughput of neural network operations.
Innovation Solution
The proposed solution involves segmenting the memory array into sub-arrays with local and global bit lines, where each sub-array has a single word line actuated during computation, enabling charge sharing between local and global bit lines to enhance precision and throughput, and supporting multi-bit weights through capacitive or switched coupling, allowing for efficient digital signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the memory array is segmented into sub-arrays with local and global bit lines, then precision of output sensing is improved, but device complexity increases
Solution Approach 1:
The memory array is divided into multiple sub-arrays, each with its own local bit lines that connect to shared global bit lines. This segmentation allows charge from multiple memory cells to be accumulated and sensed together, improving measurement precision while managing complexity through hierarchical organization
Solution Approach 2:
Local bit lines from multiple sub-arrays are merged into shared global bit lines that connect to the sensing circuitry. This combining approach enables parallel processing of multiple memory rows while using a reduced number of global bit lines, improving precision without proportionally increasing device complexity
2Productivity
If row parallelism is enhanced through segmented memory architecture, then computational throughput is improved, but device complexity increases
Solution Approach 1:
The memory array is segmented into multiple sub-arrays that can be accessed in parallel, enabling simultaneous computation operations across multiple rows. This increases computational throughput by allowing parallel processing while managing complexity through the structured sub-array organization
Solution Approach 2:
The segmented architecture with shared global bit lines provides multi-functionality, allowing the same global bit line infrastructure to serve multiple sub-arrays for different computation operations. This universal approach enables enhanced row parallelism and throughput without requiring dedicated infrastructure for each sub-array
3Measurement precision
If multi-bit weights are supported through capacitive or switched coupling, then precision of computation is improved, but device complexity increases
Solution Approach 1:
Capacitive coupling or switched coupling mechanisms are introduced as intermediaries between the memory cells and global bit lines. These intermediaries enable multi-bit weight representation by allowing controlled charge transfer and accumulation, improving computation precision while managing complexity through the use of simple capacitive or switch-based elements
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach increases the precision of output sensing, enhances row parallelism, and improves computational throughput while managing large neural network layers with minimal circuit area impact, effectively addressing the inefficiencies in existing technologies.
Implementation Method 1
enabling charge sharing between local and global bit lines to enhance precision and throughput
Data Source
AI summary
A memory array includes sub-arrays with memory cells arranged in a row-column matrix where each row includes a word line and each sub-array column includes a local bit line. A control circuit supports a first operating mode where only one word line in the memory array is actuated during memory access and a second operating mode where one word line per sub-array is simultaneously actuated during an in-memory computation performed as a function of weight data stored in the memory and applied feature data. Computation circuitry coupling each memory cell to the local bit line for each column of the sub-array logically combines a bit of feature data for the in-memory computation with a bit of weight data to generate a logical output on the local bit line which is charge shared with the global bit line.


