In-Memory Compute Cell Using Ternary Weights for Low-Variation MAC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current in-memory computing devices for machine learning applications face challenges in scalability and energy efficiency due to limited dynamic range of analog word line voltages and variations in bit cell currents, especially in larger array sizes, which hinder accurate reproduction of model weights and increase energy consumption.
Innovation Solution
The implementation of a compute cell with a memory unit storing ternary weights and a logic unit that selectively enables conductive paths for charging and discharging read bit lines based on the signs of the weights and input data, allowing for more dense storage and wider multiplication operations, thereby enhancing energy efficiency and reducing variations across compute cells.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If standard 6T SRAM cells are operated in sub-threshold mode to aggregate low bit cell currents, then energy efficiency is improved, but variations in bit cell currents increase and dynamic range of analog word line voltages is limited
Solution Approach 1:
The patent changes the operating parameters by using standard threshold voltage mode instead of sub-threshold mode, and implements a two-stage computation approach with digital preprocessing. This allows the access transistors to operate in their optimal threshold region, reducing variations in bit cell currents while maintaining energy efficiency through the in-memory computing architecture.
Solution Approach 2:
The computation is segmented into two stages: a digital preprocessing stage that computes partial results, and an analog in-memory stage that completes the multiplication-accumulate operations. This segmentation allows each stage to operate in its optimal mode, with the digital stage handling sign-bit operations and the analog stage handling magnitude computations, thereby reducing overall variations and improving reliability.
2Productivity
If RRAM technology is used for in-memory computing, then scalability is improved, but voltage drops on read bit lines become critical for larger array sizes
Solution Approach 1:
The patent introduces digital preprocessing circuitry as an intermediary that computes partial results before the analog in-memory computation. This intermediary digital stage reduces the complexity and magnitude of operations required in the RRAM array, thereby reducing voltage drops on read bit lines and enabling better scalability for larger array sizes.
3Adaptability or versatility
If model data is stored at a location physically separated from the computational site, then flexibility is improved, but energy efficiency and throughput deteriorate due to constant data retrieval and transfer
Solution Approach 1:
The patent merges the storage and computation functions by implementing an in-memory computing architecture where model weights are stored directly in the RRAM array at the same location where multiplication-accumulate operations are performed. This eliminates the need for constant data retrieval and transfer between separate storage and computational sites, thereby dramatically improving energy efficiency and throughput while maintaining the flexibility to load different models.
Data Source
AI summary
A compute cell for in-memory multiplication of a digital data input and a balanced ternary weight, and an in-memory computing device including an array of the compute cells, are provided. In one aspect, the compute cell includes a set of input connectors for receiving modulated input signals representative of a sign and a magnitude of the data input, and a memory unit configured to store the ternary weight. A logic unit connected to the set of input connectors and the memory unit receives the data input and the ternary weight. The logic unit selectively enables one of a plurality of conductive paths for supplying a partial charge to a read bit line during a compound duty cycle of the set of input signals as a function of the respective signs of data input and ternary weight, and disables each of the plurality of conductive paths if at least one of the ternary weight and data input have zero magnitude.


