CIM Memory Unit With AGMI and CVSS for Low-Latency MAC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional memory units for non-volatile computing-in-memory applications face challenges such as long input latency, limited system-level inference accuracy due to small signal margin, and high power consumption, particularly in battery-powered AI edge devices requiring high precision MAC computing.
Innovation Solution
A memory unit with an asymmetric group-modulated input (AGMI) scheme and a current-to-voltage signal stacking (CVSS) scheme, which includes non-volatile memory cells, a source line, a bit line, a controller, and a CVSS converter, splits multi-bit input signals into sub-groups, generates switching signals, and converts bit-line currents into output voltages, reducing latency and energy consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If conventional fully-decoded wordline pulse-count input scheme is used, then memory unit can perform basic computing operations, but computing latency is long due to multiple cycles required for applying inputs
Solution Approach 1:
The patent segments the input signal into multiple sub-signals that are applied to different word lines simultaneously. This allows parallel processing of input data across multiple memory cells in a single cycle, eliminating the sequential multiple-cycle approach of conventional schemes and thereby reducing computing latency while improving productivity
Solution Approach 2:
The patent introduces a new dimension by simultaneously utilizing multiple word lines for input application rather than sequential activation. This dimensional expansion from single-word-line sequential processing to multi-word-line parallel processing enables inputs to be applied in parallel across different memory cells, significantly reducing the time required for computing operations
2Measurement precision
If conventional input scheme is used, then memory unit operates with simple architecture, but signal margin is small leading to limited inference accuracy
Solution Approach 1:
The patent segments the input signal into multiple sub-signals distributed across different word lines. This segmentation allows each word line to carry a portion of the input data, increasing the overall signal margin when multiple word lines are activated simultaneously. The enhanced signal margin improves inference accuracy while the modular segmentation approach keeps the complexity manageable
Solution Approach 2:
The patent merges multiple word line signals into a unified computing operation. By combining the contributions from multiple simultaneously activated word lines, the system achieves enhanced signal margin and improved inference accuracy. The merging of parallel word line inputs allows the system to leverage multiple data paths concurrently, boosting overall signal strength without proportionally increasing complexity
3Use of energy by moving object
If conventional scheme is used, then memory unit consumes acceptable power, but power consumption increases significantly in readout circuit due to large amount of DC current
Solution Approach 1:
The patent employs periodic action by utilizing pulse-width modulated signals instead of continuous DC current. The input signals are applied in controlled time windows within each computing cycle, and the readout operation uses pulsed current rather than continuous current. This periodic action reduces the average power consumption in the readout circuit while maintaining the necessary signal strength for accurate computation, directly addressing the high DC current consumption problem
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The AGMI and CVSS schemes significantly reduce computing latency, increase signal margin, and decrease energy consumption, achieving high precision MAC computing with improved energy efficiency and inference accuracy.
Implementation Method 1
The non-volatile memory cells are controlled by a plurality of word lines to generate a plurality of memory cell currents and storing the weights
Implementation Method 2
The CVSS converter is electrically connected to the non-volatile memory cells via the bit line. The CVSS converter is electrically connected to the controller and converts the bit-line current into a plurality of converted voltages according to the input sub-groups and the switching signals
Data Source
AI summary
A memory unit with an asymmetric group-modulated input scheme and a current-to-voltage signal stacking scheme for a plurality of non-volatile computing-in-memory applications is configured to compute a plurality of multi-bit input signals and a plurality of weights. A controller splits the multi-bit input signals into a plurality of input sub-groups and generates a plurality of switching signals according to the input sub-groups, and the input sub-groups are sequentially inputted to the word lines. The current-to-voltage signal stacking converter converts the bit-line current from a plurality of non-volatile memory cells into a plurality of converted voltages according to the input sub-groups and the switching signals, and the current-to-voltage signal stacking converter stacks the converted voltages to form an output voltage. The output voltage is corresponding to a sum of a plurality of multiplication values which are equal to the multi-bit input signals multiplied by the weights.


