In-Memory MAC Architecture With Shared DAC and Charge Accumulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current in-memory computing (IMC) technologies for AI applications face power efficiency issues due to latencies and high power overheads associated with converting digital byte streams to analog for MAC operations, particularly in neural network architectures.
Innovation Solution
An analog IMC architecture with a MAC core that includes an array of multilevel non-volatile memory cells, shared bit-lines, and charge-storage banks, along with analog-to-digital converters, which allows for high-speed conversion of digital byte streams to analog and efficient accumulation and conversion of weighted bit-line currents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If digital-to-analog conversion is performed using DAC for each row in memory array, then analog MAC operations can be executed, but power overhead increases significantly
Solution Approach 1:
The patent extracts the digital-to-analog conversion function from individual row DACs and consolidates it into a single shared DAC. This eliminates redundant conversion circuits across multiple rows, significantly reducing power overhead while maintaining the capability to perform analog MAC operations in the memory array.
Solution Approach 2:
The shared DAC serves all rows in the memory array simultaneously, making a single conversion unit universal for the entire array. This multi-functional approach allows one DAC to provide analog signals to multiple rows, reducing the total number of DAC instances and their associated power consumption.
2Productivity
If digital-to-analog conversion is performed before accessing memory rows, then analog MAC operations can be performed, but latency increases due to conversion time
Solution Approach 1:
The patent prepares digital input data in advance and buffers it before the actual MAC operation begins. This preliminary preparation allows the shared DAC to convert data continuously without waiting for row-by-row processing, thereby reducing conversion latency and improving overall data processing speed.
Solution Approach 2:
The shared DAC operates continuously to convert digital input data into analog signals throughout the MAC operation, rather than performing discrete conversions for each row. This continuous operation eliminates idle conversion time and maintains steady data flow to the memory array, reducing latency.
3Adaptability or versatility
If multiple DACs are used for each row to enable analog MAC operations, then conversion capability is improved, but device complexity increases
Solution Approach 1:
The patent removes the complexity of multiple row-specific DACs by extracting and consolidating the conversion function into a single shared DAC. This simplification reduces device complexity while maintaining full conversion capability for all rows through time-multiplexed or parallel access mechanisms.
Solution Approach 2:
The shared DAC is designed to serve multiple rows universally, providing the same conversion capability that would otherwise require separate DACs for each row. This universal design reduces the number of components and interconnections, thereby lowering overall device complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution enhances power efficiency and reduces latency by enabling concurrent shifting, multiplying, and accumulating of data within the memory array, improving data processing speed and power efficiency in AI applications.
Implementation Method 1
Each memory cell includes a multilevel, non-volatile memory (NVM) device... each capable of storing a weight value in an analog format
Implementation Method 2
activate the NVM devices based on a state of the bit to produce a weighted bit-line current from each activated NVM device proportional to a product of the bit and a weight stored in the NVM device
Implementation Method 3
A plurality of first charge-storage banks... configured receive a sum of weighted bit-line currents and to accumulate for each bit of the input bytes charge produced by the sum of weighted bit-line currents
Implementation Method 4
plurality of second charge-storage banks coupled to a number of analog-to-digital converters (ADCs)... configured to concurrent with the shifting and accumulating, to provide scaled voltages for each bit of previously received second input bytes to the ADC for conversion into an output byte
Data Source
AI summary
A method of operation of a semiconductor device that includes the steps of coupling each of a plurality of digital inputs to a corresponding row of non-volatile memory (NVM) cells that stores an individual weight, initiating a read operation based on a digital value of a first bit of the plurality of digital inputs, accumulating along a first bit-line coupling a first array column weighted bit-line current, in which the weighted bit-line current corresponds to a product of the individual weight stored therein and the digital value of the first bit, and converting and scaling, an accumulated weighted bit-line current of the first column, into a scaled charge of the first bit in relation to a significance of the first bit.


