In-Memory MAC Architecture With Shared DAC and Charge Accumulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current in-memory computing (IMC) technologies for AI applications face power efficiency issues due to latencies and high power overheads associated with converting digital byte streams to analog for MAC operations, particularly in neural network architectures.

Innovation Solution

An analog IMC architecture with a MAC core that includes an array of multilevel non-volatile memory cells, shared bit-lines, and charge-storage banks, along with analog-to-digital converters, which allows for high-speed conversion of digital byte streams to analog and efficient accumulation and conversion of weighted bit-line currents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If digital-to-analog conversion is performed using DAC for each row in memory array, then analog MAC operations can be executed, but power overhead increases significantly

Engineering Contradiction:
Improvepower efficiencyVSAvoidpower overhead
Core Design Contradiction:
PowerVSUse of energy by stationary object

Solution Approach 1:

The patent extracts the digital-to-analog conversion function from individual row DACs and consolidates it into a single shared DAC. This eliminates redundant conversion circuits across multiple rows, significantly reducing power overhead while maintaining the capability to perform analog MAC operations in the memory array.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The shared DAC serves all rows in the memory array simultaneously, making a single conversion unit universal for the entire array. This multi-functional approach allows one DAC to provide analog signals to multiple rows, reducing the total number of DAC instances and their associated power consumption.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If digital-to-analog conversion is performed before accessing memory rows, then analog MAC operations can be performed, but latency increases due to conversion time

Engineering Contradiction:
Improvedata processing speedVSAvoidconversion latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent prepares digital input data in advance and buffers it before the actual MAC operation begins. This preliminary preparation allows the shared DAC to convert data continuously without waiting for row-by-row processing, thereby reducing conversion latency and improving overall data processing speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The shared DAC operates continuously to convert digital input data into analog signals throughout the MAC operation, rather than performing discrete conversions for each row. This continuous operation eliminates idle conversion time and maintains steady data flow to the memory array, reducing latency.

Inventive Principle:
Principle #20Continuity of useful action

3Adaptability or versatility

If multiple DACs are used for each row to enable analog MAC operations, then conversion capability is improved, but device complexity increases

Engineering Contradiction:
Improveconversion capabilityVSAvoidarchitecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent removes the complexity of multiple row-specific DACs by extracting and consolidating the conversion function into a single shared DAC. This simplification reduces device complexity while maintaining full conversion capability for all rows through time-multiplexed or parallel access mechanisms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The shared DAC is designed to serve multiple rows universally, providing the same conversion capability that would otherwise require separate DACs for each row. This universal design reduces the number of components and interconnections, thereby lowering overall device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution enhances power efficiency and reduces latency by enabling concurrent shifting, multiplying, and accumulating of data within the memory array, improving data processing speed and power efficiency in AI applications.

Implementation Method 1

Each memory cell includes a multilevel, non-volatile memory (NVM) device... each capable of storing a weight value in an analog format

Methodology Applied
Scientific EffectAnalog storage:

Implementation Method 2

activate the NVM devices based on a state of the bit to produce a weighted bit-line current from each activated NVM device proportional to a product of the bit and a weight stored in the NVM device

Methodology Applied
Scientific EffectOhm's law: Ohm's Law

Implementation Method 3

A plurality of first charge-storage banks... configured receive a sum of weighted bit-line currents and to accumulate for each bit of the input bytes charge produced by the sum of weighted bit-line currents

Methodology Applied
Scientific EffectCharge accumulation: Capacitance

Implementation Method 4

plurality of second charge-storage banks coupled to a number of analog-to-digital converters (ADCs)... configured to concurrent with the shifting and accumulating, to provide scaled voltages for each bit of previously received second input bytes to the ADC for conversion into an output byte

Methodology Applied
Scientific EffectAnalog-to-digital conversion:

Data Source

PatentUS12045714B2In-memory computing architecture and methods for performing MAC operations
Publication Date: 2024.07.23 INFINEON TECHNOLOGIES LLC
  • US12045714B2 patent drawing
  • US12045714B2 patent drawing
  • US12045714B2 patent drawing

AI summary

A method of operation of a semiconductor device that includes the steps of coupling each of a plurality of digital inputs to a corresponding row of non-volatile memory (NVM) cells that stores an individual weight, initiating a read operation based on a digital value of a first bit of the plurality of digital inputs, accumulating along a first bit-line coupling a first array column weighted bit-line current, in which the weighted bit-line current corresponds to a product of the individual weight stored therein and the digital value of the first bit, and converting and scaling, an accumulated weighted bit-line current of the first column, into a scaled charge of the first bit in relation to a significance of the first bit.