Compute-in-Memory Circuit Design for MAC Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning applications face significant bottlenecks due to the high energy consumption and time required for data transfer between memory and processing elements, which limits the performance of machine learning computations.

Innovation Solution

The implementation of a compute-in-memory (CIM) circuit that performs operations like dot-product and absolute difference of vectors locally within an array of memory cells, reducing the need for data transfer and enhancing memory bandwidth, using techniques such as current summing and analog output interpretation to compute MAC values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in separate memory resources, then data capacity is improved, but data transfer time and energy consumption increase significantly

Engineering Contradiction:
Improvedata capacityVSAvoiddata transfer time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent merges memory and processing elements into a unified compute-in-memory structure where memory cells directly perform computational operations. This eliminates the separate memory-processor architecture, allowing data to be processed in-place without transfer, thus resolving the contradiction between maintaining large data capacity and reducing data transfer time.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces bitline voltages as intermediaries that carry both data storage and computational functions. By using voltage levels on bitlines to represent data and simultaneously perform MAC operations, the system eliminates the need for separate data transfer pathways, addressing the time loss from data movement while preserving data capacity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data is stored in separate memory resources, then data capacity is improved, but energy consumption for data transfer increases

Engineering Contradiction:
Improvedata capacityVSAvoiddata transfer energy
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent combines memory storage and computational processing into a single integrated structure. Memory cells perform MAC operations directly using bitline voltages, eliminating the energy-intensive data transfer between separate memory and processing units while maintaining full data capacity for deep learning operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The memory cells serve themselves by performing computational operations internally without external processing assistance. Each memory cell uses its stored data and bitline voltages to autonomously compute MAC values, eliminating the need for energy-consuming data movement to separate processors.

Inventive Principle:
Principle #25Self-service

3Speed

If compute operations are performed separately from memory, then processing speed is improved, but data movement overhead increases

Engineering Contradiction:
Improveprocessing speedVSAvoiddata movement overhead
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent merges computation and memory access into a single simultaneous operation. MAC computations are performed directly within the memory array using bitline voltages during the same time cycle as data access, eliminating sequential data movement overhead and achieving high processing speed without time loss from data transfer.

Inventive Principle:
Principle #5Merging (Combining)

4Speed

If large caches are placed close to processor, then data access speed is improved, but system cost increases prohibitively

Engineering Contradiction:
Improvedata access speedVSAvoidsystem cost
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent extracts the computational function from separate processing units and places it directly within the memory array. This eliminates the need for large expensive caches near processors, as data access and computation occur simultaneously in-place, achieving fast data access without the prohibitive cost of large on-processor caches.

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach reduces data movement and energy consumption, accelerating deep-learning algorithms and improving performance by embedding compute operations within memory, thereby overcoming traditional processing limitations.

Implementation Method 1

An operational amplifier may be coupled with the bitline. The operational amplifier may cause the selected bitcells to output a current independent of a state stored in the selected bitcells.

Methodology Applied
Scientific EffectOperational amplifier feedback: Feedback

Implementation Method 2

analog processor to sense an analog voltage output from the operational amplifier and convert the analog voltage to a digital value

Methodology Applied
Scientific EffectAnalog-to-digital conversion:

Data Source

PatentUS10877752B2Techniques for current-sensing circuit design for compute-in-memory
Publication Date: 2020.12.29 INTEL CORP
  • US10877752B2 patent drawing
  • US10877752B2 patent drawing
  • US10877752B2 patent drawing

AI summary

A compute-in-memory (CIM) circuit that enables a multiply-accumulate (MAC) operation based on a current-sensing readout technique. An operational amplifier coupled with a bitline of a column of bitcells included in a memory array of the CIM circuit to cause the bitcells to act like ideal current sources for use in determining an analog voltage value outputted from the operational amplifier for given states stored in the bitcells and for given input activations for the bitcells. The analog voltage value sensed by processing circuitry of the CIM circuit and converted to a digital value to compute a multiply-accumulate (MAC) value.