Compute-in-Memory MAC Circuit for Low-Transfer AI Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computer systems face inefficiencies in processing large amounts of data due to excessive data transfers between memory and processor, leading to high power consumption and slow compute times, particularly in machine learning applications like neural networks.
Innovation Solution
Implementing Compute-in-Memory (CiM) systems that perform compute operations directly on data within memory arrays, reducing data transfer by using macros and control circuits to efficiently output MAC values while minimizing bit toggling and power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If data is processed using conventional computer hardware with separate memory and processor, then data storage capacity is sufficient, but power consumption increases and compute time slows due to excessive data transfers between memory and processor
Solution Approach 1:
The patent merges memory and processing functions into a single integrated structure where processing units are embedded within the memory array. This allows data to be processed in-place without being transferred to a separate processor, simultaneously reducing power consumption from data transfers and maintaining high compute speed through direct processing at the memory location.
Solution Approach 2:
The patent introduces control circuits as intermediaries that manage data flow and processing operations within the memory array. These control circuits coordinate between the processing units and memory cells, enabling efficient in-memory processing while minimizing the need for data movement between separate memory and processing components.
2Loss of energy
If data transfer between memory and processor is reduced through Compute-in-Memory architecture, then power consumption decreases, but device complexity increases due to integrated processing units within memory arrays
Solution Approach 1:
The patent divides the memory array into multiple segments or banks, each with its own embedded processing units. This segmentation allows the complex processing functionality to be distributed across multiple smaller, manageable units rather than requiring a single complex processing structure, making the overall system more implementable while maintaining the in-memory processing benefits.
Solution Approach 2:
The processing units embedded within the memory array are designed to perform multiple functions including data processing, data storage, and control operations. This multi-functionality reduces the need for separate dedicated components for each function, thereby managing device complexity while achieving the energy efficiency benefits of in-memory processing.
3Speed
If processing operations are performed directly within memory arrays, then compute speed increases by eliminating data transfer, but manufacturing precision requirements increase due to integration of processing units with memory cells
Solution Approach 1:
The patent employs dynamic control circuits that can adaptively manage the processing operations within the memory array. These control circuits dynamically configure data flow paths and processing operations based on the specific computational task, allowing the system to achieve high compute speed while managing manufacturing precision requirements through flexible, reconfigurable operation rather than rigid fixed-function structures.
Data Source
AI summary
An integrated circuit includes a first logic gate configured to receive a first input signal and a second input signal, and generate a first control signal based on a first bit of first input signal and a first bit of the second input signal obtained in a current cycle. The integrated circuit includes a first backup storage component configured to store a second bit of the first input signal and a second bit of the second input signal obtained in a previous cycle. The integrated circuit includes a plurality of first macros each configured to selectively compute, based on the first control signal, a first multiply-accumulate (MAC) value for the first bit of the first input signal and the first bit of the second input signal.


