PIM Device MAC Circuit Zero-Point Selection for Neural Network Speed
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing-in-memory (PIM) systems face limitations in deep learning applications due to the separation of memory and processor, leading to degraded performance from limited data communication between them, necessitating an integrated solution for improved neural network computation.
Innovation Solution
A PIM device with a data selection circuit, multiplying-and-accumulating (MAC) circuit, and accumulative adding circuit, which generates selection data, performs MAC operations, and accumulatively adds MAC sign data based on zero-point selection signals, enabling efficient arithmetic operations within the device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If memory and processor are separated in general hardware systems, then device complexity is reduced and manufacturing is easier, but data communication between memory and processor is limited and AI performance is degraded
Solution Approach 1:
The patent merges memory and processor into a single integrated PIM device, where memory cells store data and dedicated MAC circuits perform arithmetic operations directly on the stored data. This integration eliminates the need for separate memory and processor components, resolving the contradiction by combining previously separate functions into one unified device that achieves both high AI performance and manageable complexity.
2Productivity
If the number of layers in neural networks is increased to improve AI performance, then computation capability is enhanced, but the amount of computation required increases exponentially
Solution Approach 1:
The patent extracts the arithmetic computation function from traditional processors and places it directly within the memory device. By embedding MAC circuits inside the memory array, the system performs computations in-place without requiring data to be moved to separate processing units. This extraction of computation from memory and its integration into the memory fabric enables deep neural networks to run with reduced power consumption and without requiring external high-performance processors.
3Productivity
If data is frequently communicated between separate memory and processor, then processing capability is maintained, but communication limitations degrade AI performance
Solution Approach 1:
The memory device performs arithmetic operations on its own stored data using integrated MAC circuits, eliminating the need to transfer data to external processors. The memory serves both as storage and computation unit, performing MAC operations directly on the data residing in memory cells. This self-service capability eliminates data communication bottlenecks and enables high-speed processing for deep learning applications.
Data Source
AI summary
A processing-in-memory (PIM) device includes a data selection circuit, a multiplying-and-accumulating (MAC) circuit, and an accumulative adding circuit. The data selection circuit generates selection data from input data and zero-point data based on a zero-point selection signal. The MAC circuit performs a MAC arithmetic operation for the selection data to generate MAC result data. The accumulative adding circuit accumulatively adds MAC sign data based on a MAC output latch signal to generate MAC latch data. A sign of the MAC sign data is determined by the zero-point selection signal.


