Processing-in-memory device for deep learning latency reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processing-in-memory (PIM) devices face limitations in performing deterministic arithmetic operations efficiently, particularly in deep learning processes, due to the separation of memory and processor units, leading to degraded performance and increased data communication latency.
Innovation Solution
A PIM device is designed with integrated memory and processor units, featuring memory banks, multiplication/accumulation (MAC) operators, and a command decoder to perform deterministic MAC arithmetic operations, allowing for simultaneous data processing and operation in an accelerator mode, thereby reducing latency and enhancing performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If memory and processor are separated in a general hardware system, then device complexity is reduced and ease of manufacture is improved, but data processing speed deteriorates and latency increases due to data communication limitations between memory and processor
Solution Approach 1:
The patent merges memory and processor functions into a single integrated device. The memory device includes memory cells for data storage and arithmetic logic units for performing arithmetic operations directly on the stored data, eliminating the need for separate memory and processor components and reducing data communication latency.
Solution Approach 2:
The memory device is designed to perform multiple functions: it can store data in normal mode and perform arithmetic operations in accelerator mode. The same memory cells serve both as storage elements and as operands for arithmetic operations, making the device versatile for both memory and computing tasks.
2Productivity
If the number of layers in neural network is increased to improve artificial intelligence performance, then computational capability is enhanced, but the amount of computations required increases exponentially leading to degraded performance due to data communication limitations
Solution Approach 1:
The patent extracts the arithmetic operation capability directly from the processor and places it within the memory device itself. This allows arithmetic operations to be performed on data while it resides in memory, eliminating the need for repeated data transfers between memory and processor that would otherwise occur in multi-layer neural network computations.
Solution Approach 2:
The memory device performs arithmetic operations on data while it is stored in memory, before the data would otherwise need to be transferred to a separate processor. This preliminary action reduces the total time required for computational tasks by eliminating subsequent data communication steps.
3Speed
If arithmetic operations are performed directly in the PIM device using integrated memory and processor, then data processing speed is improved, but device complexity increases compared to separated memory and processor systems
Solution Approach 1:
The patent combines memory cells and arithmetic logic units into a single integrated structure. The arithmetic logic units are directly coupled to the memory cells, allowing arithmetic operations to be performed on data while it remains stored in the memory, thereby eliminating data transfer overhead and improving processing speed despite the increased internal complexity.
Data Source
AI summary
A processing-in-memory (PIM) device includes memory banks configured to perform a read operation and a write operation in a normal mode, and to perform a first data providing operation in an accelerator mode, a global buffer configured to perform a second data providing operation in the accelerator mode, processing elements configured to perform at least one of a first arithmetic operation and a second arithmetic operation using at least one of the first data and the second data in the accelerator mode, a command decoder configured to output a normal mode control signal or an accelerator mode start signal, and a processor unit configured to store an operation instruction set transmitted from an external device, to transmit the operation instruction set to the processing elements, and to transmit the accelerator mode control signal to the processing elements.


