Stacked PIM Memory Channel Transfer via Base Die Buffering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of neural networks in artificial intelligence requires more computation, which is hindered by the separation of memory and processor in traditional hardware systems, leading to performance degradation due to limited data communication between them.
Innovation Solution
A processing-in-memory (PIM) system with a stacked memory device and a controller that performs sequential data move control operations to enhance data transmission efficiency by integrating arithmetic operations within the memory device, utilizing a MAC operator for matrix calculations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a general hardware system with separated memory and processor is used, then the system structure is simple and ease of manufacture is improved, but data communication between memory and processor is limited causing performance degradation
Solution Approach 1:
The patent merges the processor and memory into a single integrated device. The PIM device includes a memory unit with multiple banks and a processing unit with MAC operators that directly access the memory banks, eliminating the need for separate memory and processor components. This integration enables high-speed data processing by allowing the processing unit to perform arithmetic operations directly on data stored in the memory banks without external data transmission.
2Productivity
If the number of layers in neural network is increased to improve AI performance, then computation capability is improved, but the amount of computation required increases exponentially
Solution Approach 1:
The patent segments the computation process into multiple parallel MAC (Multiply-Accumulate) operators within the processing unit. Each MAC operator can independently perform arithmetic operations on different data sets simultaneously. This segmentation enables the system to handle complex neural network computations with multiple layers by distributing the computational load across multiple parallel operators, thereby managing exponential computation requirements through parallel processing.
Solution Approach 2:
The patent transitions from sequential processing to parallel processing by introducing multiple MAC operators that operate simultaneously on different data streams. This dimensional change from single-threaded to multi-threaded computation allows the system to process neural network layers in parallel, significantly reducing the time and power required for deep learning computations.
3Adaptability or versatility
If data is transmitted between separated memory and processor, then data storage and processing are independent, but data communication is limited causing latency
Solution Approach 1:
The patent combines memory storage and data processing functions into a single PIM device. The processing unit with MAC operators is directly integrated with the memory banks, allowing arithmetic operations to be performed on data while it resides in the memory. This eliminates the need for data to be transmitted between separate memory and processor components, thereby reducing latency while maintaining the independence of storage and processing functions through the unified architecture.
Data Source
AI summary
A memory system includes a stacked memory device and a controller. The stacked memory device includes a base die and a plurality of memory dies stacked on the base die. Each of the plurality of memory dies has a plurality of channels, and the base die is configured to function as an interface for transmitting signals and data of the pluralities of channels. The controller controls the stacked memory device such that first and second data move control operations are sequentially performed to transmit moving data from a target channel of the pluralities of channels to a destination channel of the pluralities of channels. The first data move control operation is performed to store the moving data in the target channel into the base die, and the second data move control operation is performed to write the moving data stored in the base die into the destination channel.


