Channel-Level Processing Element for In-Memory Neural Network Computation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory devices face inefficiencies in processing neural network operations due to high latency and external bus traffic, as they rely heavily on host processors for computation, which limits performance and increases power consumption.
Innovation Solution
A memory device with a channel-level processing element (PE) that generates in-memory computation results by performing operations on partial results from multiple memory banks, using operators and an adder to sum outputs, and transmitting these results to a host processor, thereby reducing the computational load on the host and enhancing internal bandwidth utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If memory devices rely on host processors for computation, then device complexity is reduced, but processing speed and performance deteriorate
Solution Approach 1:
The patent merges computation functions with memory storage functions by integrating processing elements directly into the memory device. The channel-level PEs perform computations on data stored in memory banks, combining what were previously separate memory and processing components into a unified system, thereby improving processing speed without proportionally increasing overall system complexity
Solution Approach 2:
The patent introduces a new dimensional organization by dividing memory banks into multiple channels (first channel, second channel, etc.) and placing channel-level processing elements at the channel level. This channel-level dimension enables parallel processing across channels while maintaining a manageable hierarchical structure that doesn't linearly increase device complexity
2Device complexity
If computations are performed externally via host processor, then memory device simplicity is maintained, but power consumption increases
Solution Approach 1:
The memory device performs computations itself through integrated channel-level processing elements rather than relying on external host processors. The PEs execute operations on data locally stored in memory banks, enabling the memory device to serve its own processing needs and reduce the computational burden on external systems, thereby reducing overall power consumption
Solution Approach 2:
The patent segments the processing function into channel-level units that are distributed across multiple memory channels. This segmentation allows computations to be performed in parallel across channels, improving processing efficiency while keeping each individual processing element simple and power-efficient
3Device complexity
If data is processed through external buses, then memory device internal structure is simplified, but latency increases
Solution Approach 1:
By merging computation and memory access operations within the same device, the patent eliminates the need to transfer intermediate data through external buses. The channel-level PEs can directly access data in memory banks and write results back without external bus transactions, significantly reducing latency while maintaining a relatively simple internal hierarchical structure
Solution Approach 2:
The channel-level processing elements perform computations on data while it is already present in the memory banks, before any external bus transfer is needed. This preliminary computation action reduces the amount of data that needs to be transferred externally and eliminates waiting time for bus transactions, thereby reducing overall latency
4Device complexity
If host processor handles all computations, then memory device functionality is simple, but productivity decreases
Solution Approach 1:
The patent combines memory storage functionality with computation functionality in a single integrated device. The channel-level processing elements enable the memory device to perform computations on stored data, effectively doubling its functional capabilities and significantly improving productivity by eliminating the need for constant data transfer to external processors
Solution Approach 2:
By organizing processing elements at the channel level rather than requiring a single external processor, the patent creates a parallel processing dimension. Multiple channels can process data simultaneously through their respective PEs, dramatically increasing overall processing throughput while keeping each channel's processing element functionally simple
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
A memory device includes: a plurality of memory banks divided by a plurality of channels comprising a first channel and a second channel; and a channel-level processing element (PE) configured to generate an in-memory computation result by performing an operation using a first partial result generated based on data stored in a memory bank of the first channel among the plurality of memory banks and a second partial result generated based on data stored in a memory bank of the second channel among the plurality of memory banks.