PIM Memory Parallel Window Optimization for Convolution Cycle Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Processing-In-Memory (PIM) based arrays face inefficiencies in computing convolution layers due to fixed-size parallel windows, which are not adaptable to varying sizes of PIM arrays and convolution layers, leading to suboptimal cycle counts in deep neural network computations.
Innovation Solution
A method and memory device that dynamically determine the size of a parallel window based on the PIM array, input data, and kernel size to minimize the number of computation cycles by calculating shifts, inputs, and outputs, thereby optimizing cycle efficiency for convolution layer computations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a fixed-size parallel window is used for convolution computation in PIM array, then the computation can be simplified, but the adaptability to different PIM array sizes and convolution layer sizes is reduced
Solution Approach 1:
The patent applies dynamics by making the parallel window size configurable and adaptable rather than fixed. The system dynamically adjusts the parallel window size based on the PIM array dimensions and convolution layer parameters, allowing the same PIM array to efficiently handle various convolution configurations without hardware reconfiguration.
Solution Approach 2:
The patent changes the parameter of parallel window size from a fixed value to a variable that can be adjusted according to the specific computation requirements. By modifying this parameter based on PIM array size and convolution layer characteristics, the system achieves both computational efficiency and adaptability to different configurations.
2Productivity
If the PIM array size is large, then more computation can be performed in parallel, but the number of cycles increases when the convolution layer size is small
Solution Approach 1:
The patent applies partial action by using only the necessary portion of the PIM array for each computation cycle. Instead of always utilizing the full PIM array capacity, the system dynamically adjusts the parallel window size to match the actual computation requirements, avoiding wasted cycles from processing excessive data when the convolution layer is small.
Solution Approach 2:
The system dynamically adjusts the parallel window size based on the relationship between PIM array dimensions and convolution layer size. This dynamic adaptation allows the system to optimize the balance between parallel computation capacity and the actual number of cycles required, preventing both underutilization and over-provisioning.
3Productivity
If the parallel window size is increased to reduce computation cycles, then the computational efficiency improves, but the memory bandwidth requirement increases
Solution Approach 1:
The patent optimizes the parallel window size parameter to achieve the best trade-off between computational efficiency and memory bandwidth requirements. By calculating the optimal window size based on PIM array dimensions and convolution layer parameters, the system maximizes computational throughput while keeping memory bandwidth usage within acceptable limits.
Solution Approach 2:
The patent ensures continuous useful action by optimizing the parallel window size to maintain high utilization of the PIM array throughout the computation process. This optimization prevents idle cycles and ensures that memory bandwidth is efficiently utilized for productive computations rather than being wasted on excessive data transfer.
Data Source
AI summary
There is a method of controlling a memory device. The method comprises acquiring a size of a PIM array provided to compute a convolution layer included in a deep neural network, a size of input data input to the convolution layer, and a size of a kernel filtering the input data; and determining a size of a parallel window such that a number of times of cycles of the PIM array for the convolution layer is minimized based on the size of the PIM array, the size of the input data, and the size of the kernel.


