DNN Weight Assignment to 3D Crossbar Array Tiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep Neural Network (DNN) models face contention issues due to limited memory capacity in weight-stationary architectures, leading to increased latency and throughput limitations when handling large models or batches, especially in multi-tier memory systems where poor layer assignment to 3D crossbar array tiles results in inefficient resource utilization.
Innovation Solution
The system employs an efficient allocation strategy for assigning DNN weight matrices to 2D tiers of 3D crossbar array tiles, optimizing the assignment of neural network model layers across multiple tiers to minimize contention, latency, and maximize throughput by ensuring that each tile processes a batch until all assigned layers are completed, thereby reducing the number of tiles required and improving resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple weight matrices are assigned to the same 3D memory tile to reduce the number of tiles required, then resource utilization improves, but contention increases leading to longer execution time and latency
Solution Approach 1:
The patent segments the neural network layers into different groups that can be assigned to different tiers within the same 3D memory tile. By dividing layers into segments that can be processed in different tiers, the system reduces contention while maintaining high resource utilization. Each tier can handle specific layer segments simultaneously, avoiding the bottleneck of shared resource contention.
Solution Approach 2:
The patent utilizes the third dimension (Z-axis) of the 3D crossbar array by implementing multiple tiers within each tile. Instead of assigning multiple weight matrices to the same 2D plane, the system distributes them across different tiers in the vertical dimension. This dimensional expansion allows parallel processing of different layer segments without contention, as each tier operates independently.
2Productivity
If more 3D memory tiles are allocated to handle large models or batches, then throughput increases, but the system requires more memory capacity than available
Solution Approach 1:
The patent makes each 3D memory tile universal by enabling it to handle multiple different weight matrices across its tiers. Instead of dedicating entire tiles to specific layers, each tile can service multiple layers by switching between tiers. This multi-functionality increases the effective memory capacity available for large models without requiring additional physical tiles.
Solution Approach 2:
The patent implements a tier switching mechanism where tiers can be dynamically allocated to different layers based on current processing needs. When a tier completes processing for one layer, it can be recovered and reassigned to another layer, maximizing the utilization of existing memory capacity. This dynamic allocation allows the system to handle large models and batches by efficiently reusing the same memory resources.
3Ease of operation
If layers are assigned to tiers without optimization, then implementation is simpler, but contention leads to longer dead-time latency before next input
Solution Approach 1:
The patent applies preliminary action by pre-optimizing the assignment of layers to tiers based on the specific neural network model and batch size. The system performs an initial analysis to determine the optimal tier assignment that minimizes contention and dead-time latency. This pre-computed assignment strategy is then applied during execution, achieving low latency without requiring complex real-time decision-making.
Solution Approach 2:
The patent implements dynamic layer-to-tier assignment that adapts to different neural network models and batch sizes. Rather than using a fixed static assignment, the system can reconfigure which layers are assigned to which tiers based on the current workload characteristics. This dynamic approach optimizes performance for different scenarios while maintaining ease of operation through automated configuration.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces latency and dead-time, enhances throughput, and optimizes resource usage by ensuring that each tile processes a batch until all assigned layers are completed, effectively addressing the memory constraints and contention issues in weight-stationary architectures.
Implementation Method 1
When these devices are arranged in a crossbar configuration, it allows to perform a matrix-vector multiplication in a single time step, exploiting the advantages of storage capability and Kirchhoff's circuits laws.
Data Source
AI summary
A system, method and computer program product for assigning deep neural network (DNN) weight matrices to a Compute-in-Memory (CiM) accelerator system, and particularly, efficient allocation strategies for assigning DNN model weight-layers to two-dimensional (2D) tiers of three-dimensional (3D) crossbar array tiles. Such efficient allocation strategies for assigning DNN model weight-layers to tiers and tiles of a CiM accelerator are optimized to minimize contention, latency and dead-time, and to maximize accelerator throughput. In one scenario, efficient allocation strategies include assigning DNN weight matrices to the 2D tiers of a 3D crossbar array tile to maximize throughput and minimize completion latency for a finite-batch-size example of an incoming workflow. In a further scenario, efficient allocation strategies assign DNN weight matrices to the 2D tiers of a 3D crossbar array tile to minimize dead-time-latency-before-next-batch-member-can-be-input in an infinite-batch-size or a continuous workflow scenario.


