Sparsity-Aware Neural Processing Unit for Constant Probability Index Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Sparse convolutional neural networks face inefficiencies in matrix multiplication due to varying densities of input activation and weight matrices, leading to inconsistent performance and bottlenecks in multiplier utilization across different layers.
Innovation Solution
A sparsity-aware neural processing unit is designed to maintain constant probability matching between input activations and weights by using an index matching unit with a comparator array, weight buffer memory, and IA buffer memory, ensuring efficient alignment and delivery of non-zero values, regardless of matrix density, through a p-way priority encoder and FIFO system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If sparse matrix multiplication is performed by pruning zero values, then the number of memory communications and multiplications is reduced, but the multiplier utilization rate varies significantly with matrix density
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing non-zero element positions and values in compressed sparse row (CSR) format before multiplication. The index matching unit pre-identifies matching positions between input activation and weight matrices, preparing the data structure in advance to enable efficient multiplication while maintaining consistent resource utilization regardless of density variations
Solution Approach 2:
The patent changes the parameter representation by transforming dense matrices into sparse matrix format with explicit non-zero element storage. By representing matrices in terms of non-zero elements only (using row indices, column indices, and values), the system adapts the data structure parameters to match the computational requirements, enabling efficient processing across varying density conditions
2Adaptability or versatility
If the IA and weight matrices have different densities in different layers, then sparsity optimization is needed, but maintaining constant performance across layers becomes difficult
Solution Approach 1:
The patent applies universality by designing a unified sparse matrix multiplication architecture that handles varying densities across different layers through the same mechanism. The index matching unit and compressed sparse row format work consistently regardless of the specific density values, providing a universal solution that maintains reliable performance across diverse neural network layers with different sparsity patterns
3Loss of energy
If zero values are removed by approximation, then memory communications are reduced, but the index matching complexity increases
Solution Approach 1:
The patent applies the taking out principle by extracting and separately storing the positions and values of non-zero elements from the dense matrices. The compressed sparse row format extracts only the essential information (row indices, column indices, and values) needed for computation, removing unnecessary zero value representations and reducing memory communication overhead while enabling efficient index matching
Data Source
AI summary
A method of processing of a sparsity-aware neural processing unit includes receiving a plurality of input activations (IA); obtaining a weight having a non-zero value in each weight output channel; storing the weight and the IA in a memory, and obtaining an input channel index comprising a memory address location in which the weight and the IA are stored; and arranging the non-zero weight of each weight output channel according to a row size of an index matching unit (IMU) and matching the IA to the weight in the IMU comprising a buffer memory storing the input channel index.


