Neural Network Weight Compression via ReRAM Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial neural networks require significant computational power and memory, leading to increased complexity and difficulty in miniaturization and commercialization, especially due to overfitting and the challenge of handling complex input data, which necessitates an efficient compression approach to maintain performance and reduce system costs.
Innovation Solution
A processor-implemented method that reorders filters, compresses weights by identifying zero-value weights, and generates operation unit maps to be mapped onto a resistive random access memory (ReRAM) crossbar array, allowing for sub-filter granularity compression and efficient mapping of non-zero weights, thereby reducing the number of weights and improving array compression ratios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the learning capacity of the neural network is increased to handle complex input data, then the neural network can process more complex data, but the connections within the neural network become more complex and require larger computational power and memory
Solution Approach 1:
The patent segments the filter weights into multiple groups and applies different compression ratios to different groups. This allows the neural network to maintain high learning capacity while reducing overall complexity by selectively compressing less important weight groups more aggressively
Solution Approach 2:
The patent applies different compression strategies to different parts of the weight matrix based on their importance. By identifying and preserving critical weight connections while compressing less important ones, the system maintains necessary complexity for handling complex data while reducing overall device complexity
2Productivity
If the number of filters and weights is increased to improve neural network performance, then the processing capability is enhanced, but the system cost and difficulty of miniaturization increase
Solution Approach 1:
The patent merges multiple weight values into single storage units by applying compression ratios. Multiple weight groups are combined and stored more efficiently, reducing the total number of storage units required while maintaining processing capability
Solution Approach 2:
The patent changes the compression ratio parameter applied to different weight groups. By dynamically adjusting compression ratios based on weight importance, the system achieves better processing capability with reduced system cost through optimized parameter selection
3Device complexity
If traditional weight compression methods are used without sub-filter granularity, then the compression process is simpler, but the array compression ratio is lower and more weights remain uncompressed
Solution Approach 1:
The patent segments weights into fine-grained groups at the sub-filter level, allowing for more precise compression control. This segmentation enables higher compression ratios by identifying and compressing more zero and near-zero weights that would be missed at coarser granularity levels
Data Source
AI summary
A method and apparatus are provided. The method includes reordering a plurality of filters, then based on a result of the reordering, compressing weights, among a plurality of weights of the plurality of filters, resulting in some of the plurality of weights being uncompressed weights, generating a plurality of operation unit maps by mapping the uncompressed weights to respective operation units according to a predetermined bulk unit, and mapping the plurality of operation unit maps to an array.


