Neural Network Weight Compression via ReRAM Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Artificial neural networks require significant computational power and memory, leading to increased complexity and difficulty in miniaturization and commercialization, especially due to overfitting and the challenge of handling complex input data, which necessitates an efficient compression approach to maintain performance and reduce system costs.

Innovation Solution

A processor-implemented method that reorders filters, compresses weights by identifying zero-value weights, and generates operation unit maps to be mapped onto a resistive random access memory (ReRAM) crossbar array, allowing for sub-filter granularity compression and efficient mapping of non-zero weights, thereby reducing the number of weights and improving array compression ratios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the learning capacity of the neural network is increased to handle complex input data, then the neural network can process more complex data, but the connections within the neural network become more complex and require larger computational power and memory

Engineering Contradiction:
Improvelearning capacityVSAvoidconnections complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the filter weights into multiple groups and applies different compression ratios to different groups. This allows the neural network to maintain high learning capacity while reducing overall complexity by selectively compressing less important weight groups more aggressively

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different compression strategies to different parts of the weight matrix based on their importance. By identifying and preserving critical weight connections while compressing less important ones, the system maintains necessary complexity for handling complex data while reducing overall device complexity

Inventive Principle:
Principle #3Local quality

2Productivity

If the number of filters and weights is increased to improve neural network performance, then the processing capability is enhanced, but the system cost and difficulty of miniaturization increase

Engineering Contradiction:
Improveprocessing capabilityVSAvoidsystem cost
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges multiple weight values into single storage units by applying compression ratios. Multiple weight groups are combined and stored more efficiently, reducing the total number of storage units required while maintaining processing capability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the compression ratio parameter applied to different weight groups. By dynamically adjusting compression ratios based on weight importance, the system achieves better processing capability with reduced system cost through optimized parameter selection

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If traditional weight compression methods are used without sub-filter granularity, then the compression process is simpler, but the array compression ratio is lower and more weights remain uncompressed

Engineering Contradiction:
Improvecompression process complexityVSAvoidnumber of weights
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent segments weights into fine-grained groups at the sub-filter level, allowing for more precise compression control. This segmentation enables higher compression ratios by identifying and compressing more zero and near-zero weights that would be missed at coarser granularity levels

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240037377A1Method and apparatus with weight compression
Publication Date: 2024.02.01 SAMSUNG ELECTRONICS CO LTD
  • US20240037377A1 patent drawing
  • US20240037377A1 patent drawing
  • US20240037377A1 patent drawing

AI summary

A method and apparatus are provided. The method includes reordering a plurality of filters, then based on a result of the reordering, compressing weights, among a plurality of weights of the plurality of filters, resulting in some of the plurality of weights being uncompressed weights, generating a plurality of operation unit maps by mapping the uncompressed weights to respective operation units according to a predetermined bulk unit, and mapping the plurality of operation unit maps to an array.