ReRAM Crossbar Weight Sparsity Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep neural networks (DNNs) face challenges in efficiently exploiting weight sparsity due to the tightly coupled crossbar structure of ReRAM-based DNN accelerators, which limits energy efficiency and computation performance.

Innovation Solution

A computation method and apparatus that map weights to cells in a crossbar architecture, compress rows with zero weights, and use indexing to input non-zero weight inputs, allowing for efficient multiply-and-accumulate operations by skipping computations with zero weights and rearranging input order, thereby enhancing performance and energy efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If weights are mapped to all cells in the crossbar structure, then the computation can be performed using the tightly coupled crossbar architecture, but computations with zero weights are wasted and energy efficiency deteriorates

Engineering Contradiction:
Improvecomputation performanceVSAvoidenergy efficiency
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent extracts and removes rows corresponding to zero weights from the crossbar computation units. By identifying and eliminating these ineffectual computation rows before execution, the system avoids wasting energy on computations that would produce zero results, while maintaining the integrity of the crossbar architecture for the remaining non-zero weight computations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent dynamically reconfigures the crossbar computation units by adjusting which rows are active based on the sparsity pattern of the weight matrix. The system adapts the computation structure to match the actual data characteristics, enabling variable row configurations that optimize energy consumption for different computational workloads with varying sparsity levels.

Inventive Principle:
Principle #15Dynamics

2Use of energy by stationary object

If the crossbar structure is used for MAC operations, then data movement is reduced and energy savings are achieved, but weight sparsity cannot be efficiently exploited due to the tightly coupled structure

Engineering Contradiction:
Improveenergy savingsVSAvoidsparsity exploitation capability
Core Design Contradiction:
Use of energy by stationary objectVSAdaptability or versatility

Solution Approach 1:

The patent segments the crossbar computation units into independent row groups that can be selectively activated. By dividing the computation space into manageable segments corresponding to rows with non-zero weights, the system can independently control which segments are active, enabling efficient exploitation of weight sparsity while maintaining the energy-saving benefits of the crossbar architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different operational characteristics to different rows within the crossbar structure. Rows containing non-zero weights are configured for full computation with proper voltage activation, while rows with zero weights are configured to remain inactive or skip computation. This local differentiation allows the system to optimize energy consumption at the row level while preserving the overall crossbar architecture's energy-efficient data movement capabilities.

Inventive Principle:
Principle #3Local quality

3Reliability

If all rows are processed in the crossbar unit, then complete computation is performed, but computation time is increased due to processing zero-weight rows

Engineering Contradiction:
Improvecomputation completenessVSAvoidcomputation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary identification and marking of zero-weight rows before the actual MAC computation begins. By pre-processing the weight matrix to identify which rows contain only zero weights, the system can prepare the crossbar units in advance to skip these rows during computation, eliminating unnecessary processing time while ensuring that all necessary non-zero weight computations are completed accurately.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly improves computation performance and energy efficiency by eliminating ineffectual computations and leveraging weight sparsity, particularly in ReRAM-based DNN accelerators.

Implementation Method 1

the emerged electric currents I1, I2, I3 of each ReRAM cell induced by conductance G1, G2, G3 are accumulated to a total current I on the bitline BL instantaneously. The results of the MAC operations are retrieved simultaneously by sense amplifier (not shown) connected to the bitline BL, where the value of I equals to X1×W1+X2×W2+Bias.

Methodology Applied
Scientific EffectOhm's Law: Ohm's Law

Data Source

PatentUS11526328B2Computation method and apparatus exploiting weight sparsity
Publication Date: 2022.12.13 MACRONIX INTERNATIONAL CO LTD
  • US11526328B2 patent drawing
  • US11526328B2 patent drawing
  • US11526328B2 patent drawing

AI summary

A computation method and a computation apparatus exploiting weight sparsity, adapted for a processor to perform multiply-and-accumulate operations on a memory including multiple input and output lines crossing each other. In the method, weights are mapped to the cells of each operation unit (OU) in the memory. The rows of the cells of each OU are compressed by removing at least one row of the cells each mapped with a weight of 0, and an index including values each indicating a distance between every two rows of the cells including at least one cell mapped with a non-zero weight for each OU is encoded. Inputs are inputted to the input lines corresponding to the rows of each OU excluding the rows of the cells with the weight of 0 according to the index and outputs are sensed from the output lines corresponding to the OU to compute a computation result.