Hardware Accelerator for Sparse Accumulation in SpGEMM

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The sparse accumulation operation in column-wise sparse general matrix-matrix multiplication (SpGEMM) algorithms is a performance bottleneck due to data-dependent branches, hash probing latency, and hash collisions, which are difficult to predict and optimize effectively using conventional software implementations.

Innovation Solution

The proposed ASA architecture extends the ISA to execute partial sum search and accumulation with a single instruction, uses a dedicated on-chip cache for pipelined sparse accumulation, and relies on parallel search capabilities to reduce latency, eliminating data-dependent branches and delaying overflow merging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a hash table is used for partial sum search in sparse accumulation, then the operation can be performed, but data-dependent branches cause difficulty for branch prediction and increase execution time

Engineering Contradiction:
Improvesparse accumulation speedVSAvoidhash probing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the software-based hash table with a hardware-based content-addressable memory (CAM) structure. This substitution eliminates the need for sequential hash probing and branch prediction, as the CAM performs parallel key matching in hardware, directly resolving the bottleneck caused by data-dependent branches in software hash tables.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces a dedicated hardware buffer and CAM structure as intermediary components between the input matrices and the accumulation process. This intermediary hardware structure pre-loads and caches partial sums, enabling fast parallel lookup and eliminating the time-consuming hash probing operation that plagues software implementations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If hash probing is used to find partial sums, then accumulation can proceed, but the dependency on hash probing results prevents hiding latency

Engineering Contradiction:
Improveaccumulation throughputVSAvoidhash probing latency
Core Design Contradiction:
ProductivityVSDuration of action of moving object

Solution Approach 1:

The patent performs preliminary action by pre-loading partial sums into the hardware buffer and CAM structure before the main accumulation loop begins. This pre-computation and pre-caching of partial sums allows the main processing pipeline to proceed without waiting for hash probing, effectively hiding the latency that would otherwise block throughput.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the accumulation process into independent parallel streams, each with its own hardware buffer and CAM lookup. This segmentation allows multiple partial sum searches to occur simultaneously in hardware, eliminating the sequential dependency on hash probing results and enabling latency to be hidden through parallel execution.

Inventive Principle:
Principle #1Segmentation

3Ease of manufacture

If a software hash table is used for sparse accumulation, then the implementation is simple, but hash collisions require time-consuming linear search

Engineering Contradiction:
Improveimplementation simplicityVSAvoidcollision resolution time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent replaces the software hash table with a hardware-based content-addressable memory (CAM) structure. This substitution eliminates the need for sequential hash probing and branch prediction, as the CAM performs parallel key matching in hardware, directly resolving the bottleneck caused by data-dependent branches in software hash tables.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces a dedicated hardware buffer and CAM structure as intermediary components between the input matrices and the accumulation process. This intermediary hardware structure pre-loads and caches partial sums, enabling fast parallel lookup and eliminating the time-consuming hash probing operation that plagues software implementations.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If conventional software implementation is used, then hardware overhead is avoided, but sparse accumulation becomes the performance bottleneck

Engineering Contradiction:
Improvehardware overheadVSAvoidSpGEMM performance
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent applies local quality by implementing a targeted hardware acceleration unit specifically for the sparse accumulation operation, rather than converting the entire SpGEMM system to hardware. This localized hardware buffer and CAM structure focuses resources only on the bottleneck operation, achieving performance improvement with minimal overall hardware overhead.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the SpGEMM computation into distinct functional units, with a dedicated hardware-accelerated sparse accumulation module separated from the rest of the software implementation. This segmentation allows the critical path operation to be accelerated in hardware while keeping the overall system implementation simple and maintaining ease of software deployment.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240411834A1Hardware accelerator for sparse accumulation in column-wise sparse general matrix-matrix multipliction algorithms
Publication Date: 2024.12.12 RGT UNIV OF CALIFORNIA
  • US20240411834A1 patent drawing
  • US20240411834A1 patent drawing
  • US20240411834A1 patent drawing

AI summary

A system and method of performing sparse accumulation in column-wise sparse general matrix-matrix multiplication (SpGEMM) algorithms. The method includes receiving a request to perform SpGEMM based on a first matrix and a second matrix. The method includes accumulating, in a hardware buffer, a hash key and an intermediate multiplication result of the first matrix and the second matrix. The method includes performing a probe search of a hardware cache to identify a match between the hash key and a partial sum associated with the first matrix and the second matrix. The method includes generating, by a hardware adder, a multiplication result based on the partial sum and the intermediate multiplication result from the accumulation waiting buffer.