Heterogeneous In-Memory Computing for Graph Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing graph-computing accelerators face inefficiencies due to irregular workload patterns, leading to ineffective data storage, high power consumption, and reduced parallelism in both digital and analog in-memory computing architectures, which degrade computational performance.

Innovation Solution

A heterogeneous in-memory computing apparatus with a sliding-window-execution model that dynamically assigns loads to digital or analog signal processing units based on graph data characteristics and hardware occupancy, optimizing load balance and execution efficiency by integrating complementary digital and analog computation units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If graph data is stored in traditional memory architecture with separate computing and storage layers, then data access bandwidth is sufficient, but access latency increases and computing efficiency decreases due to frequent data movement between memory and processor

Engineering Contradiction:
Improvedata access speedVSAvoidaccess latency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent merges computing units directly into the memory structure by integrating digital signal processing units and analog signal processing units within the memory chip architecture. This allows computation to be performed where data is stored, eliminating the need for frequent data movement between separate memory and processor components, thereby reducing access latency and improving data access speed.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from traditional two-dimensional memory organization to a three-dimensional stacked memory architecture with through-silicon vias connecting multiple memory layers to computing units. This vertical integration adds a spatial dimension to the computing-memory interface, enabling parallel data access paths and reducing access latency through shortened signal paths.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Use of energy by moving object

If digital signal processing units are used for graph computing, then power consumption is reduced compared to traditional architectures, but parallelism is limited due to heat dissipation constraints in stacked structures

Engineering Contradiction:
Improvepower consumptionVSAvoidparallelism
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The patent employs a hybrid architecture that combines both digital signal processing units and analog signal processing units within the same memory structure. The analog units provide high parallelism for computationally intensive operations, while digital units handle control and less parallelizable tasks. This composite approach leverages the strengths of both paradigms to achieve high parallelism without excessive power consumption.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent implements dynamic workload distribution between digital and analog processing units based on the characteristics of the graph computing tasks. The system can adaptively assign operations to analog units when high parallelism is needed and to digital units when precision or control is prioritized, thereby optimizing the balance between power consumption and parallelism according to real-time computational demands.

Inventive Principle:
Principle #15Dynamics

3Productivity

If analog signal processing units are used for graph computing, then parallelism is significantly improved, but power consumption increases due to continuous data writing and analog-digital conversion overheads

Engineering Contradiction:
ImproveparallelismVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments the processing workload by dividing graph computing tasks into portions suitable for analog processing and portions requiring digital processing. By segmenting the computational workload and assigning it to appropriate processing units, the system achieves high parallelism for analog operations while minimizing the overhead of continuous data conversion and processing, thereby reducing overall power consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces digital signal processing units as intermediaries between the analog processing units and the external memory/system. These digital units buffer and pre-process data before it reaches the analog units, reducing the frequency of analog-digital conversions and minimizing the power consumption associated with continuous conversion operations while maintaining high parallelism in the analog computing core.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If ReRAM crossbar architecture is used for in-memory computing, then matrix vector multiplication can be executed in O(1) complexity, but data access irregularity and memory bandwidth utilization are degraded due to sparse workload patterns

Engineering Contradiction:
Improvecomputation complexityVSAvoidmemory bandwidth utilization
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent designs a universal processing architecture that can handle both dense and sparse graph computing workloads efficiently. The system includes both analog signal processing units for O(1) matrix operations and digital signal processing units for irregular access patterns. This multi-functional architecture adapts to different workload characteristics, maintaining high computation efficiency while optimizing memory bandwidth utilization regardless of data sparsity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic workload routing that adapts to the sparsity characteristics of graph data in real-time. When dense subgraphs are detected, the system routes operations to analog units for efficient O(1) computation. When sparse patterns are detected, the system routes to digital units that can handle irregular access more efficiently. This dynamic adaptation optimizes both computation complexity and memory bandwidth utilization according to actual data characteristics.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances computing efficiency, reduces power consumption, and improves bandwidth utilization, enabling higher performance and lower overheads for graph algorithms across varying data scales while ensuring load balance and minimizing synchronization and remote data access overheads.

Implementation Method 1

at least comprising a plurality of dynamic random access storage devices stacked with each other and vertically connected by means of through-silicon vias

Methodology Applied
Scientific EffectElectrical Conduction: Conduction (electrical)

Implementation Method 2

By writing graph edge data into the resistor of the ReRAM unit, when a set of vertex data is converted into analog voltage signals actin on word lines, the ReRAM crossbar architecture can immediately execute matrix vector multiplication operations, and generate analog current signals on its bit lines

Methodology Applied
Scientific EffectOhm's Law: Ohm's Law

Data Source

PatentUS11176046B2Graph-computing-oriented heterogeneous in-memory computing apparatus and operational method thereof
Publication Date: 2021.11.16 HUAZHONG UNIV OF SCI & TECH
  • US11176046B2 patent drawing
  • US11176046B2 patent drawing
  • US11176046B2 patent drawing

AI summary

The present invention relates to a graph-computing-oriented heterogeneous in-memory computing apparatus, comprising a memory control unit, a digital signal processing unit, and a plurality of analog signal processing units using the memory control unit.