Heterogeneous In-Memory Computing for Graph Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing graph-computing accelerators face inefficiencies due to irregular workload patterns, leading to ineffective data storage, high power consumption, and reduced parallelism in both digital and analog in-memory computing architectures, which degrade computational performance.
Innovation Solution
A heterogeneous in-memory computing apparatus with a sliding-window-execution model that dynamically assigns loads to digital or analog signal processing units based on graph data characteristics and hardware occupancy, optimizing load balance and execution efficiency by integrating complementary digital and analog computation units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If graph data is stored in traditional memory architecture with separate computing and storage layers, then data access bandwidth is sufficient, but access latency increases and computing efficiency decreases due to frequent data movement between memory and processor
Solution Approach 1:
The patent merges computing units directly into the memory structure by integrating digital signal processing units and analog signal processing units within the memory chip architecture. This allows computation to be performed where data is stored, eliminating the need for frequent data movement between separate memory and processor components, thereby reducing access latency and improving data access speed.
Solution Approach 2:
The patent transitions from traditional two-dimensional memory organization to a three-dimensional stacked memory architecture with through-silicon vias connecting multiple memory layers to computing units. This vertical integration adds a spatial dimension to the computing-memory interface, enabling parallel data access paths and reducing access latency through shortened signal paths.
2Use of energy by moving object
If digital signal processing units are used for graph computing, then power consumption is reduced compared to traditional architectures, but parallelism is limited due to heat dissipation constraints in stacked structures
Solution Approach 1:
The patent employs a hybrid architecture that combines both digital signal processing units and analog signal processing units within the same memory structure. The analog units provide high parallelism for computationally intensive operations, while digital units handle control and less parallelizable tasks. This composite approach leverages the strengths of both paradigms to achieve high parallelism without excessive power consumption.
Solution Approach 2:
The patent implements dynamic workload distribution between digital and analog processing units based on the characteristics of the graph computing tasks. The system can adaptively assign operations to analog units when high parallelism is needed and to digital units when precision or control is prioritized, thereby optimizing the balance between power consumption and parallelism according to real-time computational demands.
3Productivity
If analog signal processing units are used for graph computing, then parallelism is significantly improved, but power consumption increases due to continuous data writing and analog-digital conversion overheads
Solution Approach 1:
The patent segments the processing workload by dividing graph computing tasks into portions suitable for analog processing and portions requiring digital processing. By segmenting the computational workload and assigning it to appropriate processing units, the system achieves high parallelism for analog operations while minimizing the overhead of continuous data conversion and processing, thereby reducing overall power consumption.
Solution Approach 2:
The patent introduces digital signal processing units as intermediaries between the analog processing units and the external memory/system. These digital units buffer and pre-process data before it reaches the analog units, reducing the frequency of analog-digital conversions and minimizing the power consumption associated with continuous conversion operations while maintaining high parallelism in the analog computing core.
4Productivity
If ReRAM crossbar architecture is used for in-memory computing, then matrix vector multiplication can be executed in O(1) complexity, but data access irregularity and memory bandwidth utilization are degraded due to sparse workload patterns
Solution Approach 1:
The patent designs a universal processing architecture that can handle both dense and sparse graph computing workloads efficiently. The system includes both analog signal processing units for O(1) matrix operations and digital signal processing units for irregular access patterns. This multi-functional architecture adapts to different workload characteristics, maintaining high computation efficiency while optimizing memory bandwidth utilization regardless of data sparsity.
Solution Approach 2:
The patent implements dynamic workload routing that adapts to the sparsity characteristics of graph data in real-time. When dense subgraphs are detected, the system routes operations to analog units for efficient O(1) computation. When sparse patterns are detected, the system routes to digital units that can handle irregular access more efficiently. This dynamic adaptation optimizes both computation complexity and memory bandwidth utilization according to actual data characteristics.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances computing efficiency, reduces power consumption, and improves bandwidth utilization, enabling higher performance and lower overheads for graph algorithms across varying data scales while ensuring load balance and minimizing synchronization and remote data access overheads.
Implementation Method 1
at least comprising a plurality of dynamic random access storage devices stacked with each other and vertically connected by means of through-silicon vias
Implementation Method 2
By writing graph edge data into the resistor of the ReRAM unit, when a set of vertex data is converted into analog voltage signals actin on word lines, the ReRAM crossbar architecture can immediately execute matrix vector multiplication operations, and generate analog current signals on its bit lines
Data Source
AI summary
The present invention relates to a graph-computing-oriented heterogeneous in-memory computing apparatus, comprising a memory control unit, a digital signal processing unit, and a plurality of analog signal processing units using the memory control unit.


