Near-Memory Compute Operators for Disaggregated Memory Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data processing approaches that move data to a CPU for computation have become a significant performance bottleneck due to limited data reuse and slow interconnect performance, especially in scale-out data-intensive applications.
Innovation Solution
Near-memory computing (NMC) is adopted, where compute units are placed close to memory, utilizing 3D integration technologies to reduce memory access latency, power consumption, and increase bandwidth by performing data operations close to memory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is moved to CPU for computation, then computation can be performed, but data movement overhead increases and performance decreases
Solution Approach 1:
The patent transitions from a traditional von Neumann architecture where computation and memory are separated to a near-memory computing architecture where compute units are placed in close proximity to memory arrays. This spatial reorganization eliminates the need for data to traverse long interconnect distances to reach CPU cores, thereby reducing data movement time while maintaining computation capability.
Solution Approach 2:
The patent introduces near-memory compute units as intermediary processing elements between the memory arrays and remote CPU cores. These compute units can perform computations directly on data residing in memory, acting as intermediaries that reduce the burden of data movement to remote processors while still enabling computation.
2Productivity
If data is moved to CPU for computation, then computation can be performed, but interconnect bandwidth is consumed and energy increases
Solution Approach 1:
By reorganizing the computational architecture to place compute units near memory arrays rather than relying on distant CPU cores, the patent eliminates energy-intensive data traversals across the interconnect fabric. This spatial reorganization enables computation to occur where data resides, dramatically reducing the energy required for data movement.
3Productivity
If compute units are placed close to memory, then data movement is reduced, but device complexity increases
Solution Approach 1:
The patent segments the computational workload by distributing compute units across multiple memory arrays rather than concentrating all computation in centralized CPU cores. This segmentation allows each memory-compute complex to operate semi-independently, managing complexity through modular distribution while achieving high data processing efficiency.
4Reliability
If conventional memory hierarchy is used, then data access is supported, but bandwidth demands of multicore processors cannot be met
Solution Approach 1:
The patent fundamentally changes the data access model by placing compute units directly adjacent to memory arrays, enabling data to be accessed and processed in the same local domain. This eliminates the hierarchical progression through multiple cache levels and interconnect traversals, providing both reliable data access support and high-speed access that meets multicore processor bandwidth demands.
Data Source
AI summary
The disclosure provides for systems and methods for improving bandwidth and latency associated with executing data requests in disaggregated memory by leveraging usage indicators (also referred to as usage value), such as “freshness” of data operators and processing “gravity” of near memory compute functions. Examples of the systems and methods disclosed herein generate data operators comprising near memory compute functions offloaded proximate to disaggregated memory nodes, assign a usage value to each data operator based on at least one of: (i) a freshness indicator for each data operators, and (ii) a gravity indicator for each near memory compute function; and allocate data operations to the data operators based on the usage value.


