Near Data Processor for Multi-Core Memory Bandwidth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multi-core architectures struggle to realize bandwidth advantages of new memory technologies like high bandwidth memory (HBM) due to inefficiencies in data movement, particularly for applications with low spatial and temporal locality, leading to increased memory access latency.

Innovation Solution

The implementation of energy-efficient near data accelerators or processors with reduced compute and caching capabilities, which are configured to directly access high bandwidth memory via multiple processing engines and memory buffers, thereby overcoming cache hierarchy and data movement inefficiencies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If traditional cache hierarchy techniques are used to move data closer to processors, then data movement cost is reduced for applications with high spatial and temporal locality, but performance deteriorates for applications with low spatial and temporal locality due to increased memory access latency

Engineering Contradiction:
Improvedata movement costVSAvoidmemory access performance
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The system segments the processing workload by introducing dedicated near data processors (NDPs) that handle specific compute-intensive tasks directly at memory location, while traditional processors handle other workloads. This segmentation allows each processor type to be optimized for its specific function, with NDPs having reduced compute capabilities but direct memory access for tasks requiring low latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Near data processors act as intermediaries between memory and traditional processors. These NDPs are positioned close to memory resources and handle data access and initial processing, reducing the burden on traditional processors and enabling more efficient data movement patterns for applications with low spatial and temporal locality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If multiple memory controllers are introduced to support high bandwidth memory, then memory bandwidth is increased, but device complexity increases due to additional control logic and data movement management

Engineering Contradiction:
Improvememory bandwidthVSAvoidcontrol logic complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

Multiple memory controllers are merged with near data processors, combining the memory control functionality with the data processing functionality in a single integrated unit. This merging reduces the overall system complexity by eliminating separate control logic and simplifying data movement management while still supporting high bandwidth memory interfaces.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The near data processors are designed with multi-functionality, serving both as data processors and as memory controllers. This universal design allows a single component to handle multiple functions (data access, computation, and memory management), reducing the need for separate dedicated components and simplifying the overall system architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If near data processors with reduced compute capabilities are used, then data movement efficiency is improved for low locality applications, but processing capability is reduced compared to traditional processors

Engineering Contradiction:
Improvedata movement efficiencyVSAvoidprocessing capability
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

Different parts of the computing system are given different quality characteristics. Traditional processors maintain full compute capability for general-purpose workloads, while near data processors have reduced compute capabilities optimized specifically for data movement and memory-access intensive tasks. This local differentiation allows each processor type to be optimized for its intended function.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Near data processors provide partial processing capability sufficient for their specific function (data movement and initial computation) but not full processing capability. This partial action approach is acceptable because NDPs handle only specific workloads where their reduced capabilities are adequate, while traditional processors handle workloads requiring full capability.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12204478B2Techniques for near data acceleration for a multi-core architecture
Publication Date: 2025.01.21 INTEL CORP
  • US12204478B2 patent drawing
  • US12204478B2 patent drawing
  • US12204478B2 patent drawing

AI summary

Examples include techniques for near data acceleration for a multi-core architecture. A near data processor included in a memory controller of a processor may access data maintained in a memory device coupled with the near data processor via one or more memory channels responsive to a work request to execute a kernel, an application or a loop routine using the accessed data to generate values. The near data processor provides an indication to the requestor of the work request that values have been generated.