Near Memory Compute Circuitry for Memory-Bound Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for accelerating memory-bound AI workloads, such as those involving AI workloads like CNNs, GNNs, and transformer workloads, face challenges related to scalability, limited spatial-temporal locality, and reliability issues due to thermal and droop effects. Additionally, they often require custom memory solutions and lack adequate error correction control.

Innovation Solution

The implementation of programmable compute logic distributed across one or more I/O switches, such as CXL switches, coupled with CXL-attached memories. This setup allows for better performance in a scale-up model by leveraging higher off-chip memory bandwidth without sacrificing memory capacity, and it includes standard memory controllers for managing error correction and reliability tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If on-die or on-chip caching is used for processing units, then processing speed may be improved, but memory capacity is limited and spatial-temporal locality is insufficient for effective use

Engineering Contradiction:
Improveprocessing speedVSAvoidmemory capacity
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent transitions from on-chip caching (within the processor die) to near-memory computing (at the memory interface/CXL switch level), effectively moving the compute functionality to a different spatial dimension. This allows access to much larger memory capacity while maintaining processing speed through proximity to the memory interface.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If custom memory solutions with integrated compute engines are used, then processing performance may be improved, but reliability deteriorates due to thermal and droop effects

Engineering Contradiction:
Improveprocessing performanceVSAvoidreliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an intermediary layer (the CXL switch with near-memory compute circuitry) between the host processor and the memory devices. This intermediary handles compute operations close to the memory interface without requiring custom integrated memory solutions, thereby maintaining reliability while improving processing performance for memory-bound workloads.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If memory capacity is increased to handle large data footprints, then scalability is improved, but memory bandwidth becomes the limiting factor

Engineering Contradiction:
Improvememory capacityVSAvoidmemory bandwidth
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent implements preliminary compute operations (such as aggregation, filtering, or preprocessing) in the near-memory compute circuitry before data is transferred to the host processor. This preliminary action reduces the volume of data that needs to be transferred, effectively increasing the usable memory bandwidth without changing the physical memory interface speed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250156356A1Techniques to utilize near memory compute circuitry for memory-bound workloads
Publication Date: 2025.05.15 INTEL CORP
  • US20250156356A1 patent drawing
  • US20250156356A1 patent drawing
  • US20250156356A1 patent drawing

AI summary

Examples include techniques to utilize near memory compute circuitry for memory-bound workloads. Examples include the near memory compute circuitry being resident on an input/output (I/O) arranged to couple with a plurality of memory devices configured as a memory pool that is accessible to a host central processing unit (CPU) through the I/O switch. The near memory compute circuitry may receive a request to obtain data from the memory pool and generate a result that is made available to the host CPU to facilitate acceleration of a memory-bound workload.