Near-Memory Lattice Accelerator Mapping for Parallel HPC and AI

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing coarse-grained reconfigurable arrays (CGRAs) face challenges in efficiently accelerating high-performance computing (HPC) and artificial intelligence (AI) applications due to suboptimal compiler mapping methods for parallel operations and high precision requirements, leading to performance bottlenecks and increased physical size.

Innovation Solution

A reconfigurable dataflow accelerator (RDA) with a lattice structure of computing and memory modules, each equipped with functional units, is used to tile operations and assign tasks based on a cost model, enabling efficient parallel processing and reduction operations within the RDA, utilizing memory modules for both operations and data transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If compiler mapping methods are optimized for parallel operations and high precision requirements, then accelerator performance is improved, but device complexity increases

Engineering Contradiction:
Improveaccelerator performanceVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the accelerator into computing modules and memory modules arranged in a lattice structure. Each module is independently configurable, allowing the compiler to map parallel operations to multiple computing modules simultaneously while maintaining manageable individual module complexity. The lattice structure divides the overall system into repeating unit cells that can be scaled without proportionally increasing individual module complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic reconfigurability where the connection topology between modules can be changed at runtime via the compiler. This allows the same physical hardware to adapt its logical structure to match the parallel operation requirements of different algorithms, achieving high performance without permanently increasing device complexity. The reconfigurable interconnect enables runtime optimization of dataflow paths.

Inventive Principle:
Principle #15Dynamics

2Speed

If memory modules are equipped with integrated functional units for operations, then operation speed is enhanced, but area of memory modules increases

Engineering Contradiction:
Improveoperation speedVSAvoidarea of memory modules
Core Design Contradiction:
SpeedVSArea of moving object

Solution Approach 1:

The patent makes memory modules multi-functional by integrating functional units that can perform both memory operations and computational operations. The same memory module structure serves dual purposes: storing data and executing operations on that data. This eliminates the need for separate compute modules for certain operations, reducing overall system area while maintaining enhanced operation speed through in-memory computation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the functional units with memory modules, combining what would traditionally be separate components into a single integrated unit. This consolidation allows data to be processed directly where it is stored, eliminating data transfer overhead and improving operation speed. The merged structure reduces total chip area compared to having separate memory and compute components.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12524236B2Near-memory operator and method with accelerator performance improvement
Publication Date: 2026.01.13 SAMSUNG ELECTRONICS CO LTD
  • US12524236B2 patent drawing
  • US12524236B2 patent drawing
  • US12524236B2 patent drawing

AI summary

A system configured to perform an operation includes: a hardware device comprising a plurality of computing modules and a plurality of memory modules arranged in a lattice form, each of the computing modules comprising a coarse-grained reconfigurable array and each of the memory modules comprising a static random-access memory and a plurality of functional units connected to the static random-access memory; and a compiler configured to divide a target operation and assign the divided target operation to the computing modules and the memory modules such that the computing modules and the memory modules of the hardware device perform the target operation.