Near-Memory Lattice Accelerator Mapping for Parallel HPC and AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing coarse-grained reconfigurable arrays (CGRAs) face challenges in efficiently accelerating high-performance computing (HPC) and artificial intelligence (AI) applications due to suboptimal compiler mapping methods for parallel operations and high precision requirements, leading to performance bottlenecks and increased physical size.
Innovation Solution
A reconfigurable dataflow accelerator (RDA) with a lattice structure of computing and memory modules, each equipped with functional units, is used to tile operations and assign tasks based on a cost model, enabling efficient parallel processing and reduction operations within the RDA, utilizing memory modules for both operations and data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If compiler mapping methods are optimized for parallel operations and high precision requirements, then accelerator performance is improved, but device complexity increases
Solution Approach 1:
The patent segments the accelerator into computing modules and memory modules arranged in a lattice structure. Each module is independently configurable, allowing the compiler to map parallel operations to multiple computing modules simultaneously while maintaining manageable individual module complexity. The lattice structure divides the overall system into repeating unit cells that can be scaled without proportionally increasing individual module complexity.
Solution Approach 2:
The patent implements dynamic reconfigurability where the connection topology between modules can be changed at runtime via the compiler. This allows the same physical hardware to adapt its logical structure to match the parallel operation requirements of different algorithms, achieving high performance without permanently increasing device complexity. The reconfigurable interconnect enables runtime optimization of dataflow paths.
2Speed
If memory modules are equipped with integrated functional units for operations, then operation speed is enhanced, but area of memory modules increases
Solution Approach 1:
The patent makes memory modules multi-functional by integrating functional units that can perform both memory operations and computational operations. The same memory module structure serves dual purposes: storing data and executing operations on that data. This eliminates the need for separate compute modules for certain operations, reducing overall system area while maintaining enhanced operation speed through in-memory computation.
Solution Approach 2:
The patent merges the functional units with memory modules, combining what would traditionally be separate components into a single integrated unit. This consolidation allows data to be processed directly where it is stored, eliminating data transfer overhead and improving operation speed. The merged structure reduces total chip area compared to having separate memory and compute components.
Data Source
AI summary
A system configured to perform an operation includes: a hardware device comprising a plurality of computing modules and a plurality of memory modules arranged in a lattice form, each of the computing modules comprising a coarse-grained reconfigurable array and each of the memory modules comprising a static random-access memory and a plurality of functional units connected to the static random-access memory; and a compiler configured to divide a target operation and assign the divided target operation to the computing modules and the memory modules such that the computing modules and the memory modules of the hardware device perform the target operation.


