Configurable Spatial Accelerator Memory Interface Circuit Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Exascale computing requires high system-level floating point performance within a tight power budget, which classical von Neumann architectures struggle to achieve due to out-of-order scheduling, simultaneous multi-threading, and complex register files, leading to high energy costs.
Innovation Solution
A configurable spatial accelerator (CSA) with a spatial array of low-complexity, energy-efficient processing elements connected by lightweight communication networks, where dataflow graphs are directly executed, and network dataflow endpoint circuits perform operations instead of traditional processing elements, reducing control overheads and improving energy efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If classical von Neumann architectures are used to achieve high system-level floating point performance, then computational capability is improved, but energy consumption increases due to out-of-order scheduling, simultaneous multi-threading, and complex register files
Solution Approach 1:
The system is segmented into multiple processing elements (PEs) that can be independently configured and activated. Each PE is a simple, energy-efficient unit that can be selectively enabled based on workload requirements, avoiding the energy overhead of complex general-purpose processors. The segmentation allows the system to achieve high floating point performance through parallel simple units rather than fewer complex units.
Solution Approach 2:
The architecture employs dynamic configuration where processing elements can be reconfigured at runtime to perform different operations. This dynamic nature allows the system to adapt to varying computational requirements without the energy penalty of static complex hardware, achieving high performance through flexible deployment of simple processing units.
2Use of energy by moving object
If configurable spatial accelerator with simple processing elements is used, then energy efficiency is improved, but device complexity increases due to spatial array configuration and network dataflow endpoint circuits
Solution Approach 1:
The network dataflow endpoint circuits automatically perform operations based on incoming dataflow instructions without requiring complex external control logic. Each endpoint circuit is self-contained and can independently execute operations, reducing the overall control complexity of the system while maintaining energy efficiency.
Solution Approach 2:
The processing elements and network endpoint circuits are designed as universal, multi-functional units that can perform various operations through configuration rather than dedicated hardware. This universality reduces device complexity by avoiding the need for specialized circuits for each function, while still achieving high energy efficiency through optimized simple architectures.
3Device complexity
If network dataflow endpoint circuits perform operations instead of traditional processing elements, then control overhead is reduced, but the number of circuit elements increases
Solution Approach 1:
Control logic is extracted from centralized processing elements and distributed to the network dataflow endpoint circuits. Each endpoint circuit contains minimal local control logic to perform its specific operation, eliminating the need for complex centralized control and reducing overall control overhead despite the increased number of circuit elements.
Solution Approach 2:
The network dataflow endpoint circuits act as intermediaries between the simple processing elements and the memory system. These intermediary circuits handle the complexity of data movement and operation execution, allowing the processing elements to remain simple while distributing control functionality across multiple lightweight circuits.
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
Systems, methods, and apparatuses relating to memory interface circuit allocation in a configurable spatial accelerator are described. In one embodiment, a configurable spatial accelerator (CSA) includes a plurality of processing elements; a plurality of request address file (RAF) circuits, and a circuit switched interconnect network between the plurality of processing elements and the RAF circuits. As a dataflow architecture, embodiments of CSA have a unique memory architecture where memory accesses are decoupled into an explicit request and response phase allowing pipelining through memory. Certain embodiments herein provide for an improved memory sub-system design via the improvements to allocation discussed herein.