Configurable Spatial Accelerator for Exascale Energy Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Exascale computing requires enormous system-level floating point performance within a tight power budget, and classical von Neumann architectures face challenges in simultaneously improving performance and energy efficiency, leading to high energy costs and inefficiencies in control overheads.
Innovation Solution
A spatial array of processing elements connected by lightweight, back-pressured communication networks, where each processing element operates only when input data is available and there is space for output, eliminating the need for triggered instructions and reducing control overheads, and utilizing network dataflow endpoint circuits to perform dataflow operations instead of traditional arithmetic-logic units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If classical von Neumann architectures are used to improve system-level floating point performance, then processing power increases, but energy consumption increases proportionally
Solution Approach 1:
The system is divided into multiple independent processing elements (PEs) organized in a spatial array, each capable of autonomous operation. This segmentation allows parallel execution of dataflow operations across multiple PEs, achieving high throughput while each PE consumes minimal energy independently.
Solution Approach 2:
Processing elements dynamically activate only when input data is available and output space exists, rather than continuous operation. This dynamic activation pattern reduces energy consumption by eliminating idle processing cycles while maintaining high utilization during active computation phases.
2Productivity
If traditional arithmetic-logic units are used for dataflow operations, then computational capability is achieved, but control overheads increase
Solution Approach 1:
Each processing element autonomously determines when to execute operations based on local data availability signals from input buffers and space availability in output buffers. This self-service mechanism eliminates the need for centralized control units, instruction decoders, and complex scheduling logic, significantly reducing control overheads.
Solution Approach 2:
The control functionality is extracted from traditional arithmetic-logic units and distributed to individual processing elements. Each PE contains minimal control logic only sufficient for autonomous operation, removing the burden of complex control overheads from the computational units.
3Use of energy by moving object
If spatial array of processing elements is implemented, then energy efficiency improves, but area consumption increases
Solution Approach 1:
Multiple processing elements are merged into a compact spatial array architecture with shared communication infrastructure. The regular interconnect pattern and resource sharing among PEs reduce overall area consumption compared to distributed architectures, achieving high energy efficiency without proportional area increase.
Solution Approach 2:
Processing elements are arranged in a two-dimensional spatial array rather than linear or hierarchical configurations. This dimensional organization enables efficient nearest-neighbor communication and reduces interconnect length, improving energy efficiency per unit area while maintaining compact footprint.
4Device complexity
If back-pressured communication networks are used, then dataflow control is simplified, but latency may increase
Solution Approach 1:
The back-pressured communication network implements feedback mechanisms where downstream buffers signal space availability upstream. This feedback enables automatic flow control without complex arbitration, simplifying dataflow control while maintaining predictable latency through regulated data movement.
Solution Approach 2:
Output buffers pre-allocate space for incoming data before production occurs. This preliminary action prevents stalls during data transfer and enables continuous flow, reducing latency variability while maintaining simplified back-pressure control logic.
Data Source
AI summary
Systems, methods, and apparatuses relating to a configurable spatial accelerator are described. In one embodiment, a processor includes a synchronizer circuit coupled between an interconnect network of a first tile and an interconnect network of a second tile and comprising storage to store data to be sent between the interconnect network of the first tile and the interconnect network of the second tile, the synchronizer circuit to convert the data from the storage between a first voltage or a first frequency of the first tile and a second voltage or a second frequency of the second tile to generate converted data, and send the converted data between the interconnect network of the first tile and the interconnect network of the second tile


