Configurable Spatial Accelerator for Exascale Energy Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Exascale computing requires enormous system-level floating point performance within a tight power budget, and classical von Neumann architectures face challenges in simultaneously improving performance and energy efficiency, leading to high energy costs and inefficiencies in control overheads.

Innovation Solution

A spatial array of processing elements connected by lightweight, back-pressured communication networks, where each processing element operates only when input data is available and there is space for output, eliminating the need for triggered instructions and reducing control overheads, and utilizing network dataflow endpoint circuits to perform dataflow operations instead of traditional arithmetic-logic units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If classical von Neumann architectures are used to improve system-level floating point performance, then processing power increases, but energy consumption increases proportionally

Engineering Contradiction:
Improvesystem-level floating point performanceVSAvoidenergy consumption
Core Design Contradiction:
PowerVSUse of energy by moving object

Solution Approach 1:

The system is divided into multiple independent processing elements (PEs) organized in a spatial array, each capable of autonomous operation. This segmentation allows parallel execution of dataflow operations across multiple PEs, achieving high throughput while each PE consumes minimal energy independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Processing elements dynamically activate only when input data is available and output space exists, rather than continuous operation. This dynamic activation pattern reduces energy consumption by eliminating idle processing cycles while maintaining high utilization during active computation phases.

Inventive Principle:
Principle #15Dynamics

2Productivity

If traditional arithmetic-logic units are used for dataflow operations, then computational capability is achieved, but control overheads increase

Engineering Contradiction:
Improvecomputational capabilityVSAvoidcontrol overheads
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Each processing element autonomously determines when to execute operations based on local data availability signals from input buffers and space availability in output buffers. This self-service mechanism eliminates the need for centralized control units, instruction decoders, and complex scheduling logic, significantly reducing control overheads.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The control functionality is extracted from traditional arithmetic-logic units and distributed to individual processing elements. Each PE contains minimal control logic only sufficient for autonomous operation, removing the burden of complex control overheads from the computational units.

Inventive Principle:
Principle #2Taking out (Extraction)

3Use of energy by moving object

If spatial array of processing elements is implemented, then energy efficiency improves, but area consumption increases

Engineering Contradiction:
Improveenergy efficiencyVSAvoidarea consumption
Core Design Contradiction:
Use of energy by moving objectVSArea of stationary object

Solution Approach 1:

Multiple processing elements are merged into a compact spatial array architecture with shared communication infrastructure. The regular interconnect pattern and resource sharing among PEs reduce overall area consumption compared to distributed architectures, achieving high energy efficiency without proportional area increase.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

Processing elements are arranged in a two-dimensional spatial array rather than linear or hierarchical configurations. This dimensional organization enables efficient nearest-neighbor communication and reduces interconnect length, improving energy efficiency per unit area while maintaining compact footprint.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Device complexity

If back-pressured communication networks are used, then dataflow control is simplified, but latency may increase

Engineering Contradiction:
Improvedataflow controlVSAvoidlatency
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The back-pressured communication network implements feedback mechanisms where downstream buffers signal space availability upstream. This feedback enables automatic flow control without complex arbitration, simplifying dataflow control while maintaining predictable latency through regulated data movement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Output buffers pre-allocate space for incoming data before production occurs. This preliminary action prevents stalls during data transfer and enables continuous flow, reducing latency variability while maintaining simplified back-pressure control logic.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10515046B2Processors, methods, and systems with a configurable spatial accelerator
Publication Date: 2019.12.24 INTEL CORP
  • US10515046B2 patent drawing
  • US10515046B2 patent drawing
  • US10515046B2 patent drawing

AI summary

Systems, methods, and apparatuses relating to a configurable spatial accelerator are described. In one embodiment, a processor includes a synchronizer circuit coupled between an interconnect network of a first tile and an interconnect network of a second tile and comprising storage to store data to be sent between the interconnect network of the first tile and the interconnect network of the second tile, the synchronizer circuit to convert the data from the storage between a first voltage or a first frequency of the first tile and a second voltage or a second frequency of the second tile to generate converted data, and send the converted data between the interconnect network of the first tile and the interconnect network of the second tile