Spatial Array Processor Privileged Configuration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Exascale computing requires high system-level floating point performance within a tight power budget, which classical von Neumann architectures struggle to achieve due to high energy costs associated with out-of-order scheduling, simultaneous multi-threading, and complex register files, making it difficult to improve both performance and energy efficiency simultaneously.

Innovation Solution

A spatial array of processing elements connected by lightweight communication networks, where each processing element operates as a dataflow operator that consumes input data only when available, and a configuration controller enables pipelined configuration and extraction operations to reduce latency and enhance energy efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If classical von Neumann architectures are used to achieve high system-level floating point performance, then processing capability is improved, but energy consumption increases significantly

Engineering Contradiction:
Improvesystem-level floating point performanceVSAvoidenergy consumption
Core Design Contradiction:
PowerVSUse of energy by moving object

Solution Approach 1:

The spatial array architecture divides the processing system into multiple independent processing elements (PEs) that operate in parallel. Each PE handles specific dataflow operations independently, eliminating the need for complex centralized control structures like out-of-order scheduling and simultaneous multi-threading. This segmentation allows the system to achieve high floating point performance through parallel computation while reducing energy consumption by removing inefficient centralized control mechanisms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The spatial array employs dynamic configuration capabilities where processing elements can be reconfigured at runtime to optimize for different computational workloads. The configuration controller enables pipelined configuration and extraction operations, allowing the system to adapt its structure dynamically based on the specific dataflow graph being executed. This dynamic adaptability ensures optimal energy-efficiency-performance tradeoff for varying computational requirements.

Inventive Principle:
Principle #15Dynamics

2Productivity

If complex control mechanisms like out-of-order scheduling and simultaneous multi-threading are implemented, then processing performance is improved, but device complexity increases

Engineering Contradiction:
Improveprocessing performanceVSAvoidcontrol mechanism complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Processing elements in the spatial array operate autonomously based on dataflow semantics. Each PE consumes input data only when available and produces output when computation is complete, without requiring complex centralized scheduling mechanisms. This self-service approach eliminates out-of-order scheduling and simultaneous multi-threading complexity while maintaining high productivity through parallel autonomous execution of dataflow operations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces complex mechanical control systems (out-of-order scheduling, simultaneous multi-threading) with a dataflow-driven control model. Instead of using intricate control logic to manage instruction execution, the system uses simple data availability signals to trigger computations. This substitution dramatically reduces device complexity while preserving processing performance through the natural parallelism of dataflow execution.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If configuration operations are performed traditionally, then processing elements can be reconfigured, but configuration time is excessive

Engineering Contradiction:
Improvereconfiguration capabilityVSAvoidconfiguration time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The configuration controller performs preliminary configuration actions by preparing configuration data in advance and using pipelined configuration operations. Instead of configuring processing elements sequentially when reconfiguration is needed, the system pre-configures elements in a pipeline fashion, allowing configuration work to be distributed across multiple cycles. This preliminary action significantly reduces the effective configuration time while maintaining full reconfiguration capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The pipelined configuration mechanism ensures continuous useful action during reconfiguration. While some processing elements are being reconfigured, others continue to execute their current dataflow graphs. The configuration controller orchestrates overlapping configuration and execution phases, ensuring that the system maintains productive work throughout the reconfiguration process rather than halting completely. This continuity reduces the effective configuration time by an order of magnitude.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS10445098B2Processors and methods for privileged configuration in a spatial array
Publication Date: 2019.10.15 INTEL CORP
  • US10445098B2 patent drawing
  • US10445098B2 patent drawing
  • US10445098B2 patent drawing

AI summary

Methods and apparatuses relating to privileged configuration in spatial arrays are described. In one embodiment, a processor includes processing elements; an interconnect network between the processing elements; and a configuration controller coupled to a first subset and a second, different subset of the plurality of processing elements, the first subset having an output coupled to an input of the second, different subset, wherein the configuration controller is to configure the interconnect network between the first subset and the second, different subset of the plurality of processing elements to not allow communication on the interconnect network between the first subset and the second, different subset when a privilege bit is set to a first value and to allow communication on the interconnect network between the first subset and the second, different subset of the plurality of processing elements when the privilege bit is set to a second value.