Multi-Tile Memory Mapped I/O Routing via Root Tile Endpoint Controller

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of routing and performance scaling in computing systems with multiple accelerator devices, particularly due to each device being configured as a peripheral device with its own bus function identifier, increases routing complexities and reduces efficiency.

Innovation Solution

Implementing a root tile and remote tiles architecture where the root tile acts as a primary interface for incoming messages and directs them to appropriate tiles, allowing for a single PCIe device software view while enabling performance scaling by connecting multiple silicon dies, with endpoint controllers managing device decode, routing, and configuration across tiles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple accelerator devices are incorporated into a system, then performance capability is improved, but routing complexity increases

Engineering Contradiction:
Improveperformance capabilityVSAvoidrouting complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges multiple accelerator devices into a single PCIe device by sharing a common BAR space. Multiple tiles (accelerator devices) are combined such that they present a unified interface to the system, with their memory-mapped I/O spaces consolidated into a single address space managed by the root tile's endpoint controller. This eliminates the need for separate PCIe devices and reduces routing complexity while maintaining enhanced performance capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The root tile's endpoint controller serves multiple functions: it manages PCIe transactions for the entire multi-tile device, decodes BAR addresses for all tiles, routes transactions to appropriate tiles, and handles configuration space access. This universal controller consolidates what would otherwise require separate controllers for each accelerator device, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If each accelerator device is configured as a peripheral device with its own bus function identifier, then device independence is improved, but routing complexity increases

Engineering Contradiction:
Improvedevice independenceVSAvoidrouting complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the accelerator functionality into multiple independent tiles (e.g., compute tiles, memory tiles, I/O tiles) that can be configured and operated independently. Each tile has its own IP blocks and can be independently initialized, yet they share a common PCIe interface and BAR space. This segmentation maintains device independence while avoiding the routing complexity of multiple separate PCIe devices.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The root tile acts as an intermediary between the system bus and the remote tiles. It receives PCIe transactions, decodes the BAR addresses, determines the destination tile, and routes the transaction accordingly. This intermediary approach allows remote tiles to be addressed through a unified BAR space without requiring direct routing paths from the system bus to each individual tile, thereby reducing routing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If multiple accelerator devices are used, then functional capability is improved, but system complexity increases

Engineering Contradiction:
Improvefunctional capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple accelerator devices into a single integrated device that presents one BAR space to the system. The root tile aggregates the memory-mapped I/O spaces of all tiles into a unified address space, allowing the system to access any tile through a single PCIe device interface. This merging approach enhances functional capability while avoiding the system complexity of managing multiple separate accelerator devices.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The root tile's endpoint controller is designed as a universal interface that handles all types of transactions for all tiles. It can decode various BAR addresses, route to different tile types (compute, memory, I/O), and manage configuration space access for the entire multi-tile device. This multi-functional design enables enhanced functional capability without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11157431B2System, apparatus and method for multi-die distributed memory mapped input/output support
Publication Date: 2021.10.26 INTEL CORP
  • US11157431B2 patent drawing
  • US11157431B2 patent drawing
  • US11157431B2 patent drawing

AI summary

In one embodiment, a method includes: receiving, in a root tile of an accelerator device having a plurality of tiles, a message from a processor, the message comprising a register write request to a register of a first remote tile of the plurality of remote tiles; decoding, in an endpoint controller of the root tile, a system address of the message to identify a destination tile for the message, based at least in part on a base address register decode of the system address; and in response to identifying the first remote tile as the destination tile, updating a first portion of an address offset field of the system address to a predetermined value and directing the message to the first remote tile coupled to the root tile via a sideband interconnect. Other embodiments are described and claimed.