Bandwidth Allocation Across Accelerator Cluster Interconnect Switch

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In modern computer systems, especially those with multiple accelerators for AI applications, the bandwidth of hardware interconnect buses becomes a bottleneck, as all accelerators are treated equally, leading to competition for memory bandwidth and latency issues, with latency-sensitive applications like image recognition suffering due to shared resources.

Innovation Solution

Implementing a system with virtual communication channels and traffic class identifiers to prioritize memory access requests, where higher-priority applications are assigned more bandwidth and resources, using a dual-tiered virtual communication channel structure with separate lanes for high-priority and low-priority traffic, ensuring efficient memory communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple accelerators share the same hardware interconnect bus without priority differentiation, then device complexity is reduced and ease of operation is improved, but bandwidth utilization efficiency deteriorates and latency-sensitive applications suffer performance degradation

Engineering Contradiction:
Improvebandwidth utilization efficiencyVSAvoidinterconnect structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The hardware interconnect bus is segmented into multiple virtual communication channels, each capable of carrying different types of traffic with different priority levels. This segmentation allows the system to differentiate between latency-sensitive traffic and other traffic, improving bandwidth utilization efficiency without requiring physically separate buses for each application type.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of priority differentiation by adding traffic class identifiers and virtual channel dimensions to the existing bus architecture. Instead of creating multiple physical buses (which would increase complexity), the solution adds logical layers (virtual channels and priority levels) to the existing single bus, thereby improving productivity without proportionally increasing device complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If all accelerators are treated equally with equal bandwidth allocation, then device complexity is minimized and ease of operation is improved, but latency-sensitive applications experience performance degradation due to bandwidth competition

Engineering Contradiction:
ImprovelatencyVSAvoidtraffic management complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent applies local quality by assigning different priority levels and bandwidth allocation characteristics to different virtual communication channels based on their specific traffic requirements. Latency-sensitive traffic channels receive higher priority and guaranteed bandwidth allocation, while other channels operate with standard allocation. This localized differentiation improves latency performance for critical applications without requiring complex global reconfiguration of the entire interconnect system.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts bandwidth allocation parameters (such as bandwidth grants, priority levels, and channel assignments) based on traffic class identifiers and application requirements. By changing these parameters rather than the fundamental hardware architecture, the system achieves improved latency performance while keeping the underlying device complexity manageable through software-controlled parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If separate physical interconnect buses are provided for each accelerator, then bandwidth utilization efficiency and application performance are improved, but device complexity and cost increase significantly

Engineering Contradiction:
ImprovethroughputVSAvoidinterconnect architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The hardware interconnect bus is designed with multi-functionality to handle multiple types of traffic simultaneously through virtual communication channels. A single physical bus infrastructure serves multiple accelerators and multiple application types by dynamically allocating bandwidth and priority levels based on traffic class identifiers. This universal approach achieves throughput improvements comparable to multiple dedicated buses while avoiding the complexity and cost of physically separate interconnect structures for each accelerator.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10848440B2Systems and methods for allocating bandwidth across a cluster of accelerators
Publication Date: 2020.11.24 CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
  • US10848440B2 patent drawing
  • US10848440B2 patent drawing
  • US10848440B2 patent drawing

AI summary

The present disclosure provides methods and systems directed to providing quality of service to cluster of accelerators. The system can include a root connector; an interconnect switch communicatively coupled to the root connector over a plurality of lanes comprising a first set of lanes and a second set of lanes, wherein the first set of lanes are associated with a first virtual communication channel and a second set of lanes are associated with a second virtual communication channel; a first accelerator communicatively coupled to the interconnect switch and associated with a first traffic class identifier corresponding to first communication traffic communicated over the first set of lanes; and a plurality of accelerators communicatively coupled to the interconnect switch and associated with a second traffic class identifier that corresponds to second communication traffic having lower priority than the first communication traffic and that is communicated over the second set of lanes.