QoS Virtual Fabric Switching for HPC Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-performance computing (HPC) systems face performance bottlenecks due to data transfer latencies across interconnects between compute nodes, which are exacerbated by the use of different protocols and physical layers at various hierarchy levels, requiring effective Quality of Service (QoS) management and flexible interconnect architectures to support advanced protocols and topologies.

Innovation Solution

The proposed architecture defines a message passing, switched server interconnection network spanning OSI Network Model Layers 1 and 2, leveraging IETF Internet Protocol for Layer 3 and combining new and leveraged specifications for Layer 4, with Host Fabric Interfaces, full-duplex point-to-point links, switches, and a comprehensive management model to implement QoS and support multiple virtual lanes, ensuring efficient data transfer and adaptability across heterogeneous hardware capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If different protocols and physical layers are used at various interconnect hierarchy levels, then protocol flexibility and adaptability are improved, but data transfer latency increases due to protocol conversion and bridging requirements

Engineering Contradiction:
Improveprotocol flexibilityVSAvoiddata transfer latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements a universal credit-based flow control mechanism that operates across multiple interconnect hierarchy levels (Intra-SoC, Inter-SoC, Inter-Rack, Inter-Fabric) regardless of the specific protocol or physical layer being used. This single flow control approach can be applied to different protocols (InfiniBand, Ethernet, proprietary) and physical layers (optical, electrical) without requiring protocol-specific implementations, thereby reducing latency while maintaining adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the flow control mechanism into virtual channels (VCs) that can be independently configured and managed at each interconnect level. Each VC can be assigned specific QoS parameters, bandwidth guarantees, and latency constraints, allowing the system to handle different protocol requirements through segmented virtual channels rather than requiring complete protocol conversion, thus reducing overall latency.

Inventive Principle:
Principle #1Segmentation

2Reliability

If credit-based flow control is implemented across multiple virtual lanes, then QoS management and data transfer reliability are improved, but interconnect complexity increases

Engineering Contradiction:
Improvedata transfer reliabilityVSAvoidinterconnect complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal credit-based flow control mechanism that operates across multiple interconnect hierarchy levels (Intra-SoC, Inter-SoC, Inter-Rack, Inter-Fabric) regardless of the specific protocol or physical layer being used. This single flow control approach can be applied to different protocols (InfiniBand, Ethernet, proprietary) and physical layers (optical, electrical) without requiring protocol-specific implementations, thereby reducing latency while maintaining adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the flow control mechanism into virtual channels (VCs) that can be independently configured and managed at each interconnect level. Each VC can be assigned specific QoS parameters, bandwidth guarantees, and latency constraints, allowing the system to handle different protocol requirements through segmented virtual channels rather than requiring complete protocol conversion, thus reducing overall latency.

Inventive Principle:
Principle #1Segmentation

3Productivity

If packet preemption and interleaving mechanisms are added to enhance QoS, then data transfer efficiency is improved, but system complexity and processing overhead increase

Engineering Contradiction:
Improvedata transfer efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements dynamic packet preemption and interleaving mechanisms that adapt to real-time QoS requirements and network conditions. High-priority packets can preempt lower-priority packets in flight, and interleaving dynamically adjusts based on bandwidth allocation and latency constraints. This dynamic behavior improves data transfer efficiency by ensuring critical data receives necessary bandwidth while maintaining manageable system complexity through adaptive rather than static configurations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3087710B1Method, apparatus and system for QOS within high performance fabrics
Publication Date: 2019.11.20 INTEL CORP
  • EP3087710B1 patent drawingFigure 1
  • EP3087710B1 patent drawingFigure 2~13
  • EP3087710B1 patent drawingFigure 3~4

AI summary

Method, apparatus, and systems for implementing Quality of Service (QoS) within high performance fabrics. A multi-level QoS scheme is implemented including virtual fabrics, Traffic Classes, Service Levels (SLs), Service Channels (SCs) and Virtual Lanes (VLs). SLs are implemented for Layer 4 (Transport Layer) end-to-end transfer of fabric packets, while SCs are used to differentiate fabric packets at the Link Layer. Fabric packets are divided into flits, with fabric packet data transmitted via fabric links as flits streams. Fabric switch input ports and device receive ports detect SC IDs for received fabric packets and implement SC-to-VL mappings to determine VL buffers to buffer fabric packet flits in. An SL may have multiple SCs, and SC-to-SC mapping may be implemented to change the SC for a fabric packet as it is forwarded through the fabric, while maintaining its SL. A Traffic Class may include multiple SLs, enabling request and response traffic for an application to employ separate SLs.