Collective Switch Architecture for HPC Network Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High Performance Computing (HPC) systems require network hardware that can support multiple collective operations simultaneously with low latency for short collectives and high bandwidth for long collectives, while ensuring reproducible results for floating point reductions, which existing technologies fail to achieve effectively.

Innovation Solution

A collective switch hardware architecture that includes input and output arrangement circuits, collective reduction logic with arithmetic logic units (ALUs) and arbitration and control circuitry, capable of supporting multiple simultaneous collective operations from different classes, allowing arbitrary input and output port mapping, and enabling parallel execution of collective operations within the network switch/router.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If one collective reduction operation is supported at a time per node, then reproducible floating point results are achieved, but the network hardware cannot support multiple collective operations simultaneously

Engineering Contradiction:
Improvenumber of simultaneous collective operationsVSAvoidreproducibility of floating point results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the collective operation processing by implementing separate hardware paths: a first path for short collectives using embedded network logic, and a second path for long collectives using network hub logic. This segmentation allows simultaneous execution of multiple collective operations while maintaining reproducible results through dedicated hardware enforcement of operation ordering in each path.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If hardware supports multiple short collectives, then operational versatility improves, but reproducibility for floating point operations cannot be guaranteed

Engineering Contradiction:
Improvenumber of simultaneous short collective operationsVSAvoidreproducibility of floating point results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements preliminary action by pre-configuring the network hub logic with arbitration mechanisms that establish a fixed order of operations before execution. The hub receives collective operations, arbitrates them in advance to determine execution sequence, and enforces this predetermined order during floating point reductions, ensuring reproducibility while supporting multiple simultaneous operations.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If collective operations are processed sequentially, then simplicity is maintained, but latency for short collectives and bandwidth for long collectives are not optimized

Engineering Contradiction:
Improvecollective reduction latency and bandwidthVSAvoidnetwork switch architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transitions from sequential processing to parallel processing by adding a spatial dimension to the architecture. It implements multiple independent processing paths (embedded network logic path and network hub logic path) that operate simultaneously on different collective operations. This dimensional expansion enables concurrent execution, optimizing both latency for short collectives and bandwidth for long collectives.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If arbitrary input port and output port mapping is enabled, then versatility for different collective classes improves, but device complexity increases

Engineering Contradiction:
Improveinput output port mapping flexibilityVSAvoidswitch hardware architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements universality by designing the network switch with a unified architecture that handles multiple collective classes through a single network hub logic. The hub receives input from multiple input ports, performs arbitration, and routes to multiple output ports, providing arbitrary input-output mapping capability. This universal design achieves versatility without proportionally increasing complexity by consolidating control functions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10425358B2Network switch architecture supporting multiple simultaneous collective operations
Publication Date: 2019.09.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10425358B2 patent drawing
  • US10425358B2 patent drawing
  • US10425358B2 patent drawing

AI summary

An apparatus includes a collective switch hardware architecture, including an input arrangement circuit including multiple input ports and multiple outputs. The input arrangement circuit routes its multiple input ports to selected ones of its outputs. The collective switch hardware architecture includes collective reduction logic coupled to the multiple outputs of the input arrangement circuit and having multiple outputs. The collective reduction logic includes ALU(s) and arbitration and control circuitry. The ALU(s) and arbitration and control circuitry support multiple simultaneous collective operations from different collective classes, and support arbitrary input port and output port mapping to different collective classes. The collective switch hardware architecture further includes an output arrangement circuit including a multiple inputs coupled to the multiple outputs of the collective reduction logic and including multiple output ports. The output arrangement circuit is configured to route its multiple inputs to selected ones of its output ports.