Sideband Performance Tracing for HPC Network Congestion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-performance computing (HPC) systems face challenges in detecting and correcting communication performance issues due to complexities in large-scale computing environments, such as network congestion and system noise, which can lead to performance degradation, especially in partitioned global address space applications.

Innovation Solution

A system for sideband performance tracing of network traffic that includes a source endpoint node, a network computing device, and a target endpoint node, where trace data is recorded and transmitted separately from the network packet, allowing for the reconstruction of performance traces and identification of performance issues within the HPC network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If trace data is embedded in every network packet, then performance monitoring completeness is improved, but network overhead and complexity increase

Engineering Contradiction:
Improveperformance monitoring completenessVSAvoidnetwork overhead
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the tracing system into two independent components: (1) a sideband tracing channel that captures performance data separately from network packets, and (2) the main network data channel that transmits only application data. This segmentation allows performance monitoring without embedding trace data in every packet, reducing network overhead while maintaining monitoring completeness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a sideband channel as an intermediary mechanism that captures performance metrics (timestamps, queue depths, latency) separately from the main network traffic. This intermediary channel enables comprehensive performance tracing without adding complexity to the primary data transmission path.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Difficulty of detecting and measuring

If performance tracing is enabled in HPC networks, then detection of performance issues is improved, but system complexity and difficulty of operation increase

Engineering Contradiction:
Improvedetection of performance issuesVSAvoidsystem complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSEase of operation

Solution Approach 1:

The sideband tracing system automatically captures performance metrics (timestamps, queue depths, latency measurements) without requiring manual configuration or intervention. The system self-services by continuously monitoring and recording performance data in the sideband channel, making detection of performance issues easier while maintaining manageable system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback mechanisms where traced performance data is continuously monitored and can trigger alerts or notifications when performance thresholds are exceeded. This feedback loop enables automatic detection and reporting of performance issues, reducing the complexity of manual monitoring while improving detection capability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10135711B2Technologies for sideband performance tracing of network traffic
Publication Date: 2018.11.20 INTEL CORP
  • US10135711B2 patent drawing
  • US10135711B2 patent drawing
  • US10135711B2 patent drawing

AI summary

Technologies for tracing network performance include a network computing device configured to receive a network packet from a source endpoint node, process the received network packet, capture trace data corresponding to the network packet as it is processed by the network computing device, and transmit the received network packet to a target endpoint node. The network computing device is further configured to generate a trace data network packet that includes at least a portion of the captured trace data and transmit the trace data network packet to the destination endpoint node. The destination endpoint node is configured to monitor performance of the network by reconstructing a trace of the network packet based on the trace data of the trace data network packet. Other embodiments are described herein.