Scalable InfiniBand Diagnostic Tool for Switch Performance Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large-scale computer systems with switched fabric topologies, component failures can significantly degrade performance, making it crucial to detect and locate failures quickly and accurately to maintain system efficiency.
Innovation Solution
A method is implemented where a control node in a distributed memory computer system performs data transfer tests between pairs of switches to assess their performance, comparing results against minimum standards, and identifies underperforming components by grouping switches based on topology data and conducting individual component tests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional diagnostic tools are run on each individual switch or node to detect port errors, then failure detection capability is provided, but the complexity of the diagnostic process increases and system performance degrades due to the scale of the network
Solution Approach 1:
The patent segments the large-scale network into hierarchical groups (pods, tiers, layers) based on topology data. Instead of diagnosing each switch individually, the system organizes switches into manageable segments that can be tested collectively. This reduces diagnostic complexity by grouping switches into hierarchical structures where tests can be performed at multiple levels of abstraction.
Solution Approach 2:
The patent creates a universal diagnostic framework that works across different network topologies and scales. The same high-level test methodology can be applied regardless of network size or specific configuration, making the diagnostic process scalable and topology-agnostic. This multi-functional approach allows a single diagnostic system to handle various network configurations without requiring topology-specific customizations.
2Measurement precision
If comprehensive data transfer tests are performed across all switches to ensure system performance, then measurement precision of switch performance is improved, but the time required for diagnostics increases
Solution Approach 1:
The patent performs preliminary actions by first conducting high-level aggregate tests across groups of switches before drilling down to individual switch diagnostics. The system uses topology data to pre-organize switches into testable groups and performs initial broad assessments. Only when anomalies are detected does the system proceed to more time-consuming detailed individual switch tests, significantly reducing overall diagnostic time while maintaining measurement precision for problematic areas.
Solution Approach 2:
The patent applies partial action by performing comprehensive tests on only the necessary subset of switches rather than all switches in the network. The high-level diagnostic framework identifies specific groups or individual switches that require detailed testing based on aggregate performance metrics. This selective approach maintains measurement precision for critical components while avoiding unnecessary testing of healthy switches, thereby reducing total diagnostic time.
3Measurement precision
If individual port error detection is performed on each switch, then failure location precision is improved, but the overall system productivity decreases due to the large number of processors and switches
Solution Approach 1:
The patent segments the network into hierarchical groups (pods, tiers, layers) that can be tested collectively. Instead of individually testing every switch and port, the system organizes switches into segments that share common characteristics or physical groupings. This allows failure detection to be performed at the segment level first, then narrowed down to specific switches and ports only when necessary, maintaining failure location precision while preserving system productivity.
Solution Approach 2:
The patent performs preliminary aggregate tests at the group level before conducting individual port error detection. By first testing groups of switches collectively using topology-based methodologies, the system can identify problematic segments without immediately invoking time-consuming individual port tests across the entire network. This preliminary filtering action maintains the ability to locate failures precisely while minimizing the impact on system productivity.
Data Source
AI summary
In accordance with some implementations, a method for evaluating large scale computer systems based on performance is disclosed. A large scale, distributed memory computer system receives topology data, wherein the topology data describes the connections between the plurality of switches and lists the nodes associated with each switch. Based on the received topology data, the system performs a data transfer test for each of the pair of switches. The test includes transferring data between a plurality of nodes and determining a respective overall test result value reflecting overall performance of a respective pair of switches for a plurality of component tests. The system determines that the pair of switches meets minimum performance standards by comparing the overall test result value against an acceptable test value. If the overall test result value does not meet the minimum performance standards, the system reports the respective pair of switches as underperforming.


