Dragonfly Routing With Global Fault Tables for Incomplete Connectivity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large networks face issues with resource sharing among different communications, leading to unpredictable latency and bandwidth, especially in high-performance computing (HPC) environments, where congestion can cause significant delivery delays and packet drops, affecting scalability and network performance.
Innovation Solution
The network is subdivided into groups with edge and global ports, using a global fault table or list to manage routing around faulty links, and implementing flow channels with dynamic flow setup and congestion control based on flow-specific acknowledgments to ensure fair and efficient traffic management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If traffic classes provide priority isolation for important communications, then delivery latency for high-priority traffic is improved, but lower-priority traffic becomes crippled by resource exhaustion
Solution Approach 1:
The patent applies local quality by differentiating buffer allocation and congestion control parameters across different traffic classes. Each traffic class receives tailored buffer space allocation and congestion notification thresholds, allowing high-priority traffic to receive preferential treatment while ensuring low-priority traffic maintains sufficient resources and fairness guarantees, preventing complete resource exhaustion for any single class
Solution Approach 2:
The patent implements feedback mechanisms where each traffic class receives congestion notifications based on its specific buffer occupancy thresholds. This feedback loop allows the network to dynamically adjust transmission rates for each class independently, enabling high-priority traffic to maintain low latency while low-priority traffic receives appropriate resource allocation and congestion management, preventing either class from being completely crippled
2Productivity
If packets are dropped to avoid head-of-line-blocking, then bandwidth utilization is improved, but packet loss increases causing retransmissions and latency spikes
Solution Approach 1:
The patent applies preliminary action by pre-configuring per-traffic-class buffer space allocation and congestion thresholds before congestion occurs. This allows the network to proactively manage buffer occupancy for each class, enabling packets to be dropped in a controlled manner based on class-specific criteria rather than uniform policies, thereby maintaining bandwidth utilization while minimizing unnecessary packet loss and retransmissions
Solution Approach 2:
The patent implements local quality by applying differentiated buffer management and congestion control strategies for each traffic class. Each class has its own buffer occupancy thresholds and congestion notification mechanisms, allowing the system to make granular decisions about packet dropping based on class-specific characteristics, thereby optimizing bandwidth utilization while minimizing the impact on delivery latency for any single class
3Reliability
If network buffers are increased to prevent packet drops, then delivery reliability is improved, but buffer exhaustion still occurs causing congestion and latency increases
Solution Approach 1:
The patent applies segmentation by dividing the network buffer into separate per-traffic-class buffers rather than using a single shared buffer pool. This segmentation allows independent management of buffer occupancy for each class, preventing any single class from exhausting all buffer resources and causing system-wide congestion. Each class maintains its own buffer occupancy thresholds and congestion control mechanisms, ensuring reliable packet delivery while managing latency through class-specific buffer management
Solution Approach 2:
The patent implements feedback mechanisms where each traffic class receives congestion notifications based on its specific buffer occupancy status. This per-class feedback loop enables dynamic adjustment of transmission rates independently for each class, allowing the network to maintain high packet delivery reliability through adequate buffer allocation while preventing buffer exhaustion-induced latency spikes by proactively managing buffer occupancy at the class level
Data Source
AI summary
Systems and methods are provided for managing a data communication within a multi-level network having a plurality of switches organized as groups, with each group coupled to all other groups via global links, including: at each switch within the network, maintaining a global fault table identifying the links which lead only to faulty global paths, and when the data communication is received at a port of a switch, determine a destination for the data communication and, route the communication across the network using the global fault table to avoid selecting a port within the switch that would result in the communication arriving at a point in the network where its only path forward is across a global link that is faulty; wherein the global fault table is used for both a global minimal routing methodology and a global non-minimal routing methodology.


