Global Bandwidth-Aware Adaptive Routing for Ethernet Switches
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Ethernet communication systems face challenges in efficiently managing bandwidth and adapting routing in response to changes in network conditions, such as link failures or congestion, which can lead to reduced performance and congestion.
Innovation Solution
The implementation of global bandwidth-aware adaptive routing in Ethernet communications, which involves determining events associated with changes in network bandwidth and using routing protocols to modify adaptive routing algorithms, allowing for selection from different routes based on downstream capacity information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If adaptive routing is based on local switch states (queue length and port utilization), then network utilization is maximized through load balancing, but upstream switches cannot shift traffic away from downstream devices experiencing reduced bandwidth capacity due to link failures or congestion
Solution Approach 1:
The patent implements a feedback mechanism where downstream switches monitor their own bandwidth capacity and send telemetry data upstream to inform routing decisions. This allows the network to react to changes in downstream conditions by dynamically adjusting routing paths, resolving the contradiction between maximizing utilization and maintaining reliability under varying network conditions.
Solution Approach 2:
The patent introduces an intermediary telemetry system that bridges the information gap between downstream switch states and upstream routing decisions. This intermediary mechanism enables upstream switches to make informed routing choices based on actual downstream capacity, allowing them to shift traffic away from congested or failed links while maintaining overall network utilization.
2Productivity
If AI training workloads operate at high utilization, then processing efficiency is improved, but link failures cause congestion and performance degradation
Solution Approach 1:
The patent implements dynamic routing that continuously adapts to changing network conditions. When link failures or congestion occur, the system dynamically reroutes traffic to maintain high utilization on available paths while avoiding congested ones, thus preserving both processing efficiency and performance stability in AI training workloads.
Solution Approach 2:
The patent changes routing parameters based on real-time network conditions, specifically using downstream bandwidth capacity information to adjust path selection. This allows the system to maintain high utilization by dynamically selecting optimal paths that avoid failures and congestion, thereby preserving both efficiency and stability.
3Reliability
If fabric-based flow control mechanisms (PFC) are used to manage congestion, then bandwidth is protected, but head-of-line blocking and network spread occur causing performance degradation
Solution Approach 1:
The patent extracts the congestion control function from the fabric-based PFC mechanism and relocates it to the routing layer. By making routing bandwidth-aware of downstream conditions, the system prevents congestion before it occurs rather than reacting to it, thereby protecting bandwidth without inducing head-of-line blocking or network spread.
Data Source
AI summary
Systems and methods herein are for global bandwidth-aware adaptive routing in a network communication and include at least one switch to determine an event associated with a change in network bandwidth between a local host and a remote host, where the at least one switch is further to provide routing protocols for the network communication, and where the routing protocols is to be used to modify an adaptive routing in the at least one switch for selection from different routes for the network communication between the local host and the remote host.


