OPTIMIZATION OF THE SELECTION OF RIVERS TO BE DIVIDED

The system optimizes flow rerouting in network fabrics by using midpoint and ingress devices to detect and manage congestion, addressing load imbalances and improving network efficiency.

DE102025106957A1Pending Publication Date: 2026-03-26HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Network fabrics experience load imbalances and congestion due to persistent flows, leading to inefficient rerouting that affects overall network performance and cost.

Method used

A system that optimizes flow rerouting by using midpoint and ingress network devices to detect congestion and generate load metrics, allowing for intelligent redirection based on various parameters and conditions, including probabilistic models and pause mechanisms.

Benefits of technology

Improves network efficiency by reducing congestion and optimizing flow selection, leading to enhanced performance and cost-effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system generates a load metric for each flow from a first set of received flows through a network device acting as an intermediary device. The system sends a redirect acknowledgment (ACK) containing the load metric for that flow to a first input network device in response to the load metric being greater than a load value. The system then forwards a second set of flows through a network device acting as a second input network device. The system receives redirect ACKs from multiple intermediary devices, each corresponding to a multiple of flows from the second set. Each redirect ACK contains a load metric for that specific flow from the multiple of flows. Based on a set of redirection conditions, the system selects a first flow to redirect. The system then redirects the first flow to a new path.
Need to check novelty before this filing date? Find Prior Art

Description

STATEMENT ON STATE-FUNDED RESEARCH

[0001] This application was filed with government support under contract number H98230-15-D-0022 / 0003, issued by the Maryland Office of Procurement. The government holds certain rights to this invention. BACKGROUND

[0002] A network fabric can include entry network devices, intermediate or "midpoint" network devices, and exit network devices. The paths through the network fabric for ordered flows can be selected based on the load. Some flows, such as persistent flows, can lead to a load imbalance over time, and some paths may be more heavily used than others. Congestion can be detected by a midpoint network device when a packet for a flow is received. The midpoint network device can forward the detected "midpoint congestion" to the entry network device, allowing the entry network device to reroute the flow to a new path. However, rerouting flows can impact the cost and efficiency of the network fabric. BRIEF DESCRIPTION OF THE FIGURES Fig. Figure 1 shows an environment that facilitates the optimization of the selection of streams to be redirected according to one aspect of the present application. Fig. Figure 2 shows an environment that facilitates the optimization of the selection of streams to be redirected according to one aspect of the present application. Fig. Figure 3A shows a flowchart illustrating a procedure for optimizing the selection of streams to be redirected, including a network device acting as an intermediate network device, in accordance with one aspect of the present application. Fig. Figure 3B shows a flowchart illustrating a procedure for optimizing the selection of streams to be redirected, including a network device acting as an input network device, in accordance with an aspect of the present application. Fig. Figure 3C shows a flowchart representing a procedure that facilitates the optimization of the selection of currents to be diverted, including stopping a diverted current, according to one aspect of the present application. Fig. 3D represents a flowchart illustrating a procedure that facilitates the optimization of the selection of currents to be diverted, including the diversion of a current, in accordance with an aspect of the present application. Fig. Figure 4 shows a computer system that facilitates the optimization of the selection of currents to be diverted according to one aspect of the present application. Fig. Figure 5 shows a computer-readable medium that facilitates the optimization of the selection of currents to be diverted according to one aspect of the present application.

[0003] In the illustrations, identical numbers refer to the same elements of the illustration. DETAILED DESCRIPTION

[0004] Aspects of the present application provide a system that facilitates the optimization of the selection of flows to be rerouted, including the decision of whether or not to reroute a flow. The system can be based on congestion detected by an intermediate network device and on congestion managed by an entry network device.

[0005] A network fabric can include input network devices, intermediate network devices, and output network devices. Paths through the network fabric for ordered flows can be selected based on load. A flow can follow the same selected path while data is pending in the network fabric. Some flows, such as persistent flows that run over a long period, can lead to load imbalances over time; for example, the load may change over time, with some paths being used more than others.

[0006] A congestion can occur in the middle of the network fabric (i.e., a "mid-fabric congestion" or "mid-point congestion," detected by an intermediate or mid-point network device) or at an exit point of the network fabric (i.e., an "end-point congestion," detected by an exit or endpoint network device) when a packet is received for a flow. It can happen that too many flows attempt to use the same connection, resulting in a surplus of packets waiting in a queue to receive their share of the connection's bandwidth. Redirecting a flow that is congested at an endpoint and has already reached the exit network device may not provide any benefit.In contrast, rerouting a mid-network congestion can improve overall network efficiency because the rerouted flow is likely to be directed to a different mid-network link with fewer flows and more available bandwidth to accommodate additional packets. A mid-point congestion can be detected by a mid-point network device when a packet is received for a flow, and the mid-point device can then report the detected congestion to the ingress network device, allowing the ingress device to reroute the flow to a new path. However, rerouting flows can impact the cost and efficiency of the network fabric.

[0007] The described aspects provide a system that facilitates the optimization of stream selection based on congestion detected by a midpoint network device and congestion managed by an ingress network device. A midpoint network device can detect congestion associated with a received stream sent by an ingress network device (i.e., a midpoint congestion) when a packet is received for that stream. The midpoint network device can generate a load metric for the received stream. This load metric can be based on various parameters, such as the bandwidth consumption of all streams entering the midpoint network device and the size of a packet within a particular stream.If the load metric is greater than a predefined or preconfigured load value, the middle network device can send a "Redirect Acknowledgement (ACK)" back to the incoming network device, containing the generated load metric. The determination of whether to generate and send a Redirect Acknowledgement (ACK) is discussed further below, for example, in relation to... Fig. 3A described.

[0008] After receiving multiple redirect ACKs corresponding to multiple flows, the inbound network device can select one flow (corresponding to an original path) to redirect. The inbound network device can optimize the selection of the traffic flow to be redirected based on several techniques. In one technique, the inbound network device can pause and "drain" the selected flow, i.e., wait for the return of pending ACKs. If, while waiting for the selected flow to be drained, the original path is offered as the redirection path more than a certain number of times, the inbound network device can release the flow and simply use the original path.In some cases, the input network device can release the flow to a next-hop network device on the original path, but the next-hop network device may still wait for the flow to expire before choosing a different path. Otherwise, the input network device can redirect the flow to a new path.

[0009] In another technique, the input network device can store the load metric contained in the redirect ACK corresponding to a flow (e.g., the redirected flow). When the input network device receives a second redirect ACK from the same flow (e.g., the redirected flow) on the new path, it can store the load metric contained in the second redirect ACK. The input network device can then use the stored information to determine whether to select the flow for redirection or to perform another redirection operation on the redirected flow.

[0010] In another technique, the input network device can base the decision of whether to select a flow for redirection on various redirection conditions, including, but not limited to: a time span that has elapsed since a last redirected flow; a quantity of data to be sent in a given flow; a comparison of the stored load metric of the given flow with the load metrics of the other flows; and the difference, if available, between the stored load metrics of received redirect ACKs corresponding to the same flow.

[0011] Fig. Figure 1 shows an environment 100 that facilitates the optimization of the selection of streams to be redirected according to an aspect of the present application. The environment 100 can comprise a network 110 of switches, which can be referred to as a "switch fabric" and can include switches 112, 114, 116, 118, and 120. Each switch can have a unique address or identifier within the switch fabric 110. Various types of endpoints, processing nodes, devices, and networks can be connected to a switch fabric. For example, a storage array 130 can be connected to the switch fabric 110 via switch 112; an HPC network (e.g., InfiniBand, Slingshot, or another high-performance network) 132 can be connected to the switch fabric 110 via switch 114; a number of end hosts, such asHosts 136 and 138 can be connected to switch fabric 110 via switch 118; and an Internet Protocol (IP) / Ethernet network 134 can be connected to switch fabric 110 via switch 120. The HPC network 132 can include multiple networked computer and storage devices running programs simultaneously to perform various complex and resource-intensive tasks. The IP / Ethernet network 134 can include physical Ethernet cabling and an IP-based application layer protocol between network devices, including communication via Transport Communication Protocol (TCP) / IP and User Datagram Protocol (UDP) packets. Switch fabric 110 can itself be an Ethernet network or an HPC network.

[0012] In general, a switch can have edge ports and fabric ports. An edge port can be connected to a device located outside the fabric. A fabric port can be connected to another switch within the fabric via a fabric link. Typically, traffic can enter the switch fabric via an input port of an edge switch and exit the switch fabric via an output port of another (or the same) edge switch. An input link can connect a network interface controller (NIC) of an edge device (e.g., an HPC end host) to an input edge port of an edge switch. The switch fabric can then transport the traffic to an output edge switch, which in turn can forward the traffic to a destination edge device via another NIC.A packet can be routed in Switch Fabric 110 based on its Layer 2 address (“Fabric address”), which can be considered the equivalent of a MAC (Media Access Control) address in Ethernet. The routing path for the packet can be determined based on adaptive forwarding, for example, based on the local programming of the switches in Switch Fabric 110 and the information about load, traffic, and congestion that is available to and associated with Switch Fabric 110.

[0013] In some aspects, Switch Fabric 110 or HPC Network 132 can include network devices (i.e., switches), including input network devices, mid-network or midpoint network devices, and output or endpoint network devices. A switch in Switch Fabric 110 can contain systems that perform operations related to an input network device, a mid-network device, and an output network device. For example, Switch 118 can be an input network device for data originating from Device 136 and destined for IP / Ethernet Network 134 (with Switch 120 as the output network device for such data), and Switch 118 can also be an output network device for data originating from IP / Ethernet Network 134 and destined for Device 136 (with Switch 120 as the input network device for such data). Furthermore, a switch in Switch Fabric 110 can also contain systems that perform operations related to midpoint network devices.For example, switch 118 can be an intermediate network device for data originating from IP / Ethernet 134 and destined for the HPC network 132, e.g., via a possible path that includes switch 120 (acting as the inbound network device), switch 118 (acting as the intermediate network device), and switch 114 (acting as the outbound network device). Thus, a single switch can contain systems that perform functions related to an inbound network device, an intermediate network device, and an outbound network device.

[0014] Another example: Data transferred from IP / Ethernet network 134 ("source") to HPC network 132 ("destination") can enter switch fabric 110 via inbound network device 120 and be transferred to the outbound network device 114 via intermediate network device 116. Based on this data transferred from source to destination, switch 116 can receive an initial set of flows and generate a load metric for each flow. The load metric can be based on a current load associated with switch 116, determined, for example, by the depth of an outbound queue at switch 116, which stores pending packets awaiting transmission. The load can be expressed as an explicit Congestion Avoidance (ECA) value. The ECA value can contain a specific number of bits (e.g.,11 bits) and can indicate a degree or severity of congestion on the link, as determined by Switch 116 in the middle of network fabric 110. The ECA can be an input used to determine whether to generate an ACK. The current load of Switch 116 can also be based on the size of a packet in a particular flow of the first set of received flows. Furthermore, the decision to generate a redirect ACK associated with Switch 116 can be based on a product of the load and the packet size. In some aspects, the load metric can, for example,The decision is based on the following: bandwidth consumption associated with the detecting switch or network device; the amount of pending data in an input buffer associated with the detecting switch or network device; information received by a NIC that is associated with the amount of pending data to be processed by the detecting switch or network device; and information associated with the state of a given flow (e.g., the flow of the first set of flows received by intermediate network device 116, or the flow of a second set of flows forwarded by switch 120). Other metrics can also be used to determine whether or not to send the redirect ACK.

[0015] Switch 116 can determine whether a load metric for a given flow from the first set of flows is greater than a predetermined load value. The predetermined load value can be a randomly generated number, another number, or a threshold. The predetermined load value can be selected or preconfigured by the system or an administrative user connected to Network Fabric 110 or Switch 116. If the load metric is greater than the predetermined load value, Switch 116 can send a redirect ACK containing the generated load metric for that flow to the ingress network device 120. If the load metric is less than the predetermined load value, Switch 116 can choose not to send the redirect ACK to the ingress network device 120. In some cases, Switch 116 can compare the load metric to the predetermined load value if the load metric is greater than a predetermined threshold, for example...a preliminary or initial threshold.

[0016] Switch 120 (acting as the input network device in the example shown in environment 100) can forward a second set of flows, including flows destined for HPC network 132 via Switch 114 (acting as the output network device). This second set of flows can be forwarded via network fabric 110, including through Switch 116 (acting as an intermediate network device) and Switch 118 (also acting as an intermediate network device). The intermediate network devices receiving the second set of flows can detect congestion in the middle when a packet for a flow is received and send redirect ACKs containing a load metric for that flow. Switch 120 can receive redirect ACKs from multiple intermediate network devices, such as Switches 116 and 118, corresponding to multiple flows in the second set of flows.

[0017] Switchboard 120 can select a first flow to be redirected from the plurality of flows corresponding to the received redirect ACKs (indicating congestion in the middle). The first flow can be associated with a first path and correspond to a first redirect ACK with a first load metric. The selection of the first flow to be redirected can be based on a set of redirection conditions, including but not limited to: a time elapsed since a last redirected flow; a data volume awaiting transmission in a given flow from the plurality of flows; a comparison of the load metric of the given flow from the plurality of flows with load metrics of other flows in the plurality of flows; or a difference, if available, between load metrics contained in received redirect ACKs corresponding to the same flow.The set of redirection conditions can be associated with a probability that a corresponding flow will be selected from the majority of flows for redirection. This probability can increase based on an increase in load (e.g., a rising ECA value returned in the redirect ACK) or an increase in packet size. For example, the system can select the first flow to be redirected using a probabilistic model based on the ECA value.

[0018] Switch 120 can redirect the first flow to a new path and store an entry for the redirected first flow in a data structure. This entry can contain the first load metric. In some cases, after redirecting the first flow, Switch 120 can receive a second redirect ACK corresponding to the redirected first flow. This second redirect ACK can be sent by an internetwork device and can contain a second load metric. Switch 120 can store this second load metric in the entry for the redirected first flow. When deciding whether to select the flow for redirection again, Switch 120 can determine a difference between the second load metric and the first load metric. Switch 120 can adjust the probability of selecting the first flow for redirection based on this difference. For example, a small difference (i.e.,A difference of less than a first predetermined value indicates that the congestion on the new path for the first diverted flow has not improved and that the first flow may be a candidate for diversion. Conversely, a large difference (i.e., greater than a second predetermined value) may indicate that the congestion on the new path has improved and that diverting the flow is less advantageous. Consequently, the probability of selecting the first flow for diversion can be adjusted by the input network device.

[0019] Before rerouting the first flow, Switch 120 can also pause the first flow and initiate a waiting period. For example, Switch 120 can wait until the first flow has "run dry," meaning until Switch 120 has received a predetermined number of pending ACKs for the first flow. During the pause or waiting period, Switch 120 can "repeatedly" offer the original path for the first flow. For example, if Switch 120 offers the original path more than a predetermined number of times (e.g., 10 times) or more than a predetermined rate (e.g., 5 times in 5 milliseconds) during a specific period (e.g., the last 10 milliseconds), Switch 120 can decide to release the first flow so that it continues to be routed along the first path. Thus, under certain circumstances, Switch 120 can refrain from rerouting the first path.The circumstances of "repeated" offers described above serve only as illustrative examples. Other metrics can also be used as thresholds for determining repeated offers that trigger the release of the first stream.

[0020] Fig. Figure 2 shows an environment 200 that facilitates the optimization of the selection of streams to be redirected according to an aspect of the present application. The environment 200 can include: input network devices 210, 220, 230, and 240; intermediate or midpoint network devices 212, 214, 216, 222, 224, 226, 232, 234, 236, 242, 244, and 246; and output network devices 218, 228, 238, and 248. The environment 200 can be the network fabric 110 of Fig. 1 are similar in that there can be multiple paths for data to travel from an input network device through one or more intermediate network devices to an output network device. The data can travel through the environment 200 via a plurality of paths, e.g.: a path 250 (represented by a solid line) from a network input 202 to a network device 210 (via a communication 250.1) to a network device 222 (via a communication 250.1) to a network device 224 (via a communication 250.2) to a network device 226 (via a communication 250.3) to a network device 218 (via a communication 250.4) and finally out to a network output 204 (via a communication 250.5); a path 280 (indicated by a dotted line) from network input 202 to network device 230 (via communication 280.0) to network device 232 (via communication 280.1) to network device 234 (via communication 280.2) to network device 236 (via a communication 280.3) to network device 248 (via a communication 280.4) and finally out to a network output 204 (via a communication 280.5); and a path 290 (indicated by an alternating dotted and dashed line) from network input 202 to network device 240 (via a communication 290.0) to network device 242 (via a communication 290.1) to network device 244 (via a communication 290.2) to network device 246 (via a communication 290.3) to network device 248 (via a communication 290.4) and finally out to a network output 204 (via a communication 290.5).

[0021] Furthermore, data can travel via a path 260 (indicated by a thick solid line) from network input 202 to network device 220 (via communication 260.0) to network device 222 (via communication 260.1) to network device 224 (via communication 260.2) to network device 226 (via communication 260.3) to network device 218 (via communication 260.4) and finally to network output 204 (via communication 260.5).

[0022] During operation, an intermediate network device can detect a midpoint congestion and an exiting network device can detect an endpoint congestion when a packet for a flow is received. For example, if a packet for a flow on path 250 or 260 is received, network device 222 (operating as an intermediate network device) can detect a midpoint congestion 206 (indicated by a bold "X") with respect to the flows originating from inbound network devices 210 and 220. If a packet for a flow on path 280 or 290 is received, network device 248 (operating as an exit network device) can detect an endpoint congestion 206 (indicated by a bold "X") with respect to the flows originating from inbound network devices 230 and 240.

[0023] Since the output network device 248 detects the endpoint overload (with respect to the currents on paths 250 and 260 originating from network devices 230 and 240) when the current has already reached the network output, rerouting these currents will not improve their performance. In such cases, the system may instead slow down the currents contributing to the overload at the network input (e.g., at 202).

[0024] Since the streams originating from network devices 210 and 220 have reached the intermediate network device 222 but have not yet arrived at the network output, redirecting these streams can improve performance. Each intermediate network device can receive streams and generate a load metric for each stream. As described above, the load metric can be based on a current load associated with a particular network device, such as the depth of an output buffer or queue on that device. For example, network device 222 can generate a load metric for the streams originating from network devices 210 and 220. Network device 222 can determine that the load metric for the stream originating from network device 220 is greater than a specific load value. This specific load value can be a preconfigured or predetermined value.In this way, network device 222 can detect a medium congestion 206. After detecting a congestion in medium 206, network device 222 can send a redirect ACK to inbound network device 220 (via a communication 265 to network device 220). In some aspects, network device 220 can be an intermediate network device that can send the redirect ACK to another inbound network device in network 202 (e.g., via a communication 266). Network device 220 (and the inbound network devices 210, 230, and 240 shown) can therefore perform functions associated with both an intermediate network device and an endpoint network device (as above with respect to switches 116 and 120). Fig. 1 described).

[0025] The input network device 220 can receive the redirect ACK from the intermediate network device 222 (via 265) indicating a mid-network congestion of flow 206 originating from network device 220 (on path 260). The input network device 220 can also receive other redirect ACKs from other intermediate network devices indicating a mid-network congestion with respect to other flows on other paths (not shown). Each redirect ACK can contain the load metric for the corresponding flow. The input network device 220 can determine a probability for selecting each flow to be redirected based on a set of redirection conditions, as described above with respect to switch 120. Fig. As described in section 1, based on probability and rerouting conditions, input network device 220 can select the outgoing stream from network device 220 (on path 260) and redirect this stream to a new path (path 270, as indicated by a dashed line), for example, from network device 220 to network device 212 (via communication 270.1), to network device 214 (via communication 270.2), to network device 216 (via communication 270.3), to network device 218 (via communication 270.4), and finally to a network output 204 (via communication 270.5). In some cases, network device 220 can be an intermediate network device and receive the redirected data on the new path 270 from network input 202 (via communication 270.0, as indicated by the dashed line).Thus, the network device 220 can perform the operations described above both with respect to the switch 120 (as a network input device) and to the switch 116 (as a network intermediate device). Fig. 1. Perform the operations performed as a network input device. The operations performed are described below in relation to the flowcharts in Fig. 3B, Fig. 3C and Fig. 3D, the overload management subsystem / instructions 430 of Fig. 4 and instructions 514-522 of Fig. 5 described. The operations performed as an internetwork device are described below with reference to the flowchart in Fig. 3A, the overload detection / instructions subsystem 420 of Fig. 4 and instructions 510-514 of Fig. 5 described.

[0026] Fig. Figure 3A shows a flowchart 300 illustrating a procedure that facilitates the optimization of the selection of flows to be redirected, including a network device acting as an intermediate network device, according to one aspect of the present application. Traffic can be routed through a system or network fabric and pass through many network devices, e.g., from input network devices through intermediate network devices to output network devices. A network device can contain instructions, subsystems, units, logic, hardware, firmware, or software components that enable the network device to perform operations as an input network device, an intermediate network device, or an output network device.

[0027] During operation, the system receives an initial set of flows (operation 302) through a network device that acts as the first intermediate network device in a network fabric. For example, the intermediate network device 222 in Fig. Two streams were received from connections 250.1 and 260.1. While in Fig. 2 where only two communications or streams to the intermediate network device 222 are shown, an intermediate network device can receive any number of streams, which can lead to the first set of streams.

[0028] The system generates a load metric for each flow from an initial set of received flows (Operation 304) via the network device that operates as the first intermediary device in the network fabric. The network device can generate the load metric based on a current load associated with the network device, indicated by the depth of its output buffer, which represents a quantity of data awaiting transmission. The decision of whether or not to generate a redirect ACK can be, for example,This may also be based on the following factors: an ECA value indicating the degree or severity of congestion on the link; the size of a packet in a given flow; the product of the load and the packet size; the current consumption of bandwidth connected to the network device; the amount of data waiting in an input buffer of the network device; and any information received by a network card or associated with the state of the given flow. If the amount of data waiting in the output buffer is greater than a predetermined threshold, the network device may determine that the load metric is greater than a load value, where this load value may be a predetermined threshold, an initial threshold, or some other limit set or determined by the system or an administrative user connected to the system or network device.

[0029] If the load metric is greater than a load value (Decision 306), the system, in response to the fact that the load metric is greater than the load value, sends a redirect acknowledgment (ACK) including the load metric for that flow to a first input network device associated with that flow (Operation 308). For example, when a packet is received for a flow, the intermediate network device 222 in Fig. 2. Detect a medium overload 206 (based on the generated load metric being greater than the load value) and send a redirect ACK to the input network device 220 (or another input network device in the input network 202), as above with respect to communication 265 from Fig. 2 described.

[0030] If the load metric is not greater than the load value (decision 306) (i.e., less than or equal to the load value), the system refrains from sending the redirect ACK to the first inbound network device in response to the load metric being less than or equal to the load value (operation 310). For example, in the case of the intermediate network device 222 in Fig. 2. To remain: If the intermediate network device 222 determines that the generated load metric is not greater than the load value, the intermediate network device 222 can refrain from sending the redirect ACK (e.g., it does not send communication 265). The process is labeled A in Fig. Continued in 3B.

[0031] Fig. Figure 3B shows a flowchart 330 illustrating a procedure that facilitates the optimization of the selection of streams to be redirected, including a network device acting as an ingress network device, in accordance with one aspect of the present application. During operation, the system routes a second set of streams through the network device acting as a second ingress network device in the network fabric (Process 332). For example, each of the network devices 210, 220, 230, and 240 in Fig. 2 act as an input network device and forward a second set of flows (which may differ from the first set of flows received by the intermediate network device in Procedure 302 of Fig. 3A is received). For the input network device 220, the second set of flows can include the flow indicated by communication path 260 (including communications 260.1-260.5).

[0032] The system receives a plurality of redirect ACKs from a plurality of internetwork devices, corresponding to a plurality of flows from the second set of flows, with each redirect ACK containing a load metric for a respective flow from the plurality of flows (Operation 334). As in Fig. As shown in Figure 2, the input network device 220 can receive the Redirect ACK 265 (as generated and transmitted by the intermediate network device 222 upon detection of congestion in the middle 206 when a packet of a stream is received). Although in Fig. Not shown, the input network device 220 can also receive other redirect ACKs generated and transmitted by other intermediate network devices upon detection of mid-point congestion for corresponding flows. Each redirect ACK can contain the generated load metric for the corresponding flow.

[0033] The system selects a first redirected flow from the plurality of flows based on a set of redirection conditions, where the first flow is associated with a first path and corresponds to a first redirect ACK with a first load metric (Operation 336). The set of redirection conditions can be used to determine a probability for selecting a particular redirected flow or to assign a ranking to the flows (e.g., an order in which the flows should be selected for redirection). The redirection conditions can, for example,The following are included: a time span that has elapsed since a last redirected flow; a data set that is pending transmission in each flow of the plurality of flows; a comparison of the load metric of each flow of the plurality of flows with load metrics of other flows in the plurality of flows; a difference, if available, between load metrics contained in received redirect ACKs corresponding to the same flow; and a ranking of the plurality of flows.

[0034] The system determines whether to stop the first flow before redirecting it to a new path or whether to redirect it to the new path (Decision 338). For example, the system may decide to stop the first flow if a configuration is set to initiate a wait based on tracked pending ACKs, and the system may decide to redirect the first flow if the probability of the first flow being redirected is greater than a threshold probability. If the system decides to stop the first flow before redirecting it to the new path (Decision 338), the operation is labeled B in Fig. 3C continues. If the system decides to redirect the first flow to the new path (decision 338), the operation is completed at label C in Fig. 3D continues. The operation can continue from operation 336 to decision 338 either to label B (pause) or to label C (redirection), simultaneously for different input network devices or flows. In some cases, the system does not execute decision 338 and instead continues from operation 336 to either label B or label C.

[0035] Fig. Figure 3C shows a flowchart 340 illustrating a procedure that facilitates the optimization of the selection of flows to be redirected, including stopping a flow that can be redirected, according to one aspect of the present application. During operation, the system, through the network device acting as the second input network device, stops the first flow before redirecting the first flow to the new path (Process 342). Fig. 2. The input network device 220 can pause or stop data associated with the flow (“first flow” via path 260) before that first flow is redirected to a new path.

[0036] The system waits until at least a predetermined number of pending ACKs associated with the first flow are received (Operation 344). The system (i.e., the network device acting as the second network input network device, such as input network device 220 of Fig. 2) The system can track the number of pending ACKs received in response to the sending of packets from the first flow. Alternatively, the system can wait until the downstream flow has deleted all packets, indicated by returned ACKs that represent the amount or quantity of data in the flow, rather than the number of packets required to send that data. The system can, but does not have to, maintain a one-to-one mapping of returned ACKs to sent packets. The predetermined number of pending ACKs can be configured to account for packet loss and can be a specific number or a percentage.For example, the input network device 220 may wait until at least twenty of the pending ACKs (or 80% or some other threshold) associated with the first flow (via path 260) have been received or returned, indicating that the data associated with the pending ACKs has been successfully transmitted to or through the output network device. In some cases, the input network device 220 may wait until almost all or all of the pending ACKs have been received.

[0037] If the specified number of pending ACKs is not received (decision 346), the operation returns to operation 344. If the specified number of pending ACKs is received (decision 346), the system determines whether the (same) first path is offered more than a specified number of times as a new path for the paused first flow that can be rerouted (decision 348).

[0038] If the (same) first path is not offered as a new path more than a predetermined number of times (e.g., 5) (decision 348), the process is terminated at label C. Fig. 3D continues. If the (same) first path is offered as a new path more than a predetermined number of times (e.g., 5) (Decision 348), the system releases the first flow so that it continues to be routed along the first path (Operation 350). Following the release of the first flow for further routing along the (same) first path, the system refrains from rerouting the first flow onto the new path (Operation 352). For example, if in Fig. 2. If the network device offers the same first path (path 260) as a new path no more than five times, the process is labeled C. Fig. 3D continued (i.e., if the network device offers the same first path (path 260) more than five times as a redirect or new path, the network device may (by tracking the offered path and the number of offers of the path) determine to release this first flow so that it continues to be routed on the original path (path 260), i.e., to release the first path 260 over which the packets received by the network device 222 triggered the originally detected congestion in the middle (206), and the network device may refrain from redirecting the flow (originally over path 260) to the new path (over path 270).

[0039] Fig. Figure 3D shows a 360-degree flowchart illustrating a procedure that facilitates the optimization of the selection of flows to be redirected, including the redirection of a flow, in accordance with an aspect of the present application. During operation, the system, through the network device acting as the second input network device, redirects the first flow to the new path (Operation 362). For example, input network device 220 can redirect the first flow (via path 260) to a new path (via path 270, as indicated by the dashed lines), as above in relation to Fig. 2 described.

[0040] The system stores an entry for the redirected first flow in a data structure of the network device acting as the second input network device, with the entry containing the first load metric (operation 364). The network device can store an entry for the first redirected flow, including identification information for the original flow (e.g., via path 260), identification information for the new or redirected path (e.g., via path 270), and initial load metric information determined or generated by the network device with respect to the first flow.

[0041] The system receives a second redirect ACK corresponding to the redirected first flow, with the second redirect ACK containing a second load metric (Operation 366). For example, the input network device 220, which is in Fig. Not shown in Figure 2, a further Redirect-ACK (second Redirect-ACK) is received from another intermediate network device, e.g., intermediate network device 234. The second Redirect-ACK may also contain identification information for the corresponding original flow (second flow), identification information for a new or redirected path, and second load metric information determined or generated by intermediate network device 234 with respect to the second flow.

[0042] The system stores the second load metric (Operation 368) in the entry for the redirected first flow. The data structure can be a table, a list, an array, or another type of data and related information storage. For example, with input network device 220 in Fig. To remain 2, which receives both the first and second redirect ACKs and stores the associated information, the input network device 220 can store the second load metric in the same entry as the first load metric.

[0043] The system calculates a difference between the second load metric contained in the second redirect ACK and the first load metric contained in the first redirect ACK (Operation 370). The network device acting as the second input network device can retain the data structure and can also perform and store the calculation of the difference in the data structure entry for the redirected first flow. The difference between the first and second load metrics can be expressed, for example, as: the difference between ECA values; the difference between bandwidth consumption; the difference between the number of pending bytes; and the difference based on how the first and second load metrics are calculated or measured.

[0044] The system adjusts the probability of selecting the first traffic flow to be rerouted based on the difference (Operation 372). For example, a small difference (e.g., less than three percent difference in measurements) might indicate that congestion has not significantly improved on the new or rerouted path. Consequently, the first flow might be marked as a strong candidate for rerouting, meaning the network device could increase the probability of selecting the first flow for rerouting. Conversely, a large difference (e.g., more than 60 percent difference in measurements) might indicate that congestion on the new or rerouted path has significantly improved. As a result, rerouting the first flow might be less advantageous, and the network device could mark the first flow as a weak candidate for rerouting.The designation as a "strong" or "weak" candidate is for illustrative purposes only. Other categories or types can also be used, such as levels, value ranges or windows, and a finite or limited number of categories assigned to each candidate from the set of received streams.

[0045] By allowing midpoint network devices to generate a metric and send redirect ACKs under certain circumstances, and by enabling ingress network devices to receive multiple redirect ACKs and make decisions about redirecting a flow based on various redirection conditions (as described here), the described aspects provide a system that can optimize the selection of flows to be redirected based on congestion detected by midpoint network devices (midpoint congestion) and congestion managed by ingress network devices. Optimizing flow selection can lead to improved performance and a more efficient overall system.

[0046] Fig. Figure 4 shows a computer system 400 that facilitates the optimization of the selection of streams to be redirected according to one aspect of the present application. The computer system 400 comprises a processor 402, a memory 404, and a storage device 406. The memory 404 can include volatile memory (e.g., random-access memory (RAM)) that serves as managed memory and can be used to store one or more memory pools. In addition, the computer system 400 can be coupled with peripheral I / O user devices 410 (e.g., a display device 411, a keyboard 412, and a pointing device 413). The storage device 406 comprises a non-transitory, computer-readable storage medium and stores an operating system 416, an overload detection subsystem / instructions 420, an overload management subsystem / instructions 430, and data 442. The computer system 400 can have fewer or more units or instructions than those shown in Figure 406. Fig. 4 shown.

[0047] Instructions 420 may contain instructions 422 and 424 which, when executed by the computer system 400, may cause the computer system 400 to perform the procedures and / or processes described in this disclosure, e.g., including the computer system 400 acting as an inter-network device. In particular, the computer system 400 may store instructions 422 to generate a load metric for a given flow from a first set of received flows, as above with respect to, e.g., the switch 116. Fig. 1 and process 304 from Fig. 3A described.

[0048] The computer system 400 can store instructions 424 to send a redirect ACK, including the load metric for the respective flow, to an incoming network device in response to the load metric being greater than a load value, as described above with respect to switches 118 and 120. Fig. 1 and process 308 from Fig. 3A described.

[0049] Instructions 430 may also include instructions 432, 434, 436, 438, and 440, which, when executed by the computer system 400, can cause the computer system 400 to perform the procedures and / or processes described in this disclosure, including, for example, the computer system 400 acting as an input network device. In particular, the computer system 400 may store instructions 432 to forward a second set of streams, as described above, for example, in relation to the forwarding of streams by the input network device 220 and the operation 332 of Fig. 3B.

[0050] The computer system 400 can further store instructions 434 to receive from a plurality of internetwork devices a plurality of redirect ACKs corresponding to a plurality of flows from the second set of flows, each redirect ACK containing a load metric for a respective flow from the plurality of flows. The reception of multiple redirect ACKs, each containing a load metric for a respective flow, is described above with respect to the input network device 220. Fig. 2 and process 334 of Fig. 3B described.

[0051] The computer system 400 can store instructions 436 to select a first flow to be redirected from the plurality of flows based on a set of redirection conditions, where the first flow is associated with a first path and corresponds to a first redirect ACK with a first load metric. The selection of a flow to be redirected can be based on a certain probability or a set of redirection conditions, as above with respect to the input network device 220. Fig. 2 and process 336 of Fig. 3B described.

[0052] The Computer System 400 can store instructions 438 to redirect the initial flow to a new path. Redirecting the initial flow can occur after the flow has been stopped, after waiting until a predetermined number of pending ACKs have been received, or after it has been determined that the same initial path has been offered a certain number of times compared to a predetermined number, as described above in relation to operations 340-348. Fig. 3C and Operation 352 of Fig. Described in 3D.

[0053] The computer system 400 can store instructions 440 to store an entry for the redirected first flow in a data structure, where the entry contains the first load metric, as above in relation to operation 364 of Fig. Described in 3D.

[0054] Commands 420 and 430 can contain more commands than those in Fig. 4 shown. For example, commands 420 and 430 can be used to perform the operations described above in relation to the environments of Fig. 1 and Fig. 2; the ones in the flowcharts of the Fig. 3A-D depicted processes; and instructions 510-522 of CRM 500 in Fig. 5.

[0055] Data 442 can contain all data required as input or generated as output by the methods, operations, communications, and / or processes described in this disclosure. In particular, the data 442 can store at least the following: a load metric; a flow; data of a flow; a load value; a predetermined value; the result of comparing a load metric with a load value; a redirect ACK; a redirect ACK corresponding to a flow and containing a load metric; a plurality of flows; a selected flow; a first path; an original path; an identical path; a new path; a path for redirecting a flow; a data structure; an entry in a data structure; a difference between load metrics; a probability of selecting a flow to be redirected; an adjusted probability; a condition; a redirection condition; a time span; a data set;a comparison between load metrics; a difference between load metrics; a ranking; a current load; a packet size; a product of load and packet size; bandwidth consumption; a quantity of data pending in an output or input buffer; information received from a NIC or associated with a flow state; a decision whether or not to send a redirect ACK; and a predetermined or preconfigured threshold.

[0056] Fig. Figure 5 shows a computer-readable medium (CRM) 500 that facilitates the optimization of the selection of flows to be redirected in accordance with an aspect of the present application. CRM 500 can be a non-transitory computer-readable medium or device that stores instructions which, when executed by a computer or processor, cause the computer or processor to perform a procedure. CRM 500 can store instructions 510 to generate a load metric for a particular flow from an initial set of received flows, as above with respect to, for example, switches 116 and 118. Fig. 1 and process 304 from Fig. 3A described.

[0057] CRM 500 can store instructions 512 to send a redirect ACK including the load metric for the respective flow in response to the load metric being greater than a load value, as above with respect to switches 118 and 120. Fig. 1 and process 308 from Fig. 3A described.

[0058] CRM 500 can store instructions 514 for forwarding a second set of streams, as described above, e.g., regarding the forwarding of streams through the input network device 220. Fig. 2 and process 332 of Fig. 3B.

[0059] CRM 500 can store instructions 516 to receive multiple redirect ACKs from multiple intermediate network devices, corresponding to multiple flows from the second set of flows, each redirect ACK containing a load metric for a specific flow from the multiple flows. Receiving multiple redirect ACKs, each containing a load metric for a specific flow, is described above with respect to the input network device 220. Fig. 2 and process 334 of Fig. 3B described.

[0060] CRM 500 can store instructions 518 to select a first redirected flow from the plurality of flows based on a set of redirection conditions, where the first flow is associated with a first path and corresponds to a first redirect ACK containing a first load metric. The selection of the first redirected flow (i.e., the "candidate flows") can be based on determining a probability for each flow or on one or more redirection conditions, including those given above as examples with respect to the input network device 220. Fig. 2 and process 336 of Fig. 3B were specified

[0061] CRM 500 can store instructions 520 to redirect the initial flow to a new path, as above in relation to the input network device 220. Fig. 2 (Representation of the detour to the path via 271-285, indicated by the dashed line) and operation 336 of Fig. 3B described.

[0062] CRM 500 can store instructions 522 to store an entry for the redirected first flow in a data structure, where the entry contains the first load metric, as above in relation to operation 364 of Fig. Described in 3D.

[0063] CRM 500 can handle more instructions than those in Fig. The 5 shown are included. CRM 500 can, for example, also provide instructions for performing the operations described above in relation to the environments of the Fig. 1 and Fig. 2; the ones in the flowcharts of the Fig. The processes shown in 3A-D and instructions 420 and 430 of computer system 400 in Fig. 4.

[0064] The term "network device" refers to a device, component, or computing unit that can provide a communication pipeline for packets sent from a "processing node" or an "endpoint node." A processing or endpoint node can refer to a device, component, or hardware part that can serve as the source or destination of data, such as a control packet or a data packet. A network device can be an inbound network device, an intermediate or middle network device, or an outbound or endpoint network device. An example of a network device is a switch, as discussed above. Fig. 1. A processing node or endpoint node can be an input node (which is an endpoint for data returning from a request) or an output node (which is an endpoint for data sent from a request). Furthermore, a network device can operate as an input network device, an intermediate network device, or an output network device, or perform the functions described herein.

[0065] In general, the disclosed aspects provide a computer system, a procedure, and a computer-readable medium that facilitate the optimization of the selection of flows to be redirected. The computer system operates in a network fabric comprising input network devices, intermediate network devices, and output network devices. The computer system includes a processor and a storage device that stores instructions for congestion detection and congestion management (also referred to as subsystems) which, when executed by the processor, are intended to perform the following operations. The congestion detection subsystem may contain instructions to: generate a load metric for a given flow from an initial set of received flows; and send a redirect acknowledgment (ACK) containing the load metric for that flow to an input network device in response to the load metric being greater than a load value.The congestion management subsystem may contain instructions to: receive an initial redirect ACK corresponding to an initial flow; and, based on a set of redirection conditions, determine whether to select the initial flow for redirection. The computer system may further contain instructions to perform the operations described herein, including those relating to: the environments of the . Fig. 1 and Fig. 2; the ones in the flowcharts of the Fig. 3A-D depicted processes; and the instructions of CRM 500 in Fig. 5.

[0066] In one variation of this aspect, the congestion management instructions are further designed to: redirect a second set of flows, including the first flow, where the first flow is associated with a first path and where the first redirect ACK indicates a first load metric; receive a plurality of redirect ACKs corresponding to a plurality of flows of the second set of flows from a plurality of intermediate network devices, where the plurality of redirect ACKs includes the first redirect ACK and where each redirect ACK contains a load metric for a corresponding flow of the plurality of flows; determine that, from the plurality of flows, the first flow is selected to be redirected based on the set of redirection conditions; and redirect the first flow to a new path.The congestion management instructions further include: storing an entry for the redirected first flow in a data structure, wherein the entry contains the first load metric; receiving a second redirect ACK corresponding to the redirected first flow, wherein the second redirect ACK contains a second load metric; and storing the second load metric in the entry for the redirected first flow.

[0067] In another variation of this aspect, the congestion management instructions also serve to determine a difference between the second load metric contained in the second redirect ACK and the first load metric contained in the first redirect ACK. The congestion management instructions further serve to adjust a probability for selecting the first current to be redirected based on this difference.

[0068] In another variant of this aspect, the set of diversion conditions is linked to a probability that a corresponding river will be selected for diversion from the majority of rivers.

[0069] In another variant, the set of redirection conditions includes at least one of the following elements: a time span that has elapsed since a last redirected flow; a data set that is pending transmission in a given flow of the plurality of flows; a comparison of the load metric of the given flow from the plurality of flows with load metrics of other flows in the plurality of flows; a difference, if available, between load metrics contained in received redirect ACKs corresponding to the same flow; or a ranking of the plurality of flows.

[0070] In another variant, the generated load metric for the respective flow from the first set of flows in the congestion detection subsystem and the load metric for the corresponding flow from the plurality of flows in the congestion management subsystem is based on at least one of the following factors: a load assigned to the congestion detection subsystem or the congestion management subsystem, expressed as an explicit congestion avoidance value (ECA); or a packet size in the respective flow from the first set of flows or in the corresponding flow from the plurality of flows.

[0071] In another variant, the generated load metric for the respective flow from the first set of flows in the congestion detection subsystem and the load metric for the corresponding flow from the plurality of flows in the congestion management subsystem comprise: a product of the load and the packet size for the respective flow in the congestion detection subsystem or in the congestion management subsystem.

[0072] In another variant, the generated load metric for the respective flow from the first set of flows in the congestion detection subsystem and the load metric for the corresponding flow from the plurality of flows in the congestion management subsystem are based on at least one of the following factors: bandwidth consumption associated with the congestion detection subsystem or the congestion management subsystem; a set of data waiting in an input buffer assigned to the congestion detection subsystem or the congestion management subsystem; information received from a network interface controller (NIC) and associated with a set of data awaiting processing by the congestion detection subsystem or the congestion management subsystem;or information that is associated with a state of the respective flow of the first set of flows in the congestion detection subsystem or the corresponding flow from the plurality of flows in the congestion management subsystem.

[0073] In another variant, the congestion management instructions are to halt the first flow before redirecting it to a new path. Furthermore, the congestion management instructions are to wait until at least a predetermined number of pending ACKs associated with the first flow are received. Finally, the congestion management instructions are to release the first flow, allowing it to continue along the first path and preventing its redirection, in response to waiting until the predetermined number of pending ACKs has been received and in response to the first path being offered more than a predetermined number of times.

[0074] In another variant, the overload detection instructions should refrain from sending the Redirect-ACK to the input network device if the load metric is smaller than the load value.

[0075] In another variant, the overload detection instructions are also used to compare the load metric with the load value if the load metric is greater than a specified threshold.

[0076] In another variant, the load value consists of a randomly generated number.

[0077] In another aspect, a computer-implemented procedure can encompass various operations performed, for example, by a system. The system generates a load metric for each flow within a first set of received flows through a network device acting as the first intermediary device in a network fabric. The system sends a redirect acknowledgment (ACK) containing the load metric for that flow to a first input network device associated with that flow, in response to the load metric being greater than a load value. The system refrains from sending the redirect ACK to the first input network device if the load metric is less than the load value. The system then forwards a second set of flows through the network device acting as the second input network device in the network fabric.The system receives multiple redirect ACKs from multiple internetwork devices, corresponding to multiple flows from the second set of flows, each redirect ACK containing a load metric for that specific flow. Based on a set of redirection conditions, the system selects a first flow to be redirected from the multiple flows. The first flow is associated with a first path and corresponds to a first redirect ACK containing a first load metric. The system then redirects the first flow to a new path. The procedure may include additional operations, including those related to the environments of the system. Fig. 1 and Fig. 2; the ones in the flowcharts of the Fig. 3A-D depicted processes; instructions 420 and 430 of computer system 400 in Fig. 4; and instructions 510-522 of CRM 500 in Fig. 5.

[0078] In another aspect, a non-transitory computer-readable storage medium (or CRM) stores instructions to generate a load metric for each flow from a first set of received flows. These instructions further serve to send a redirect acknowledgment (ACK) containing the load metric for that flow in response to the load metric being greater than a load value. The instructions further serve to forward a second set of flows. The instructions further serve to receive multiple redirect ACKs from multiple internetwork devices, corresponding to multiple flows from the second set of flows, with each redirect ACK containing a load metric for a specific flow from that multiple set of flows.The instructions further serve to select a first flow to be redirected from the plurality of flows based on a set of redirection conditions, where the first flow is associated with a first path and corresponds to a first redirect ACK containing a first load metric. The instructions further serve to redirect the first flow to a new path. The instructions further serve to store an entry for the redirected first flow in a data structure, where the entry contains the first load metric. The CRM can also store instructions to perform the operations described above with respect to the environments of the . Fig. 1 and Fig. 2; the ones in the flowcharts of the Fig. 3A-D depicted operations; instructions 420 and 430 of computer system 400 in Fig. 4; and instructions 510-522 of CRM 500 in Fig. 5.

[0079] The foregoing description is intended to enable the person skilled in the art to produce and use the aspects and examples and is given in connection with a specific application and its requirements. Various modifications of the disclosed aspects are readily apparent to the person skilled in the art, and the general principles defined herein can be applied to other aspects and applications without departing from the spirit and scope of this disclosure. Therefore, the aspects described here are not limited to those shown but have the broadest possible scope consistent with the principles and features disclosed herein.

[0080] Furthermore, the foregoing descriptions of the aspects serve only for illustration and description. They do not claim to be exhaustive and do not limit the aspects described herein to the disclosed forms. Accordingly, many modifications and variations will be obvious to those skilled in the art. Moreover, the above disclosure is not intended to limit the aspects described herein. The scope of the aspects described herein is defined by the accompanying claims.

Claims

[1] A computer system operating in a network fabric containing inbound network devices and inter-network devices, wherein the computer system comprises: a subsystem for overload detection and a subsystem for overload management; where the subsystem for overload detection is intended for: Generating a load metric for each flow of an initial group of received flows and Sending a redirect acknowledgment (ACK) containing the load metric associated with the respective flow to an input network device in response to the load metric being greater than a load value; and The overload management subsystem is designed for: Receiving an initial redirect ACK, which corresponds to an initial flow; and Determine whether to select the first flow to be diverted, based on a set of diversion conditions. [2] Computer system according to claim 1, wherein the overload management subsystem is further designed to: Redirecting a second set of flows containing the first flow, where the first flow is associated with a first path and where the first redirect ACK indicates a first load metric; Receiving a plurality of redirect ACKs corresponding to a plurality of flows from the second set of flows, from a plurality of intermediate network devices, wherein the plurality of redirect ACKs includes the first redirect ACK and wherein each redirect ACK contains a load metric for a corresponding flow from the plurality of flows; Determining the selection of the first flow to be diverted from the plurality of flows based on the set of diversion conditions; Diverting the first river onto a new path; Storing an entry for the redirected first flow in a data structure, where the entry contains the first load metric; Receiving a second redirect ACK corresponding to the redirected first flow, where the second redirect ACK contains a second load metric; and Store the second load metric in the entry for the redirected first flow. [3] Computer system according to claim 2, wherein the overload management subsystem is further designed to: Determine the difference between the second load metric contained in the second redirect ACK and the first load metric contained in the first redirect ACK; and Adjusting a probability for selecting the first flow to be diverted based on the difference. [4] Computer system according to claim 2, wherein the set of diversion conditions is associated with a probability that a particular river is selected from the plurality of rivers for diversion. [5] Computer system according to claim 2, wherein the set of redirection conditions comprises at least one of the following: a period of time that has passed since the last diverted river; a set of data to be sent in a given stream from the plurality of streams; a comparison of the load metric of the respective river from the majority of rivers with load metrics of other rivers in the majority of rivers; a difference, if available, between load metrics included in received redirect ACKs for the same flow; or a ranking of the majority of rivers. [6] Computer system according to claim 2, wherein the generated load metric for the respective flow from the first set of flows in the overload detection subsystem and the load metric for the corresponding flow from the plurality of flows in the overload management subsystem are based on at least one of the following: a load assigned to the overload detection subsystem or the overload management subsystem, expressed as an explicit overload avoidance value (ECA); or a size of a packet in the respective river from the first set of rivers or in the corresponding river from the plurality of rivers. [7] Computer system according to claim 6, wherein the generated load metric for the respective flow from the first set of flows in the overload detection subsystem and the load metric for the corresponding flow from the plurality of flows in the overload management subsystem comprise the following: a product of the load and packet size for the respective flow in the overload detection subsystem or the overload management subsystem. [8] Computer system according to claim 2, wherein the generated load metric for the respective flow from the first set of flows in the overload detection subsystem and the load metric for the corresponding flow from the plurality of flows in the overload management subsystem are based on at least one of the following: Bandwidth consumption allocated to the congestion detection subsystem or the congestion management subsystem; a set of data that is waiting in an input buffer assigned to the overload detection subsystem or the overload management subsystem; Information received from a network interface controller (NIC) and associated with a data set awaiting processing by the congestion detection subsystem or the congestion management subsystem; or Information that is associated with a state of the respective flow of the first set of flows in the congestion detection subsystem or the corresponding flow from the plurality of flows in the congestion management subsystem. [9] Computer system according to claim 2, wherein the overload management subsystem is further designed to: Pause the first river before the first river is diverted onto a new path; Wait until at least a predetermined number of pending ACKs associated with the first flow have been received; and in response to waiting until the predetermined number of pending ACKs has been received, and in response to the first path being offered more than a predetermined number of times: Releasing the first river to allow further routes on the first path; and Failure to redirect the first path. [10] Computer system according to claim 1, wherein the overload detection subsystem is further designed to: In response to the load metric being below the load value, the redirect ACK is not sent to the incoming network device. [11] Computer system according to claim 1, wherein the overload detection subsystem is further designed to: Comparing the load metric with the load value in response to the load metric being greater than a predetermined threshold. [12] Computer system according to claim 1, wherein the load value comprises a randomly generated number. [13] A computer-implemented method comprising: Generating a load metric for each flow from a first set of received flows by a network device operating as the first intermediate network device in a network fabric; Sending a redirect acknowledgment (ACK) to an initial input network device associated with the respective flow, containing the load metric associated with that flow, in response to the load metric being greater than a load value; Failure to send the Redirect-ACK to the first input network device in response to the load metric being less than the load value; Receiving an initial redirect ACK, which corresponds to an initial flow; and Determine whether to select the first flow to be diverted, based on a set of diversion conditions. [14] Computer-implemented method according to claim 13, further comprising: Redirecting a second set of flows containing the first flow, where the first flow is associated with a first path and where the first redirect ACK indicates a first load metric; Receiving a plurality of redirect ACKs corresponding to a plurality of flows from the second set of flows, from a plurality of intermediate network devices, wherein the plurality of redirect ACKs includes the first redirect ACK and wherein each redirect ACK contains a load metric for a respective flow from the plurality of flows; Determining the selection of the first river to be diverted from the majority of rivers based on the set of diversion conditions and Diverting the first river onto a new path. [15] Computer-implemented method according to claim 14, further comprising: Storing an entry for the redirected first flow in a data structure by the network device acting as the second input network device, where the entry contains the first load metric; Receiving a second redirect ACK corresponding to the redirected first flow, where the second redirect ACK contains a second load metric; Storing the second load metric in the entry for the diverted first flow; Calculate the difference between the second load metric contained in the second redirect ACK and the first load metric contained in the first redirect ACK; and Adjusting the probability of selecting the first flow to be diverted based on the difference. [16] Computer-implemented method according to claim 14, where the set of diversion conditions is associated with a probability with which a corresponding river is selected from the plurality of rivers for diversion, and wherein the set of redirection conditions includes at least one of the following: a period of time that has passed since the last diverted river; a set of data that are ready to be sent in the respective stream from the majority of streams; a comparison of the load metric of the respective river from the majority of rivers with load metrics of other rivers in the majority of rivers; a difference, if available, between load metrics contained in received redirect ACKs corresponding to the same flow; or an ordered list containing the majority of the rivers. [17] Computer-implemented method according to claim 14, wherein the generated load metric for the respective flow from the first set of flows and the load metric for the corresponding flow from the plurality of flows are based on at least one of the following: a load assigned to the respective flow from the first set of flows or the corresponding flow from the plurality of flows, expressed as an explicit Congestion Avoidance Value (ECA); or a size of a packet in the respective river from the first set of rivers or in the corresponding river from the plurality of rivers. [18] Computer-implemented method according to claim 14, wherein the generated load metric for the respective flow from the first set of flows and the load metric for the corresponding flow from the plurality of flows are based on at least one of the following: Bandwidth consumption allocated to the network device operating as the first intermediate network device or the second input network device; a set of data that is waiting in an input buffer connected to the network device that is acting as the first intermediate network device or as the second input network device; Information received from a network interface controller (NIC) and associated with a data set awaiting processing by the network device acting as the first intermediate network device or the second input network device; or Information that is assigned to a state of the respective flow of the first set of flows or the corresponding flow of the plurality of flows. [19] Computer-implemented method according to claim 14, further comprising: Pausing the first flow through the network device acting as the second input network device before the first flow is redirected to the new path; Wait until at least a predetermined number of pending ACKs associated with the first flow have been received; and in response to waiting until the predetermined number of pending ACKs has been received, and in response to the first path being offered more than a predetermined number of times: Releasing the first river to allow further routes on the first path; and Failure to redirect the first path. [20] A non-transitory, computer-readable medium that stores instructions for: Generating a load metric for each flow of a first group of received flows; Sending a redirect acknowledgment (ACK) including the load metric for the respective flow in response to the load metric being greater than a load value; Diverting a second group of rivers; Receiving a plurality of redirect ACKs corresponding to a plurality of flows from the second set of flows, from a plurality of intermediate network devices, where each redirect ACK contains a load metric for a particular flow from the plurality of flows; Selecting a first river to be diverted from the plurality of rivers based on a set of diversion conditions, where the first flow is associated with a first path and corresponds to a first redirect ACK with a first load metric; Diverting the first river onto a new path and Storing an entry for the redirected first flow in a data structure, where the entry contains the first load metric.