Optimizing selection of rerouted streams
By detecting congestion and generating load metrics through midpoint network devices, and combining redirection confirmations and rerouting conditions from ingress network devices, flow selection is optimized, resolving the load imbalance problem caused by persistent flows in the network structure, improving network efficiency and reducing costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2026-03-24
AI Technical Summary
In the existing network architecture, the load imbalance and uneven path usage caused by persistent flows lead to network efficiency and cost issues, especially when the intermediate structure is congested and the rerouting of flows is not optimized enough.
The midpoint network device detects congestion and generates load metrics. The midpoint network device generates a redirection acknowledgment (ACK). The ingress network device selects flows for rerouting based on the load metrics and rerouting conditions, including pausing and storing load metric information to optimize flow selection.
It improves the overall efficiency and performance of the network structure, reduces network costs, optimizes flow rerouting, and enhances system load balancing.
Smart Images

Figure CN121728006A_ABST
Abstract
Description
[0001] Statement of Government-Sponsored Research
[0002] This application was made with government support under contract number H98230-15-D-0022 / 0003 awarded by the Maryland Procurement Office. The government has certain rights in the application. BACKGROUND
[0003] Network structures can include ingress network devices, intermediate or "midpoint" network devices, and egress network devices. Paths through a network structure for ordered flows can be selected based on load. Some flows, such as persistent flows, can cause load imbalance over time, and some paths can be used more frequently than others. When a packet of a flow is received, congestion can be detected by a midpoint network device. The detected "midpoint congestion" can be relayed by the midpoint network device to an ingress network device and the ingress network device can be allowed to re-route the flow to a new path. However, re-routing a flow can impact the cost and efficiency of the network structure. BRIEF DESCRIPTION OF DRAWINGS
[0004] FIG. 1 An environment that facilitates optimizing the selection of flows to be re-routed is illustrated in accordance with an aspect of the present application.
[0005] FIG. 2 An environment that facilitates optimizing the selection of flows to be re-routed is illustrated in accordance with an aspect of the present application.
[0006] FIG. 3A A flow diagram illustrating a method that facilitates optimizing the selection of flows to be re-routed, including a network device operating as an intermediate network device, in accordance with an aspect of the present application is presented.
[0007] FIG. 3B A flow diagram illustrating a method that facilitates optimizing the selection of flows to be re-routed, including a network device operating as an ingress network device, in accordance with an aspect of the present application is presented.
[0008] FIG. 3C A flow diagram illustrating a method that facilitates optimizing the selection of flows to be re-routed, including pausing flows that can be re-routed, in accordance with an aspect of the present application is presented.
[0009] FIG. 3D A flow diagram illustrating a method that facilitates optimizing the selection of flows to be re-routed, including re-routing flows, in accordance with an aspect of the present application is presented.
[0010] FIG. 4A computer system that facilitates optimization of selection of flows to be rerouted is illustrated in accordance with an aspect of the present application.
[0011] FIG. 5 A computer readable medium that facilitates optimization of selection of flows to be rerouted is illustrated in accordance with an aspect of the present application.
[0012] In the drawings, like reference numerals refer to like elements throughout. DETAILED DESCRIPTION
[0013] Aspects of the present application provide a system that facilitates optimization of selection of flows to be rerouted, including whether to reroute a flow. The system can be based on congestion detected by mid-point network devices and congestion managed by ingress network devices.
[0014] A network fabric can include ingress network devices, intermediate network devices, and egress network devices. Paths through the network fabric for in-order flows can be selected based on load. Flows can follow the same selected paths while data for the flows is pending in the network fabric. Some flows, such as persistent flows that last for long periods of time, can cause load imbalance over time (e.g., load can change over time), and some paths can be used more frequently than other paths.
[0015] Congestion can occur in the middle of the network fabric (i.e., "mid-fabric congestion" or "mid-point congestion" detected by intermediate or mid-point network devices) or at the egress of the network fabric (i.e., "end-point congestion" detected by egress or end-point network devices) when a packet for a flow is received. Too many flows can attempt to share the same link, which can cause too many packets waiting in a queue to be given their share of link bandwidth. Rerouting flows that encounter end-point congestion and have already reached the egress network device can not provide a benefit. In contrast, rerouting flows that encounter mid-fabric congestion can result in overall efficiency of the network fabric improving because the rerouted flows will likely be directed onto a different mid-fabric link that has fewer flows and idle bandwidth to accommodate more packets. Mid-point network devices can detect mid-point congestion when a packet for a flow is received, and the mid-point network devices can relay the detected mid-point congestion to the ingress network devices and allow the ingress network devices to reroute the flow to a new path. However, rerouting flows can impact the cost and efficiency of the network fabric.
[0016] The described aspects provide a system that facilitates optimizing the selection of flows to be rerouted based on congestion detected by a midpoint network device and congestion managed by an ingress network device. When a packet of a flow is received, the midpoint network device can detect congestion associated with the received flow that was transmitted by the ingress network device (i.e., midpoint congestion). The midpoint network device can generate a load metric for the received flow. The load metric can be based on various parameters, such as, for example, the bandwidth consumption of all flows entering the midpoint network device and the size of the packets in the particular flow. If the load metric is greater than a predetermined or preconfigured load value, the midpoint network device can return a "redirect acknowledgment (ACK)" to the ingress network device that includes the generated load metric. The following is described with respect to, for example, FIG. 3A Determining whether to generate and transmit a redirect ACK is described.
[0017] Upon receiving multiple redirect ACKs corresponding to multiple flows, the ingress network device can select a flow (corresponding to an original path) to be rerouted. The ingress network device can optimize the selection of the flow to be rerouted based on several techniques. In one technique, the ingress network device can stop and "exclude" the selected flow, i.e., wait for the pending ACK to be returned. While waiting for the selected flow to be excluded, if the original path is provided as the path for rerouting the flow more than a certain number of times, the ingress network device can release the flow and simply use the original path. In some aspects, the ingress network device can release the flow to the next hop network device on the original path, but the next hop network device can still wait for the flow to be excluded before selecting a different path. Otherwise, the ingress network device can reroute the flow onto a new path.
[0018] In another technique, the ingress network device can store the load metric included in the redirect ACKs corresponding to a flow (e.g., a rerouted flow). If the ingress network device receives a second redirect ACK from the same flow (e.g., a rerouted flow) on a new path, the ingress network device can store the load metric included in the second redirect ACK. The ingress network device can subsequently use the stored information to determine whether to select the flow for rerouting or whether to perform another rerouting operation on the rerouted flow.
[0019] In another technique, the ingress network device can decide whether to select a flow for rerouting based on various rerouting conditions, including but not limited to, for example: the amount of time that has passed since the last rerouted flow; the amount of data pending transmission in the corresponding flow; a comparison of the stored load metric of the corresponding flow to the load metrics of other flows; and, if any, the difference between the stored load metrics of the received redirect ACKs corresponding to the same flow.
[0020] FIG. 1An environment 100 that facilitates optimization of selection of flows to be re-routed is illustrated in accordance with an aspect of the present application. The environment 100 can include a network 110 of switches (which can be referred to as a "switch fabric") and can include switches 112, 114, 116, 118, and 120. Each switch can have a unique address or identifier within the switch fabric 110. Various types of endpoints, processing nodes, devices, and networks can be coupled to the switch fabric. For example, a storage array 130 can be coupled to the switch fabric 110 via switch 112; a high performance computing (HPC) network (e.g., InfiniBand, Slingshot, or any other high performance network) 132 can be coupled to the switch fabric 110 via switch 114; a plurality of end hosts (e.g., host 136 and host 138) can be coupled to the switch fabric 110 via switch 118; and an Internet Protocol (IP) / Ethernet network 134 can be coupled to the switch fabric 110 via switch 120. The HPC network 132 can include a plurality of networked computers and storage devices running parallel programs to accomplish different complex and performance intensive tasks. The IP / Ethernet network 134 can include physical Ethernet cabling and application layer protocols between IP-based network devices, including communication via Transmission Communication Protocol (TCP) / IP and User Datagram Protocol (UDP) packets. The switch fabric 110 itself can be an Ethernet or HPC network.
[0021] Generally, a switch can have edge ports and fabric ports. An edge port can be coupled to a device that is outside the fabric. A fabric port can be coupled to another switch within the fabric via a fabric link. Generally, traffic can be injected into the switch fabric 110 via an ingress port of an edge switch and can exit the switch fabric 110 via an egress port of another (or the same) edge switch. An ingress link can couple a network interface controller (NIC) of an edge device (e.g., an HPC end host) to an ingress edge port of an edge switch. The switch fabric 110 can then transmit the traffic to an egress edge switch, which in turn can deliver the traffic to a destination edge device via another NIC. Packets can be forwarded in the switch fabric 110 based on their Layer 2 addresses ("fabric addresses"), which can be considered equivalent to Media Access Control (MAC) addresses in Ethernet. The forwarding path of a packet can be determined based on adaptive forwarding, e.g., based on local programming of switches in the switch fabric 110 and information related to load, traffic, and congestion available to and associated with the switch fabric 110.
[0022] In some aspects, switch architecture 110 or HPC network 132 may include network devices (i.e., switches) including ingress network devices, intermediate or midpoint network devices, and egress or endpoint network devices. The switches in switch architecture 110 may include systems that perform operations associated with the ingress, intermediate, and egress network devices. For example, switch 118 may be an ingress network device for data originating from device 136 and destined for IP / Ethernet 134 (with switch 120 acting as an egress network device for such data), and switch 118 may also be an egress network device for data originating from IP / Ethernet 134 and destined for device 136 (with switch 120 acting as an ingress network device for such data). Additionally, the switches in switch architecture 110 may also include systems that perform operations associated with midpoint network devices. For example, switch 118 may be an intermediate network device for data originating from IP / Ethernet 134 and destined for HPC network 132, for example, via possible paths including switch 120 (acting as an ingress network device), switch 118 (acting as an intermediate network device), and switch 114 (acting as an egress network device) Therefore, a single switch can include a system that performs functions associated with ingress network devices, intermediate network devices, and egress network devices.
[0023] As another example, data propagating from IP / Ethernet 134 (“source”) to HPC network 132 (“destination”) can enter switch structure 110 via ingress network device 120 and propagate to egress network device 114 via intermediate network device 116. Based on this data propagated from source to destination, switch 116 can receive a first set of flows and generate a load metric for each flow. The load metric can be determined based on the current load associated with switch 116 and, for example, the depth of the output queue on switch 116 storing pending packets waiting to be transmitted. The load can be represented as an Explicit Congestion Avoidance (ECA) value. ECA can include a number of bits (e.g., 11 bits) and can indicate the level or severity of congestion on the link, as determined by switch 116 at the midpoint of network structure 110. ECA can be an input for determining whether an ACK should be generated. The current load associated with switch 116 can also be based on the size of packets in a given flow of the first set of received flows. Furthermore, the decision for generating a redirect ACK associated with switch 116 can be based on the product of the load and the packet size. In some respects, load metrics may be based on, for example, the following: bandwidth consumption associated with the detection switch or network device; the amount of data pending in the input buffer associated with the detection switch or network device; information received from the NIC and associated with the amount of pending data to be processed by the detection switch or network device; and information associated with the status of the corresponding flow (e.g., a flow in a first set of flows received by intermediate network device 116 or a flow in a second set of flows forwarded by switch 120). Other metrics may also be used to determine whether to send a redirect ACK.
[0024] Switch 116 can determine whether the load metric of a corresponding flow in the first group of flows is greater than a predetermined load value. The predetermined load value can be a randomly generated number, another number, or a threshold. The predetermined load value can be selected or pre-configured by the system or a management user associated with network structure 110 or switch 116. If the load metric is greater than the predetermined load value, switch 116 can send a redirection ACK including the generated load metric of the corresponding flow to ingress network device 120. If the load metric is less than the predetermined load value, switch 116 can avoid sending a redirection ACK to ingress network device 120. In some aspects, switch 116 can compare the load metric with the predetermined load value in response to the load metric being greater than a predetermined threshold (e.g., a preliminary threshold or initial threshold).
[0025] Switch 120 (operating as an ingress network device in the continued example depicted in environment 100) can forward a second set of flows, including flows destined for HPC network 132 via switch 114 (operating as an egress network device). The second set of flows can be forwarded via network structure 110, including via switch 116 (operating as an intermediate network device) and switch 118 (also operating as an intermediate network device). An intermediate network device receiving the second set of flows can detect midpoint congestion upon receiving packets of the flows and send a redirection ACK including a load metric for the corresponding flow. Switch 120 can receive redirection ACKs corresponding to multiple flows in the second set of flows from multiple intermediate network devices (e.g., switches 116 and 118).
[0026] Switch 120 can select a first flow for rerouting from among the plurality of flows corresponding to a received redirect ACK (which indicates midpoint congestion). The first flow may be associated with a first path and may correspond to a first redirect ACK including a first load metric. Selection of the first flow for rerouting may be based on a set of rerouting conditions, including but not limited to: the amount of time elapsed since the most recently rerouted flow; the amount of data to be sent in the corresponding flow among the plurality of flows; a comparison of the load metric of the corresponding flow among the plurality of flows with the load metrics of other flows among the plurality of flows; or, if applicable, the difference between the load metrics included in the received redirect ACK corresponding to the same flow. This set of rerouting conditions may be associated with the probability that the corresponding flow from among the plurality of flows will be selected for rerouting. The probability may increase based on an increase in load (e.g., an increased ECA value returned in the redirect ACK) or an increase in packet size. For example, the system may select the first flow for rerouting based on the ECA value using a probabilistic model.
[0027] Switch 120 can reroute a first flow to a new path and can also store an entry for the rerouted first flow in a data structure. This entry may include a first load metric. In some aspects, after rerouting the first flow, switch 120 can receive a second redirect ACK corresponding to the rerouted first flow. The second redirect ACK may be sent by an intermediate network device and may include a second load metric. Switch 120 can store the second load metric in the entry for the rerouted first flow. When determining whether to re-select the flow for rerouting, switch 120 can determine the difference between the second load metric and the first load metric. Switch 120 can adjust the probability of selecting the first flow for rerouting based on this difference. For example, a small difference (i.e., less than a first predetermined value) may indicate that congestion on the new path for the flow used for the first rerouting has not improved, and the first flow may be a candidate for rerouting. On the other hand, a large difference (i.e., greater than a second predetermined value) may indicate that congestion on the new path has improved and rerouting the flow may not be beneficial. Therefore, the probability that the first flow will be selected for rerouting can be adjusted by the ingress network device.
[0028] Before rerouting the first flow, switch 120 may also pause the first flow and initiate a waiting period. For example, switch 120 may wait until the first flow is "excluded," that is, until switch 120 has received a predetermined number of pending ACKs associated with the first flow. During the pause or waiting period, switch 120 may "repeatedly" provide the original path to the first flow. For example, if switch 120 provides the original path more than a predetermined number of times (e.g., 10 times) or exceeds a predetermined rate (e.g., 5 times within 5 milliseconds) during a certain time period (e.g., the most recent 10 milliseconds), switch 120 may determine to release the first flow to continue routing on the first path. Therefore, in some cases, switch 120 may avoid rerouting the first path. The above-described "repeated" provision is provided only as an illustrative example. Other metrics may be used as thresholds for determining the repeated provision that triggers the release of the first flow.
[0029] FIG. 2 The illustration depicts an environment 200 that facilitates the selection of flows to be rerouted according to one aspect of this application. Environment 200 may include: ingress network devices 210, 220, 230, and 240; intermediate or midpoint network devices 212, 214, 216, 222, 224, 226, 232, 234, 236, 242, 244, and 246; and egress network devices 218, 228, 238, and 248. Environment 200 may be similar to... FIG. 1The network structure 110 can specifically have multiple paths for propagating data from an ingress network device through one or more intermediate network devices to an egress network device. Data can propagate through environment 200 via multiple paths, for example: path 250 (indicated by solid lines), from network ingress 202 to network device 210 (via communication 250.1) to network device 222 (via communication 250.1) to network device 224 (via communication 250.2) to network device 226 (via communication 250.3) to network device 218 (via communication 250.4), and finally output to network egress 204 (via communication 250.5); path 280 (indicated by dashed lines), from network ingress 202 to network device 230 (via communication 280.0) to network device 232 (via communication 280.1) to network device 218 (via communication 250.5). The path 234 (via communication 280.2) leads to network device 236 (via communication 280.3) to network device 248 (via communication 280.4), and finally outputs to network egress 204 (via communication 280.5); and path 290 (indicated by alternating dotted lines) leads from network ingress 202 to network device 240 (via communication 290.0) to network device 242 (via communication 290.1) to network device 244 (via communication 290.2) to network device 246 (via communication 290.3) to network device 248 (via communication 290.4), and finally outputs to network egress 204 (via communication 290.5).
[0030] Additionally, data can travel via path 260 (indicated by the thick solid line) from network entry 202 to network device 220 (via communication 260.0), then to network device 222 (via communication 260.1), then to network device 224 (via communication 260.2), then to network device 226 (via communication 260.3), then to network device 218 (via communication 260.4), and finally be output to network exit 204 (via communication 260.5).
[0031] During operation, intermediate network devices can detect midpoint congestion, and egress network devices can detect endpoint congestion when receiving packets from a flow. For example, when receiving packets from a flow on path 250 or 260, network device 222 (operating as an intermediate network device) can detect midpoint congestion 206 (indicated by a bold "X") related to the flow originating from ingress network devices 210 and 220. When receiving packets from a flow on path 280 or 290, network device 248 (operating as an egress network device) can detect endpoint congestion 206 (indicated by a bold "X") related to the flow originating from ingress network devices 230 and 240.
[0032] Because egress network device 248 detects endpoint congestion (related to flows originating from network devices 230 and 240 on paths 250 and 260) when the flows have already reached the network egress, rerouting these flows will not help improve their performance. In this case, the system can instead slow down the flows that are causing congestion at the network ingress (e.g., at 202).
[0033] Conversely, since flows originating from network devices 210 and 220 have reached intermediate network device 222 but have not yet reached the network egress, rerouting these flows may improve performance. Each intermediate network device can receive flows and generate a load metric for each flow. As mentioned above, the load metric can be based on the current load associated with the respective network device (e.g., the depth of the output buffer or queue on the respective network device). For example, network device 222 can generate a load metric for flows originating from network devices 210 and 220. Network device 222 can determine that the load metric for a flow originating from network device 220 is greater than a specific load value. The specific load value can be a pre-configured or predetermined value. Therefore, network device 222 can detect intermediate congestion 206. Upon detecting intermediate congestion 206, network device 222 can send a redirection ACK to ingress network device 220 (via communication 265 to network device 220). In some aspects, network device 220 can be an intermediate network device that can send a redirection ACK to another ingress network device in network ingress 202 (e.g., via communication 266). Therefore, network device 220 (and the depicted ingress network devices 210, 230, and 240) can perform functions associated with both intermediate and endpoint network devices (as described above). FIG. 1 (As described in switches 116 and 120).
[0034] Ingress network device 220 can (via 265) receive from intermediate network device 222 a redirection ACK indicating midpoint congestion 206 associated with a flow originating from network device 220 (on path 260). Ingress network device 220 can also receive from other intermediate network devices additional redirection ACKs indicating midpoint congestion associated with other flows on other paths (not shown). Each redirection ACK may include a load metric for the corresponding flow. Ingress network device 220 can determine the probability of selecting each flow for rerouting based on a set of rerouting conditions, as described above. FIG. 1The switch 120 is described. Based on probability and rerouting conditions, the ingress network device 220 can select flows originating from network device 220 (on path 260) from these flows and can reroute the flow to a new path (such as path 270 indicated by the dashed line), for example, from network device 220 to network device 212 (via communication 270.1) to network device 214 (via communication 270.2) to network device 216 (via communication 270.3) to network device 218 (via communication 270.4), and finally output to network egress 204 (via communication 270.5). In some aspects, network device 220 can be an intermediate network device and can receive rerouting data on the new path 270 from network ingress 202 (via communication 270.0 indicated by the dashed line). Therefore, network device 220 can perform the above-described... FIG. 3B The operation of both switch 120 (as an ingress network device) and switch 116 (as an intermediate network device) is described below. FIG. 3C , FIG. 3D and FIG. 4 Flowchart in FIG. 5 The congestion management subsystem / instruction 430, and FIG. 3A Instructions 514-522 further describe the operations performed as an ingress network device. The following is about... FIG. 4 Flowchart in FIG. 5 The congestion detection subsystem / instruction 420, and FIG. 3A Instructions 510-514 further describe the operations performed as an intermediate network device.
[0035] FIG. 2 A flowchart 300 illustrating a method for optimizing the selection of flows to be rerouted, according to one aspect of this application, includes a network device operating as an intermediate network device. Traffic can be forwarded through a system or network structure and propagated through numerous network devices, for example, from an ingress network device via intermediate network devices to an egress network device. A network device may include instructions, subsystems, units, logic, hardware, firmware, or software components that allow the network device to perform operations as an ingress network device, intermediate network device, or egress network device.
[0036] During operation, the system receives a first set of flows (operation 302) through a network device operating as a first intermediate network device in the network structure. For example, FIG. 2 Intermediate network device 222 can receive streams from communications 250.1 and 260.1. Although in FIG. 2 Only two communications or streams to intermediate network device 222 are depicted, but the intermediate network device can receive any number of streams, thus enabling the generation of the first set of streams.
[0037] The system generates a load metric for the corresponding flow in the first set of received flows through a network device operating as the first intermediate network device in the network structure (operation 304). The network device may generate the load metric based on the current load associated with the network device (such as indicated by the depth of its output buffer, which represents the amount of pending data to be sent). The decision on whether to generate a redirect ACK may also be based on, for example, the following: the ECA value indicating the level or severity of congestion on the link; the size of the packets in the corresponding flow; the product of the load and the packet size; the current bandwidth consumption associated with the network device; the amount of pending data in the network device's input buffer; and any information received from the NIC or associated with the state of the corresponding flow. If the amount of pending data to be sent in the output buffer is greater than a predetermined threshold, the network device may determine that the load metric is greater than a load value, which may be a predetermined threshold, an initial threshold, or another limit set or determined by the system or an administrative user associated with the system or network device.
[0038] If the load metric is greater than the load value (Decision 306), the system responds by sending a redirection acknowledgment (ACK) including the load metric of the corresponding flow to the first ingress network device associated with the flow (Operation 308). For example, when a data packet of a flow is received, FIG. 2 Intermediate network device 222 can detect midpoint congestion 206 (based on a generated load metric exceeding the load value) and can send a redirection ACK to ingress network device 220 (or another ingress network device in network ingress 202), as described above. FIG. 2 The communication described in 265.
[0039] If the load metric is not greater than the load value (Decision 306) (i.e., less than or equal to the load value), the system avoids sending a redirection ACK to the first ingress network device (Operation 310) in response to the load metric being less than or equal to the load value. Continue FIG. 3B In the example of intermediate network device 222, if intermediate network device 222 determines that the generated load metric is not greater than the load value, then intermediate network device 222 can avoid sending a redirect ACK (e.g., not sending communication 265). Operation in FIG. 3B Continue at label A.
[0040] FIG. 2 A flowchart 330 illustrates a method for optimizing the selection of flows to be rerouted, according to one aspect of this application, including a network device operating as an ingress network device. During operation, the system forwards a second set of flows through the network device operating as a second ingress network device in the network structure (operation 332). For example, FIG. 3AAny of network devices 210, 220, 230, and 240 can operate as an ingress network device and can forward a second set of flows (which may differ from the ingress network device). FIG. 2 The first set of flows received by the intermediate network device in operation 302). For the ingress network device 220, the second set of flows may include the flows indicated by communication path 260 (including communications 260.1-260.5).
[0041] The system receives multiple redirection ACKs from multiple intermediate network devices, each corresponding to a flow in the second group of flows. The corresponding redirection ACK includes the load metric of the corresponding flow within those flows (Operation 334). For example... FIG. 2 As shown, the ingress network device 220 can receive a redirect ACK 265 (generated and sent by the intermediate network device 222 when midpoint congestion 206 is detected upon receiving a packet from the flow). Although not in FIG. 3C As depicted, the ingress network device 220 can also receive additional redirection ACKs generated and sent by other intermediate network devices when congestion is detected at the midpoint of the corresponding flow. Each redirection ACK may include a load metric generated for the corresponding flow.
[0042] The system selects a first flow for rerouting from a plurality of flows based on a set of rerouting conditions, wherein the first flow is associated with a first path and corresponds to a first redirect ACK including a first load metric (operation 336). This set of rerouting conditions can be used to determine the probability of selecting a corresponding flow for rerouting or to order flow assignments (e.g., the order in which flows are selected for rerouting). Rerouting conditions may include, for example: the amount of time elapsed since the most recently rerouted flow; the amount of pending data to be sent in the corresponding flow among the plurality of flows; a comparison of the load metric of the corresponding flow among the plurality of flows with the load metrics of other flows among the plurality of flows;, in some cases, the difference between the load metrics included in received redirect ACKs corresponding to the same flow; and the order in which the plurality of flows are arranged.
[0043] The system determines whether to pause the first flow before rerouting it to a new path, or to reroute the first flow to a new path (Decision 338). For example, if the configuration is set to initiate a wait period based on the tracked pending ACK, the system can determine to pause the first flow, and if the probability of the first flow being rerouted is greater than a threshold probability, the system can determine to reroute the first flow. If the system determines to pause the first flow before rerouting it to a new path (Decision 338), then the operation... FIG. 3D Continue at label B. If the system determines to reroute the first stream to the new path (Decision 338), then the operation continues at... FIG. 3CThe process continues at label C. For different ingress network devices or flows, the operation can continue from operation 336 to decision 338 and then proceed in parallel to either label B (pause) or label C (rerouting). In some respects, the system may not execute decision 338, but instead continue from operation 336 to either label B or label C.
[0044] FIG. 2 A flowchart 340 illustrates a method for optimizing the selection of flows to be rerouted, according to one aspect of this application, including pausing flows that can be rerouted. During operation, the system pauses the first flow via a network device operating as a second ingress network device (operation 342) before rerouting the first flow to a new path. FIG. 2 In this context, the ingress network device 220 can pause or stop data associated with the data stream (the “first stream” via path 260) before rerouting the first stream to a new path.
[0045] The system waits until it receives at least a predetermined number of pending ACKs associated with the first flow (Operation 344). The system (i.e., the network device acting as the second network ingress device, e.g.) FIG. 3D The ingress network device 220 (operating network device) can track the number of pending ACKs received in response to the transmission of the first-order data packet. Alternatively, the system can wait until all packets in the downstream flow have been cleared, as indicated by a returned ACK, which represents the amount or quantity of data in the flow, rather than the number of packets required to transmit that data. The system may or may not have a one-to-one mapping of returned ACKs to transmitted packets. The predetermined number of pending ACKs can be configured to account for packet loss and can be a specific number or percentage. For example, the ingress network device 220 can wait until at least twenty (or 80% or another threshold) pending ACKs associated with the first data flow (via path 260) are received or have been returned, indicating that the data associated with the pending ACKs has been successfully transmitted to or by the egress network device. In some respects, the ingress network device 220 can wait until almost all or all pending ACKs have been received.
[0046] If no predetermined number of pending ACKs are received (Decision 346), the operation returns to Operation 344. If the predetermined number of pending ACKs are received (Decision 346), the system determines whether the (same) first path has been offered as a new path for the first flow that may be rerouted for suspension more than a predetermined number of times (Decision 348).
[0047] If the (same) first path is not offered as a new path more than a predetermined number of times (e.g., 5 times) (Decision 348), then the operation is performed in... FIG. 2Continue at label C. If the (same) first path is offered as a new path more than a predetermined number of times (e.g., 5 times) (Decision 348), the system releases the first flow to continue routing on the first path (Operation 350). After releasing the first flow to continue routing on the (same) first path, the system avoids rerouting the first flow on new paths (Operation 352). For example, in FIG. 3D In the process, if the network device does not provide the same first path (path 260) as a new path more than five times, then the operation is in... FIG. 3D Continue at label C (i.e., reroute the first flow to a different new path). If the network device provides the same first path (path 260) as a reroute or new path more than five times, the network device can (by tracking the provided paths and the number of times the path is provided) determine to release the first flow to continue routing on the original path (path 260) (i.e., the first path 260 on which the packets received by network device 222 triggered the initially detected midpoint congestion (206)), and the network device can avoid rerouting the flow (originally routed via path 260) on a new path (via path 270).
[0048] FIG. 2 A flowchart 360 illustrates a method for optimizing the selection of flows to be rerouted, according to one aspect of this application, including a method for rerouting flows. During operation, the system reroutes the first flow to a new path via a network device operating as a second ingress network device (operation 362). For example, ingress network device 220 can reroute the first flow (via path 260) to a new path (via path 270 as shown by the dashed line), as described above regarding... FIG. 2 As described.
[0049] The system stores rerouted first-order entries in a data structure through a network device operating as a second entry network device. These entries include a first load metric (operation 364). The network device may store rerouted first-order entries, including identification information of the original flow (e.g., via path 260), identification information of the new or rerouted path (e.g., via path 270), and first load metric information related to the first flow, determined or generated by the network device.
[0050] The system receives a second redirect ACK corresponding to the first rerouted flow, wherein the second redirect ACK includes a second load metric (Operation 366). For example, although in FIG. 2Not depicted, but the ingress network device 220 may receive another redirection ACK (second redirection ACK) from another intermediate network device (e.g., intermediate network device 234). The second redirection ACK may also include identification information of its corresponding original flow (second flow), identification information of the new or rerouted path, and second load metric information related to the second flow determined or generated by the intermediate network device 234.
[0051] The system stores the second load metric in the first-level entry of the rerouting (Operation 368). The data structure can be a table, list, array, or other way of storing data and associated information. Therefore, continue... FIG. 4 In an example where the ingress network device 220 receives both the first and second redirect ACKs and stores the associated information, the ingress network device 220 may store the second load metric in the same entry as the first load metric.
[0052] The system calculates the difference between the second load metric included in the second redirect ACK and the first load metric included in the first redirect ACK (operation 370). The network device operating as the second ingress network device can maintain a data structure and can also perform the difference calculation and store it in the data structure entry of the first flow of the rerouting. The difference between the first load metric and the second load metric can be represented by, for example: the difference between ECA values; the difference between bandwidth consumption; the difference between the number of pending bytes; and the difference in the calculation or measurement method based on the first load metric and the second load metric.
[0053] The system adjusts the probability of selecting the first flow for rerouting based on the difference (Operation 372). For example, a small difference (e.g., a measurement difference of less than 3%) may indicate that congestion is not significantly improved when using a new or rerouting path. Therefore, the first flow can be marked as a strong candidate for rerouting, i.e., the network device can increase the probability that the first flow is selected for rerouting. Conversely, a large difference (e.g., a measurement difference greater than 60%) may indicate that congestion has been significantly improved when using a new or rerouting path. Therefore, rerouting the first flow may not be beneficial, and the network device can mark the first flow as a weak candidate for rerouting. The labeling of "strong" or "weak" candidates is provided for illustrative purposes only. Other categories or types can be used, including levels, ranges or windows of values, and a limited or constrained number of categories to be assigned to each candidate in the group of received flows.
[0054] Therefore, by allowing the midpoint network device to generate metrics and send redirection ACKs under certain circumstances, and by allowing the ingress network device to receive multiple redirection ACKs and make decisions about rerouting flows based on various rerouting conditions (as described herein), the described aspects provide a system that can optimize the selection of flows to be rerouted based on congestion detected by the midpoint network device (midpoint congestion) and congestion managed by the ingress network device. Optimizing flow selection results in improved performance and a more efficient overall system.
[0055] FIG. 4 The illustration depicts a computer system 400 that facilitates the optimization of the selection of flows to be rerouted, according to one aspect of this application. The computer system 400 includes a processor 402, a memory 404, and a storage device 406. The memory 404 may include volatile memory (e.g., random access memory (RAM)) that serves as managed memory and can be used to store one or more memory pools. Furthermore, the computer system 400 may be coupled to peripheral I / O user equipment 410 (e.g., a display device 411, a keyboard 412, and a pointing device 413). The storage device 406 includes a non-transitory computer-readable storage medium and stores an operating system 416, a congestion detection subsystem / instructions 420, a congestion management subsystem / instructions 430, and data 442. The computer system 400 may include a processor 402, a memory 404, and a storage device 406. FIG. 1 The entities or instructions shown are fewer or more entities or instructions.
[0056] Instruction 420 may include instructions 422 and 424, which, when executed by computer system 400, cause computer system 400 to perform the methods and / or processes described in this disclosure, for example, including computer system 400 operating as an intermediate network device. Specifically, computer system 400 may store instruction 422 to generate load metrics for the corresponding flows in the first set of received flows, as described above regarding, for example... FIG. 3A Switch 116 and FIG. 1 As described in operation 304.
[0057] Computer system 400 may store instruction 424 to send a redirection ACK, including the load metric of the corresponding flow, to the ingress network device in response to a load metric exceeding the load value, as described above. FIG. 3A Switches 118 and 120 and FIG. 3B As described in operation 308.
[0058] Instruction 430 may also include instructions 432, 434, 436, 438, and 440, which, when executed by computer system 400, cause computer system 400 to perform the methods and / or processes described in this disclosure, for example, including computer system 400 operating as an ingress network device. Specifically, computer system 400 may store instruction 432 to forward a second set of flows, as described above, for example, with respect to ingress network device 220 and the forwarding flow. FIG. 2 The operation described in 332.
[0059] Computer system 400 may further store instructions 434 to receive from multiple intermediate network devices multiple redirection ACKs corresponding to multiple flows in the second set of flows, the corresponding redirection ACKs including load metrics of the corresponding flows in the multiple flows. (The above is about...) FIG. 3B The ingress network device 220 and FIG. 2 Operation 334 describes receiving multiple redirect ACKs, each including a load metric for the corresponding stream.
[0060] Computer system 400 may store instructions 436 to select a first flow for rerouting from the plurality of flows based on a set of rerouting conditions, the first flow being associated with a first path and corresponding to a first redirection ACK including a first load metric. The selection of a flow for rerouting may be based on a determined probability or a set of rerouting conditions, as described above. FIG. 3B The ingress network device 220 and FIG. 3C Operation 336 is described.
[0061] Computer system 400 may store instruction 438 to reroute the first flow to a new path. Rerouting the first flow may occur after pausing the flow and waiting until a predetermined number of pending ACKs have been received, or after determining that the same first path has been offered a certain number of times (relative to the predetermined number), as described above regarding... FIG. 3D Operations 340 to 348 and FIG. 3D The operation described in 352.
[0062] Computer system 400 can store instructions 440 to store the first-order entry for rerouting in a data structure, the entry including a first load metric, as described above. FIG. 4 The operation described in 364.
[0063] Instructions 420 and 430 may include more than FIG. 1 The instructions shown are further instructions. For example, instructions 420 and 430 may include instructions for performing the operations described above with respect to the following: FIG. 2 and FIG. 3A to FIG. 3D In the environment; FIG. 5The operations depicted in the flowchart; and FIG. 5 Instructions 510-522 of CRM 500.
[0064] Data 442 may include any data required as input or generated as output by the methods, operations, communications, and / or processes described in this disclosure. Specifically, data 442 may store at least: load metrics; streams; data of the streams; load values; predetermined values; the result of comparing load metrics with load values; redirection ACKs; redirection ACKs corresponding to streams and including load metrics; multiple streams; selected streams; a first path; the original path; the same path; a new path; a path used to reroute streams; a data structure; entries in the data structure; differences between load metrics; the probability of selecting a stream for rerouting; adjusted probabilities; conditions; rerouting conditions; a time quantity; a data quantity; the result of comparing load metrics; the differences between load metrics; the order of arrangement; the current load; the size of the data packets; the product of load and data packet size; bandwidth consumption; the amount of data pending in the output or input buffers; information received from the NIC or associated with the state of the streams; a decision on whether to send a redirection ACK; and predetermined or pre-configured thresholds.
[0065] FIG. 1 The illustration depicts a computer-readable medium (CRM) 500 that, according to one aspect of this application, facilitates the selection of flows to be rerouted. The CRM 500 may be a non-transitory computer-readable medium or device storing instructions that, when executed by a computer or processor, cause the computer or processor to perform a method. The CRM 500 may store instructions 510 to generate load metrics for corresponding flows in a first set of received flows, as described above regarding, for example... FIG. 3A Switches 116 and 118 and FIG. 1 As described in operation 304.
[0066] The CRM 500 can store instruction 512 to transmit a redirect ACK, including the load metric of the corresponding stream, in response to a load metric exceeding the load value, as described above. FIG. 3A Switches 118 and 120 and FIG. 2 As described in operation 308.
[0067] CRM 500 can store instruction 514 to forward a second set of streams, as shown above for example regarding... FIG. 3B The ingress network device 220 and the forwarding flow in the middle FIG. 2 The operation described in 332.
[0068] The CRM 500 can store instructions 516 to receive multiple redirection ACKs from multiple intermediate network devices corresponding to multiple flows in a second set of flows, wherein the corresponding redirection ACK includes the load metric of the corresponding flow in the multiple flows. (The above is about...) FIG. 3B The ingress network device 220 and FIG. 2 Operation 334 describes receiving multiple redirect ACKs, each including a load metric for the corresponding stream.
[0069] CRM 500 can store instructions 518 to select a first flow to be rerouted from the plurality of flows based on a set of rerouting conditions, wherein the first flow is associated with a first path and corresponds to a first redirection ACK including a first load metric. The selection of the first flow to be rerouted (i.e., the "candidate flow") can be based on determining the probability of each flow or based on one or more rerouting conditions, including those mentioned above. FIG. 3B The ingress network device 220 and FIG. 2 The operations provided in 336 are examples.
[0070] CRM 500 can store instruction 520 to reroute the first stream to a new path, as mentioned above. FIG. 3B (The paths rerouted via 271 to 285 are depicted, as indicated by the dashed lines) ingress network device 220 and FIG. 3D Operation 336 is described.
[0071] CRM 500 can store instruction 522 to store the first-order entry for rerouting in a data structure, where the entry includes the first load metric, as described above. FIG. 5 The operation described in 364.
[0072] CRM 500 can include more FIG. 1 The instructions shown are more than just those. For example, the CRM 500 can also store instructions for performing the operations described above regarding the following: FIG. 2 and FIG. 3A to FIG. 3D In the environment; FIG. 4 The operations depicted in the flowchart; and FIG. 1 Instructions 420 and 430 of computer system 400.
[0073] The term "network device" refers to any device, component, or computing entity that can provide a communication channel for data packets sent from a "processing node" or "endpoint node." A processing node or endpoint node can refer to a device, component, or hardware component that can operate as a source or destination of data, including, for example, control packets or data packets. Network devices can include ingress network devices, intermediate or midpoint network devices, or egress or endpoint network devices. An example of a network device can be a switch, as described above regarding... FIG. 1 As described herein, a processing node or endpoint node may include an ingress node (which is the endpoint from which data is returned from a request) or an egress node (which is the endpoint from which data is sent from a request). Additionally, a network device may operate or perform the functions of an ingress network device, intermediate network device, or egress network device as described herein.
[0074] In general, the disclosed aspects provide a computing system, method, and computer-readable medium that facilitates the optimization of the selection of flows to be rerouted. The computing system operates within a network architecture including ingress network devices, intermediate network devices, and egress network devices. The computing system includes a processor and a storage device storing congestion detection instructions and congestion management instructions (also referred to as subsystems), which perform the following actions when executed by the processor: The congestion detection subsystem may include instructions for: generating a load metric for a corresponding flow in a first set of received flows; and sending a redirection acknowledgment (ACK) including the load metric of the corresponding flow to the ingress network device in response to a load metric exceeding a load value. The congestion management subsystem may include instructions for: receiving a first redirection ACK corresponding to a first flow; and determining whether to select the first flow for rerouting based on a set of rerouting conditions. The computing system may further include instructions for performing the operations described herein, including those related to: FIG. 2 and FIG. 3A to FIG. 3D In the environment; FIG. 5 The operations depicted in the flowchart; and FIG. 1 The instructions in CRM 500.
[0075] In a variation of this, the congestion management instruction is further configured to: forward a second set of flows including a first flow, wherein the first flow is associated with a first path, and wherein a first redirection ACK indicates a first load metric; receive from a plurality of intermediate network devices a plurality of redirection ACKs corresponding to a plurality of flows in the second set of flows, wherein the plurality of redirection ACKs include a first redirection ACK, and wherein each redirection ACK includes a load metric of the corresponding flow in the plurality of flows; determine, based on the set of rerouting conditions, to select a first flow from the plurality of flows for rerouting; and reroute the first flow to a new path. The congestion management instruction is further configured to: store an entry for the rerouting first flow in a data structure, the entry including the first load metric; receive a second redirection ACK corresponding to the rerouting first flow, the second redirection ACK including a second load metric; and store the second load metric in the entry for the rerouting first flow.
[0076] In a further variation, the congestion management command is further used to determine the difference between the second load metric included in the second redirect ACK and the first load metric included in the first redirect ACK. The congestion management command is further used to adjust the probability of selecting the first flow for rerouting based on this difference.
[0077] In another variation of this aspect, the set of rerouting conditions is associated with the probability that a corresponding flow from the plurality of flows is selected for rerouting.
[0078] In a further variation, the set of rerouting conditions includes at least one of the following: the amount of time elapsed since the most recently rerouted flow; the amount of pending data to be sent in the corresponding flow of the plurality of flows; a comparison of the load metric of the corresponding flow of the plurality of flows with the load metrics of other flows of the plurality of flows; in some cases, the difference between the load metrics included in the received redirection ACK corresponding to the same flow; or the order in which the plurality of flows are arranged.
[0079] In a further variation, the load metric generated for the corresponding flow in the first group of flows in the congestion detection subsystem and the load metric for the corresponding flow in the plurality of flows in the congestion management subsystem are based on at least one of the following: the load associated with the congestion detection subsystem or the congestion management subsystem, which is expressed as an explicit congestion avoidance (ECA) value; or the size of packets in the corresponding flow in the first group of flows or in the corresponding flow in the plurality of flows.
[0080] In a further variation, the load metric generated by the corresponding flow in the first group of flows in the congestion detection subsystem and the load metric of the corresponding flow in the multiple flows in the congestion management subsystem include: the product of the load and the packet size of the corresponding flow in the congestion detection subsystem or the congestion management subsystem.
[0081] In a further variation, the load metric generated by the corresponding flow in the first group of flows in the congestion detection subsystem and the load metric of the corresponding flow in the plurality of flows in the congestion management subsystem are based on at least one of the following: bandwidth consumption associated with the congestion detection subsystem or the congestion management subsystem; the amount of data pending in the input buffer associated with the congestion detection subsystem or the congestion management subsystem; information received from the network interface controller (NIC) and associated with the amount of data pending to be processed by the congestion detection subsystem or the congestion management subsystem; or information associated with the state of the corresponding flow in the first group of flows in the congestion detection subsystem or the corresponding flow in the plurality of flows in the congestion management subsystem.
[0082] In a further variation, the congestion management command suspends the first flow before it is rerouted to a new path. The congestion management command is further configured to wait until at least a predetermined number of pending ACKs associated with the first flow are received. The congestion management command will then be further configured to, in response to waiting until the predetermined number of pending ACKs are received and in response to the first path being offered more than a predetermined number of times: release the first flow to allow it to continue routing on the first path; and avoid rerouting the first path.
[0083] In a further variation, the congestion detection command is used to avoid sending a redirect ACK to the ingress network device in response to a load metric being less than the load value.
[0084] In a further variation, the congestion detection instruction is further used to compare the load metric with the load value in response to a load metric being greater than a predetermined threshold.
[0085] In a further variation, the load value includes randomly generated numbers.
[0086] On the other hand, a computer-implemented method may include various operations performed by, for example, a system. The system generates a load metric for a corresponding flow in a first set of received flows via a network device operating as a first intermediate network device in a network architecture. In response to a load metric being greater than a load value, the system sends a redirection acknowledgment (ACK) including the load metric of the corresponding flow to a first ingress network device associated with the corresponding flow. In response to a load metric being less than a load value, the system avoids sending a redirection ACK to the first ingress network device. The system forwards a second set of flows via a network device operating as a second ingress network device in the network architecture. The system receives multiple redirection ACKs from multiple intermediate network devices corresponding to multiple flows in the second set of flows, wherein the corresponding redirection ACK includes the load metric of the corresponding flow in the multiple flows. The system selects a first flow from the multiple flows for rerouting based on a set of rerouting conditions, wherein the first flow is associated with a first path and corresponds to a first redirection ACK including the first load metric. The system reroutes the first flow to a new path. The method may include additional operations, including operations related to: FIG. 2 and FIG. 3A to FIG. 3D In the environment; FIG. 4 The operations depicted in the flowchart; FIG. 5 Instructions 420 and 430 of the computing system 400; and FIG. 1 Instructions 510-522 of CRM 500.
[0087] On the other hand, a non-transitory computer-readable storage medium (or CRM) stores instructions for generating load metrics for corresponding flows in a first set of received flows. These instructions are further configured to transmit a redirection acknowledgment (ACK) including the load metric of the corresponding flow in response to a load metric exceeding a load value. These instructions are further configured to forward a second set of flows. These instructions are further configured to receive multiple redirection ACKs corresponding to multiple flows in the second set of flows from multiple intermediate network devices, wherein the corresponding redirection ACK includes the load metric of the corresponding flow in the multiple flows. These instructions are further configured to select a first flow for rerouting from the multiple flows based on a set of rerouting conditions, wherein the first flow is associated with a first path and corresponds to a first redirection ACK including the first load metric. These instructions are further configured to reroute the first flow to a new path. These instructions are further configured to store an entry for the rerouted first flow in a data structure, wherein the entry includes the first load metric. The CRM may also store instructions for performing the operations described above with respect to the following: FIG. 2 and FIG. 3A to FIG. 3D In the environment; FIG. 4 The operations depicted in the flowchart; FIG. 5 Instructions 420 and 430 of computer system 400; and Instructions 510-522 of CRM 500.
[0088] The foregoing description is presented to enable any person skilled in the art to make and use the aspects and examples, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed aspects will be apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects and applications without departing from the spirit and scope of this disclosure. Therefore, the aspects described herein are not limited to those shown, but are intended to be consistent with the maximum scope of the principles and features disclosed herein.
[0089] Furthermore, the foregoing descriptions of the various aspects have been presented solely for illustrative and descriptive purposes. These descriptions are not intended to be exhaustive or to limit the aspects described herein to the disclosed forms. Accordingly, many modifications and variations will be apparent to those skilled in the art. Additionally, the foregoing disclosure is not intended to limit the aspects described herein. The scope of the aspects described herein is defined by the appended claims.
Claims
1. A computing system operating in a network structure including an ingress network device and intermediate network devices, the computing system comprising: Congestion detection subsystem and congestion management subsystem; The congestion detection subsystem is used for: Generate the load metric for the corresponding stream in the first set of received streams; as well as In response to the load metric being greater than the load value, a redirection acknowledgment (ACK) including the load metric of the corresponding flow is sent to the ingress network device; and The congestion management subsystem is used for: Receive the first redirect ACK corresponding to the first stream; and Whether to select the first flow for rerouting is determined based on a set of rerouting conditions.
2. The computing system as described in claim 1, wherein, The congestion management subsystem is further used for: The forwarding includes a second set of flows of the first flow, wherein the first flow is associated with a first path, and wherein the first redirect ACK indicates a first load metric; Receive multiple redirection ACKs corresponding to multiple flows in the second group of flows from multiple intermediate network devices, wherein the multiple redirection ACKs include the first redirection ACK, and wherein the corresponding redirection ACK includes the load metric of the corresponding flow in the multiple flows; Based on the set of rerouting conditions, the first flow is selected from the plurality of flows for rerouting; Reroute the first stream to the new path; The first-order entry for rerouting is stored in a data structure, the entry including the first load metric; Receive a second redirect ACK corresponding to the first rerouted flow, the second redirect ACK including a second load metric; and The second load metric is stored in the first-order entry of the rerouting.
3. The computing system as described in claim 2, wherein, The congestion management subsystem is further used for: Determine the difference between the second load metric included in the second redirect ACK and the first load metric included in the first redirect ACK; and The probability of selecting the first flow for rerouting is adjusted based on the aforementioned differences.
4. The computing system as described in claim 2, in, The set of rerouting conditions is associated with the probability that a corresponding flow from the plurality of flows is selected for rerouting.
5. The computing system as described in claim 2, wherein, The set of rerouting conditions includes at least one of the following: The amount of time that has elapsed since the most recently rerouted flow; The amount of data to be sent pending in the respective streams of the plurality of streams; A comparison of the load metric of the corresponding flow in the plurality of flows with the load metrics of the other flows in the plurality of flows; In some cases, the difference between the load metrics included in the received redirect ACK corresponding to the same stream; or The order in which the multiple streams are arranged.
6. The computing system as claimed in claim 2, wherein, The load metric generated by the corresponding flow in the first group of flows in the congestion detection subsystem and the load metric of the corresponding flow in the plurality of flows in the congestion management subsystem are based on at least one of the following: The load associated with the congestion detection subsystem or the congestion management subsystem, the load being represented as an explicit congestion avoidance (ECA) value; or The size of the data packets in the corresponding stream of the first group of streams or in the corresponding stream of the plurality of streams.
7. The computing system of claim 6, wherein, The load metrics generated by the corresponding flows in the first group of flows in the congestion detection subsystem and the load metrics of the corresponding flows in the plurality of flows in the congestion management subsystem include: The load is the product of the packet size of the corresponding flow in the congestion detection subsystem or the congestion management subsystem.
8. The computing system as claimed in claim 2, wherein, The load metric generated by the corresponding flow in the first group of flows in the congestion detection subsystem and the load metric of the corresponding flow in the plurality of flows in the congestion management subsystem are based on at least one of the following: Bandwidth consumption associated with the congestion detection subsystem or the congestion management subsystem; The amount of data pending in the input buffer associated with the congestion detection subsystem or the congestion management subsystem; Information received from the network interface controller (NIC) and associated with the amount of data pending processing by the congestion detection subsystem or the congestion management subsystem; or Information associated with the state of the corresponding flow in the first group of flows in the congestion detection subsystem or the corresponding flow in the plurality of flows in the congestion management subsystem.
9. The computing system as claimed in claim 2, wherein, The congestion management subsystem is further used for: Pause the first flow before rerouting it to the new path; Wait until at least a predetermined number of pending ACKs associated with the first stream are received; as well as In response to waiting until the predetermined number of pending ACKs are received and in response to the first path being provided more than the predetermined number of times: Release the first flow so that the first flow can continue to be routed on the first path; as well as Avoid rerouting the first path.
10. The computing system of claim 1, wherein, The congestion detection subsystem is further used for: In response to the load metric being less than the load value, the redirection ACK is avoided from being sent to the ingress network device.
11. The computing system of claim 1, wherein, The congestion detection subsystem is further used for: In response to the load metric being greater than a predetermined threshold, the load metric is compared with the load value.
12. The computing system of claim 1, wherein, The load value includes randomly generated numbers.
13. A computer-implemented method, comprising: The network device, operating as the first intermediate network device in the network structure, generates the load metric for the corresponding flow in the first set of received flows. In response to the load metric being greater than the load value, a redirection acknowledgment (ACK) including the load metric of the corresponding flow is sent to the first ingress network device associated with the corresponding flow; In response to the load metric being less than the load value, the redirection ACK is avoided from being sent to the first ingress network device; Receive the first redirect ACK corresponding to the first stream; as well as Whether to select the first flow for rerouting is determined based on a set of rerouting conditions.
14. The computer-implemented method of claim 13, further comprising: The forwarding includes a second set of flows of the first flow, wherein the first flow is associated with a first path, and wherein the first redirect ACK indicates a first load metric; Receive multiple redirection ACKs corresponding to multiple flows in the second group of flows from multiple intermediate network devices, wherein the multiple redirection ACKs include the first redirection ACK, and wherein the corresponding redirection ACK includes the load metric of the corresponding flow in the multiple flows; Determine whether to select the first flow for rerouting from the plurality of flows based on the set of rerouting conditions; and The first stream is rerouted to the new path.
15. The computer-implemented method of claim 14, further comprising: The network device, operating as a second entry network device, stores the first-order rerouting entries in a data structure. The entry includes the first load metric; Receive a second redirect ACK corresponding to the first rerouted flow. The second redirect ACK includes a second load metric; Store the second load metric in the first-order entry of the rerouting; Calculate the difference between the second load metric included in the second redirect ACK and the first load metric included in the first redirect ACK; and The probability of selecting the first flow for rerouting is adjusted based on the aforementioned differences.
16. The computer-implemented method as described in claim 14, in, The set of rerouting conditions is associated with the probability that a corresponding flow from the plurality of flows is selected for rerouting, and The set of rerouting conditions includes at least one of the following: The amount of time that has elapsed since the most recently rerouted flow; The amount of data to be sent pending in the respective streams of the plurality of streams; A comparison of the load metric of the corresponding flow in the plurality of flows with the load metrics of the other flows in the plurality of flows; In some cases, the difference between the load metrics included in the received redirect ACK corresponding to the same stream; or Includes an ordered list of the multiple streams.
17. The computer-implemented method as described in claim 14, in, The load metric generated by the corresponding flow in the first group of flows and the load metric of the corresponding flow in the plurality of flows are based on at least one of the following: The load associated with the corresponding flow in the first group of flows or the corresponding flow in the plurality of flows, the load being represented as an explicit congestion avoidance (ECA) value; or The size of the data packets in the corresponding stream of the first group of streams or in the corresponding stream of the plurality of streams.
18. The computer-implemented method as described in claim 14, in, The load metric generated by the corresponding flow in the first group of flows and the load metric of the corresponding flow in the plurality of flows are based on at least one of the following: Bandwidth consumption associated with the network device operating as the first intermediate network device or as the second entry network device; The amount of data pending in the input buffer associated with the network device operating as the first intermediate network device or as the second ingress network device; Information received from the network interface controller (NIC) and associated with the amount of data pending to be processed by the network device operating as the first intermediate network device or as the second ingress network device; or Information associated with the state of the corresponding flow in the first group of flows or the corresponding flow in the plurality of flows.
19. The computer-implemented method of claim 14, further comprising: The first flow is paused by the network device operating as the second ingress network device before it is rerouted to the new path; Wait until at least a predetermined number of pending ACKs associated with the first stream are received; as well as In response to waiting until the predetermined number of pending ACKs are received and in response to the first path being provided more than the predetermined number of times: Release the first flow so that the first flow can continue to be routed on the first path; as well as Avoid rerouting the first path.
20. A non-transitory computer-readable medium storing instructions, the instructions being used to: Generate the load metric for the corresponding stream in the first set of received streams; In response to the load metric being greater than the load value, a redirection acknowledgment (ACK) including the load metric of the corresponding stream is transmitted; Forward the second stream; Receive multiple redirection ACKs from multiple intermediate network devices, corresponding to multiple flows in the second group of flows. The corresponding redirection ACK includes the load metric of the corresponding flow among the plurality of flows; Based on a set of rerouting conditions, a first flow is selected from the plurality of flows for rerouting. Wherein, the first flow is associated with the first path and corresponds to the first redirect ACK, which includes the first load metric; Reroute the first stream to the new path; and Store the first-order entries for rerouting in a data structure. The entry includes the first load metric.