Network congestion control method and related device

By configuring multiple paths at the source end of the data center network and dynamically adjusting the credit and load, the problems of uneven load and network congestion in the data center network are solved, and high throughput and low congestion network performance is achieved.

CN115733799BActive Publication Date: 2025-06-13XFUSION DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110980604.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-25
Publication Date
2025-06-13
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

There are problems with uneven load and network congestion in the data center network, resulting in low overall throughput, long dynamic delay and long link failure recovery time, which cannot meet the needs of complex applications.

Method used

By configuring multiple paths at the source and assigning credit to each path, dynamically adjusting the credit and load of the path according to the feedback information, fine-grained load balancing and network congestion control are achieved.

Benefits of technology

The optimal load ratio between multiple paths is achieved, the short-board effect caused by congestion of specific paths is eliminated, the overall throughput rate is improved, and network congestion is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115733799B_ABST
    Figure CN115733799B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a network congestion control method and related device. The method includes: a network device sends multiple packets of a first data stream through multiple paths, and a credit amount is configured for each of the multiple paths, where the credit amount indicates the size of the data sending capacity of each path, and the sum of the credit amounts of the multiple paths is less than or equal to the congestion threshold maintained by the network device for the first data stream; the network device receives first feedback information including indication information on whether a first path is congested, the first feedback information is the feedback information of a first packet among the multiple packets, and the first path is the path used to send the first packet among the multiple paths; the network device re-determines the credit amount of the first path based on the first feedback information; the network device re-determines the load amount of the first path based on the credit amount of the first path. The present application can better achieve load balancing and network congestion control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a network congestion control method and related devices. Background Art

[0002] Currently, the scale of data centers continues to increase, and the services they carry are becoming more and more complex. The data center network (DCN) is under increasing pressure. In addition to high throughput requirements, many services also have relatively strict requirements for static latency and long-tail latency. However, in the current DCN, there are problems such as uneven load among multiple nodes resulting in low overall throughput, network congestion causing excessive dynamic latency, and long link failure recovery time, which have continuously troubled related equipment suppliers and operators.

[0003] Currently, Fat Tree networking is widely used in data centers, and mechanisms such as equal-cost multi-path Routing (ECMP) technology and congestion control algorithms are introduced on this basis. To a certain extent, it alleviates uneven load and network congestion. However, due to problems such as hash conflicts and poor response to bursty traffic, it performs poorly in large-scale networking and can no longer meet the requirements of complex applications based on DCN. Summary of the Invention

[0004] Embodiments of this application disclose a network congestion control method and related devices, which can better achieve load balancing of network traffic and network congestion control.

[0005] In a first aspect, this application provides a network congestion control method, which includes:

[0006] A source end sends multiple packets of a first data stream through n paths, and each of the n paths is configured with a credit amount, where the credit amount indicates the size of the capacity for each path to send data. The sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, where the first congestion threshold is the congestion threshold of the first data stream, and n is an integer greater than 1;

[0007] The source end receives first feedback information, where the first feedback information is the feedback information of a first packet, and the first feedback information includes indication information on whether the first path is congested. The first path is the path used to send the first packet among the n paths; the first packet is any one of the multiple packets of the first data stream;

[0008] The source end re-determines the credit amount of the first path based on the indication information;

[0009] The source end re - determines the load of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re - determined load of the first path decreases.

[0010] The above - mentioned source end can be a network device, or can be an intelligent network card, an on - board network card, a field programmable gate array (FPGA) with a network interface, an acceleration card, etc. in the network device.

[0011] In the embodiments of the present application, by distributing the load of a single flow to multiple paths for transmission, fine - grained load balancing is achieved, bandwidth competition at the flow level is reduced, and the load - balancing performance of the entire network is improved. In addition, by maintaining a congestion threshold for a single flow, distributing the credit amount included in the congestion threshold to multiple transmission paths of the flow, and adaptively adjusting the credit amount of the corresponding path based on the feedback information of the packets transmitted on each path, and then adjusting the load amount, that is, by combining flow - level congestion control and path - level credit management, the perception of the congestion status of multiple paths is realized, and the dynamic load of the paths is adjusted according to the congestion status of the paths, achieving the optimal load ratio between multiple paths, eliminating the short - board effect caused by congestion of a specific path, obtaining an increase in the overall throughput rate, and reducing network congestion.

[0012] In a possible implementation manner, the above - mentioned first feedback information is the information in a feedback packet received by the source end. The feedback packet includes the feedback information of m packets among the multiple packets, and m is an integer greater than 1;

[0013] The feedback information of the m packets includes indication information on whether the transmission path of each packet is congested during the transmission of the m packets.

[0014] In the embodiments of the present application, by aggregating the feedback information of multiple packets into a feedback packet for transmission, bandwidth resources can be saved. In addition, by carrying the indication information on whether the transmission paths of the multiple packets are congested in the feedback packet together, the above - mentioned network device can know the congestion situation of the transmission path of each packet, so as to adjust the credit amount and load amount of the corresponding path to better achieve network load balancing and congestion control.

[0015] In a possible implementation manner, the source end re - determines the credit amount of the first path based on the above - mentioned indication information, including:

[0016] The source end calculates a second congestion threshold based on the above - mentioned first feedback information, and the second congestion threshold is the new congestion threshold of the first data stream;

[0017] The source end adjusts the credit amount of the first path based on a first difference, where the first difference is the difference between the second congestion threshold and the congestion threshold of the first data stream before calculating the second congestion threshold.

[0018] In a possible implementation, the sum of the credit amounts and the remaining credit amounts of the above n paths is equal to the above first congestion threshold; the source end adjusts the credit amount of the first path based on the first difference, including:

[0019] When the first difference is greater than zero and the sum of the first difference and the remaining credit amount is greater than the target credit amount, the source end increases the credit amount of the first path by two target credit amounts, where the target credit amount indicates the data volume size of a packet in the first data stream; or,

[0020] When the first difference is less than zero, the source end reduces the credit amount of the first path by a first credit amount, where the first credit amount is the absolute value of the sum of the first difference and the target credit amount.

[0021] In the embodiments of the present application, after the above network device receives the feedback information of a packet of a certain flow, it can recalculate the congestion threshold of the flow. If the transmission path of the packet does not experience congestion, the new congestion threshold increases, and then the credit amount of the transmission path of the packet also increases correspondingly, so that the bandwidth of the path can be effectively utilized; if the transmission path of the packet experiences congestion, the new congestion threshold decreases, and then the credit amount of the transmission path of the packet also decreases correspondingly, so that the congestion of the path can be effectively alleviated. In addition, regardless of how the new congestion threshold changes relative to the original congestion threshold, as long as the feedback information of a packet is received, the destination end will fill back a target credit amount in the credit amount of the transmission path corresponding to the packet, that is, increase a target credit amount for sending a new packet.

[0022] In a possible implementation, the above first feedback information further includes indication information that the first packet is out-of-order at the destination end; each time the first path sends a packet, it consumes a target credit amount of the credit amount in the first path, and the target credit amount indicates the data volume size of a packet in the first data stream;

[0023] The source end re-determines the credit amount of the first path based on the first feedback information, including:

[0024] The source end maintains the current credit amount of the first path based on the indication information that the first packet is out-of-order at the destination end.

[0025] In an embodiment of the present application, a message disorder perception mechanism is introduced at the destination end to identify the messages of the above-mentioned first data stream that are seriously disordered due to path congestion, and feedback is given to the source end. After the source end receives the feedback indicating the disorder, the credit of the corresponding path can be maintained unchanged. Since each message sent consumes a target credit, if the messages received by the destination end from the path continuously have disorder exceeding the threshold, then the credit of the path in the source end will gradually decrease until it becomes zero. After the credit of the path becomes zero, the source end will no longer send messages of the first data stream through the path, thereby alleviating or even avoiding the situation where the destination end has serious disorder causing message loss, thereby reducing the number of retransmitted messages and saving bandwidth resources. In addition, the messages transmitted in the path continuously have disorder exceeding the threshold, indicating that the delay of the path is large and congestion has occurred. The present application further effectively alleviates the congestion of the path by reducing the number of messages sent on the path or even not sending messages. In addition, the present application adopts a mechanism for tracking the degree of disorder at the destination end. Compared with tracking the degree of disorder at the source end, this reduces the interference of the delay in the transmission of feedback information to the source end, and the disorder indication information obtained by this mechanism is more accurate.

[0026] In a possible implementation manner, the sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold; and the method further includes:

[0027] When the credit amount of the first path drops to zero, the source sends a detection message through the first path;

[0028] The source end receives second feedback information of the detection message; when the second feedback information indicates that the first path is not congested, the source end calculates a third congestion threshold based on the second feedback information, and the third congestion threshold is a new congestion threshold of the first data flow;

[0029] When the sum of the second difference and the remaining credit is greater than the target credit, the source end increases the credit of the first path by the target credit; the second difference is the difference between the third congestion threshold and the congestion threshold of the first data flow before calculating the third congestion threshold, and the target credit indicates the data amount of a message in the first data flow.

[0030] In the embodiment of the present application, by detecting the congestion of the path through the detection message, the available path can be restored in time, the load sharing capability can be improved, and the message of sending the first data stream is avoided from being degraded to a single path due to congestion. In addition, the load (payload) in the above detection message can be empty or 0, so that the sending of the detection message has almost no effect on the network bandwidth, and the sending frequency of the detection message is low, which will not aggravate the congestion state of the path.

[0031] In a possible implementation, the method further includes:

[0032] The source end receives third feedback information, where the third feedback information is the feedback information of the second packet among the multiple packets. The third feedback information includes information indicating that the second packet is lost, and the second packet is sent through the second path among the n paths;

[0033] The source end increases the credit amount of the second path by a target credit amount, and the target credit amount indicates the data amount size of a packet in the first data stream.

[0034] The second path and the first path above may be the same path or different paths.

[0035] In the embodiments of the present application, for a packet that is not correctly received at the destination end, that is, a lost packet, the destination end can feedback to the source end the indication information that the packet is lost. Since the destination end has consumed a target credit amount on the corresponding transmission path when sending the lost packet, after the source end receives this feedback information, it can fill back a target credit amount on this transmission path so that the credit amounts of the respective paths corresponding to the data stream match the congestion threshold of the data stream.

[0036] In a possible implementation, the source end receives the feedback information of a packet with a target sequence number, where the target sequence number includes all the sequence numbers of the packets sent through the n paths. The method further includes:

[0037] When the congestion threshold of the first data stream is greater than the actual credit amount, the source end adjusts the congestion threshold of the first data stream to a value equal to the actual credit amount, and the actual credit amount is the sum of the credit amounts of the n paths and the remaining credit amount.

[0038] In the embodiments of the present application, if the loss of the packets of the first data stream or the feedback information of the packets is caused by a link failure, but because there is a mechanism for retransmitting packets, the source end can still receive the feedback information of the packets with the above target sequence number. Then, in this case, the credit amount consumed by the packets corresponding to the unreceived feedback information cannot be filled back. In order to match the credit amounts of the respective paths corresponding to the first data stream with the congestion threshold of the first data stream for better congestion control management, the source end can adjust the congestion threshold of the first data stream to a value equal to the above actual credit amount.

[0039] In a possible implementation, the method further includes: the source end maps the source port number of the packets of the first data stream to n virtual port numbers, and the n virtual port numbers correspond to the n paths one by one.

[0040] In an embodiment of the present application, the source port number of the packets of a flow can be mapped to multiple different virtual port numbers to obtain different tuple information, so that the packets of the flow are hashed to multiple different paths for transmission based on the different tuple information, thereby achieving fine-grained load sharing.

[0041] In a second aspect, the present application provides a network device, which includes:

[0042] A sending unit, configured to send multiple packets of a first data flow through n paths, where each of the n paths is configured with a credit amount, and the credit amount indicates the size of the capacity of each path to send data. The sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, and the first congestion threshold is the congestion threshold of the first data flow. n is an integer greater than 1;

[0043] A receiving unit, configured to receive first feedback information, where the first feedback information is the feedback information of a first packet, and the first feedback information includes indication information on whether the first path is congested. The first path is the path used to send the first packet among the n paths; the first packet is any one of the multiple packets of the first data flow;

[0044] A processing unit, configured to re-determine the credit amount of the first path based on the indication information;

[0045] The processing unit is further configured to re-determine the load amount of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re-determined load amount of the first path is reduced.

[0046] In a possible implementation manner, the first feedback information is the information in a feedback packet received by the network device, and the feedback packet includes the feedback information of m packets among the multiple packets. m is an integer greater than 1;

[0047] The feedback information of the m packets includes indication information on whether the transmission path of each packet is congested during the transmission of the m packets.

[0048] In a possible implementation manner, the processing unit is specifically configured to:

[0049] Calculate a second congestion threshold based on the first feedback information, where the second congestion threshold is the new congestion threshold of the first data flow;

[0050] Adjust the credit amount of the first path based on a first difference, where the first difference is the difference between the second congestion threshold and the congestion threshold of the first data flow before calculating the second congestion threshold.

[0051] In a possible implementation manner, the sum of the credit amounts and the remaining credit amount of the n paths is equal to the first congestion threshold;

[0052] The processing unit is specifically configured to:

[0053] When the first difference is greater than zero and the sum of the first difference and the remaining credit amount is greater than the target credit amount, increase the credit amount of the first path by two times the target credit amount, where the target credit amount indicates the data volume size of a packet in the first data stream; or,

[0054] When the first difference is less than zero, reduce the credit amount of the first path by the first credit amount, where the first credit amount is the absolute value of the sum of the first difference and the target credit amount.

[0055] In a possible implementation, the first feedback information further includes indication information that the first packet is out of order at the destination end; each time the first path sends a packet, it consumes the target credit amount of the credit amount in the first path, where the target credit amount indicates the data volume size of a packet in the first data stream;

[0056] The processing unit is specifically configured to: maintain the current credit amount of the first path based on the indication information that the first packet is out of order at the destination end.

[0057] In a possible implementation, the sum of the credit amounts and the remaining credit amount of the n paths is equal to the first congestion threshold;

[0058] The sending unit is further configured to: when the credit amount of the first path drops to zero, send a probe packet through the first path;

[0059] The receiving unit is further configured to: receive second feedback information of the probe packet;

[0060] The processing unit is further configured to: when the second feedback information indicates that the first path is not congested, calculate a third congestion threshold based on the second feedback information, where the third congestion threshold is the new congestion threshold of the first data stream;

[0061] The processing unit is further configured to: when the sum of the second difference and the remaining credit amount is greater than the target credit amount, increase the credit amount of the first path by the target credit amount; the second difference is the difference between the third congestion threshold and the congestion threshold of the first data stream before calculating the third congestion threshold, and the target credit amount indicates the data volume size of a packet in the first data stream.

[0062] In a possible implementation, the receiving unit is further configured to: receive third feedback information, where the third feedback information is the feedback information of the second packet among the multiple packets, and the third feedback information includes information indicating that the second packet is lost, and the second packet is sent through the second path among the n paths;

[0063] The processing unit is further configured to increase the credit amount of the second path by a target credit amount, where the target credit amount indicates the data volume size of a packet in the first data stream.

[0064] In a possible implementation, the network device receives feedback information of a packet with a target sequence number, where the target sequence number includes all sequence numbers of packets sent through the n paths, and the processing unit is further configured to:

[0065] When the congestion threshold of the first data stream is greater than the actual credit amount, adjust the congestion threshold of the first data stream to a value equal to the actual credit amount, where the actual credit amount is the sum of the credit amounts of the n paths and the remaining credit amount.

[0066] In a possible implementation, the processing unit is further configured to:

[0067] Map the source port numbers of the packets in the first data stream to n virtual port numbers, where the n virtual port numbers correspond to the n paths one by one.

[0068] In a third aspect, an embodiment of the present application provides a network device, which may include: a memory, a processor, a sending interface, and a receiving interface coupled to the memory. Among them, the sending interface is used to support the network device to execute the sending step in the network congestion control method provided in the first aspect above. The receiving interface is used to support the network device to execute the receiving step in the network congestion control method provided in the first aspect above. Among them, the sending interface and the receiving interface may be integrated into a transceiver. The processor is used to support the network device to execute other processing steps in the network congestion control method provided in the first aspect above except for sending and receiving.

[0069] It should be noted that the sending interface and the receiving interface in the embodiments of the present invention may be integrated together or coupled through a coupler. The memory is used to store a computer program for implementing the network congestion control method described in the first aspect above, and the processor is used to execute the computer program stored in the memory. The memory and the processor may be integrated together or coupled through a coupler.

[0070] In addition, the computer program in the memory of the present application may be pre-stored or stored after being downloaded from the Internet when using the device. The present application does not specifically limit the source of the computer program in the memory. The coupling in the embodiments of the present application is an indirect coupling or connection between devices, units, or modules, which may be electrical, mechanical, or other forms, and is used for information interaction between devices, units, or modules.

[0071] In a possible implementation, the processor is used to execute the computer program stored in the memory, so that the network device performs the following operations:

[0072] Send multiple packets of a first data stream through an n number of paths via a sending interface. Each of the n paths is configured with a credit amount, which indicates the size of the capacity for each path to send data. The sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, which is the congestion threshold for the first data stream. The n is an integer greater than 1;

[0073] Receive first feedback information via a receiving interface. The first feedback information is the feedback information of a first packet, and the first feedback information includes indication information on whether the first path is congested. The first path is the path among the n paths used to send the first packet; the first packet is any one of the multiple packets of the first data stream;

[0074] Redetermine the credit amount of the first path based on the indication information; and redetermine the load of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the redetermined load of the first path is reduced.

[0075] In a possible implementation, the above network device may be a chip, such as a smart network card, an on-board network card, a field programmable gate array (FPGA), an acceleration card, a graphics processing unit (GPU), etc.

[0076] In a fourth aspect, an embodiment of the present application provides a system, which includes a first network device and a second network device. The first network device is the network device described in any item of the second aspect above, or the network device described in the third aspect above. The second network device is used to receive the packets sent by the first network device.

[0077] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the method described in any item of the first aspect above.

[0078] In a sixth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the method described in any item of the first aspect above.

[0079] Understandably, the network device described in the second aspect and the third aspect provided above, the system described in the fourth aspect, the computer storage medium described in the fifth aspect, and the computer program product described in the sixth aspect are all used to execute the method provided in any one of the first aspects above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 and Figure 2 FIG. shows a schematic diagram of the system architecture provided by an embodiment of the present application;

[0081] Figure 3 FIG. shows a schematic flowchart of a network congestion control method provided by an embodiment of the present application;

[0082] Figure 4 FIG. shows a schematic diagram of the relationship between a congestion threshold and a path credit amount provided by an embodiment of the present application;

[0083] Figure 5 FIG. shows an example diagram of the format of a feedback message provided by an embodiment of the present application;

[0084] Figure 6 FIG. shows a schematic diagram of a bitmap provided by an embodiment of the present application;

[0085] Figure 7A FIG. shows a schematic topology diagram of a three-layer fat tree network provided by an embodiment of the present application;

[0086] Figure 7B FIG. shows a schematic diagram of the change in the credit amount of a flow on each path;

[0087] Figure 8A FIG. shows a schematic flowchart of a method provided by an embodiment of the present application;

[0088] Figure 8B FIG. shows a schematic diagram of the virtual component architecture of a network device provided by an embodiment of the present application;

[0089] Figure 9 FIG. shows a schematic logical structure diagram of a device provided by an embodiment of the present application;

[0090] Figure 10 FIG. shows a schematic hardware structure diagram of a device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0092] See Figure 1 , Figure 1An exemplary schematic diagram of a system architecture applicable to the present application is shown. The system architecture includes a first device 110 and a second device 120. The first device 110 can send a message to the second device 120, that is, the first device 110 can be the source end, and the second device 120 can be the destination end.

[0093] Communication between the above-mentioned first device 110 and second device 120 is achieved through one or more network devices in the data forwarding plane of the network where they are located. The network where the first device 110 and the second device 120 are located can be, for example, a data center network (DCN), etc.

[0094] Exemplarily, the data center network can be in the networking form of a fat tree. For example, reference can be made to Figure 2 . Figure 2 An exemplary schematic diagram of a data center network with a possible fat tree topology is shown. It can be seen that the data center network includes a core layer, an aggregation layer, and an access layer. Each of the three layers includes multiple network devices. Connected to the network devices in the access layer can be end nodes, and the end nodes can be any node device with network communication capabilities such as servers, storage devices (such as disk enclosures), or dedicated computing nodes. The dedicated computing nodes include field programmable gate arrays (FPGAs), acceleration cards, and graphics processing units (GPUs), etc. The FPGAs, acceleration cards, and GPUs can all be in array form. Among them, the network devices in the aggregation layer, the network devices in the access layer, and the end nodes can be divided into m groups, and each group belongs to the devices in a container (pod), where m is an integer greater than 1. The network devices can be, for example, switches, routers, or access points, etc.

[0095] In a possible implementation manner, the above-mentioned first device 110 and the second device 120 can be end nodes in the data center network DCN, for example, any two of the above-mentioned Figure 2 end nodes.

[0096] It should be noted that the above-mentioned Figure 2 shown network system architecture is only an example and does not constitute a limitation on the network architecture applicable to the present application.

[0097] In existing communication networks, although there are mechanisms such as equal-cost multi-path routing (ECMP) technology and congestion control algorithms to alleviate load imbalance and network congestion, due to problems such as hash conflicts and poor response to burst traffic, these two mechanisms cannot meet the control needs of load imbalance and network congestion in large-scale networking.

[0098] Specifically, ECMP is a hop-by-hop flow-based load balancing strategy. When there are multiple optional paths for a flow, ECMP can use a variety of strategies to select the path to achieve a certain degree of load balancing. The routing strategies mainly include: hash routing based on five-tuples, polling to use multiple optional paths, and allocating flows based on path weights. ECMP is currently widely used because of its clear implementation rules and simple implementation. However, the ECMP mechanism has the following shortcomings:

[0099] 1. There is a hash conflict, and traffic will be concentrated on the local path, which is not good for the load balancing effect of the whole network;

[0100] 2. The perception of network congestion is limited to the current path, and does not support dynamic load adjustment based on network status;

[0101] 3. ECMP selects paths based on flows. If there are large differences between different flows, such as when elephant flows and mouse flows coexist, ECMP cannot reasonably arrange paths for different flows.

[0102] 4. There is no fault perception mechanism. Therefore, whether it is hashing, polling or weight-based routing, it is likely to aggravate congestion in the network.

[0103] In addition, in the existing solutions, there is also an adaptive routing method to deal with network failures and congestion problems. Adaptive routing is a forwarding method relative to static routing. It refers to the technology that the network forwarding unit dynamically adjusts the forwarding path of the flow to a specific destination according to the current state of the network. Compared with static routing, adaptive routing allows more active paths, can better deal with network failures and congestion, and achieve better network performance. However, on the one hand, the adaptive routing mechanism relies on the cooperation of network devices such as switches, requires hop-by-hop status notification and adaptive routing switching mechanism, and is not compatible with existing commercial switching equipment; on the other hand, the adaptive routing mechanism usually adopts a centralized routing control mechanism, which requires a dedicated management plane to manage global routing, and the end-to-end implementation is relatively complex.

[0104] Based on the above description, how to achieve better network load sharing and congestion control without the deficiencies and defects described above is a technical problem to be solved. For this purpose, the embodiments of the present application propose a network congestion control method, which can better achieve network load sharing and network congestion control compared with the existing solutions described above.

[0105] See Figure 3 , Figure 3 The network congestion control method proposed by the embodiments of the present application is shown. This method can be applied to the above Figure 1 or Figure 2 system architecture described. This method may include but is not limited to the following steps:

[0106] 301. The first network device sends multiple packets of the first data stream to the second network device through n paths. Each path among the n paths is configured with a credit amount, and the credit amount indicates the size of the data sending capacity of each path. The sum of the credit amounts of the n paths is less than or equal to the first congestion threshold, and the first congestion threshold is the congestion threshold of the first data stream. Here, n is an integer greater than 1.

[0107] Exemplarily, the above first network device may be the first device 110 in the above Figure 1 , or may be network components such as a smart network card, an on-board network card, an FPGA with a network interface, a GPU, or an accelerator in the first device 110. The above second network device may be the second device 120 in the above Figure 1 , or may be network components such as a smart network card, an on-board network card, an FPGA with a network interface, a GPU, or an accelerator in the second device 120.

[0108] The congestion threshold of the above data stream indicates the amount of packets of the data stream that can be sent currently, and the congestion threshold can be adjusted according to the congestion state of the network. The congestion threshold may be a congestion window.

[0109] The above-mentioned first data stream can be any stream sent by the first network device to the second network device. In a specific implementation, the embodiments of the present application are also applicable to congestion control of control flows. The above-mentioned n paths can be obtained by mapping the tuple information in the packets of the first data stream into n different tuple information and then performing hash-based routing based on the n different tuple information. Specifically, in hash-based routing, a hash value is obtained by performing a hash calculation on the tuple information in the packet. The hash value can correspond to a certain egress port. Then, the packet can be sent out from the certain egress port. Different egress ports correspond to different transmission paths. Therefore, after mapping the tuple information in the packets of the first data stream into n different tuple information, the packets of the first data stream can be hashed to multiple different paths for transmission, thereby achieving fine-grained load balancing and reducing bandwidth competition at the flow level.

[0110] Optionally, the above-mentioned hashing different packets to different paths for transmission based on n different tuple information can be implemented through the ECMP technology.

[0111] Optionally, the above-mentioned n paths can be regarded as n virtual paths, and the n virtual paths are mapped to m physical paths, where m is an integer greater than 1 and m is less than or equal to n. When m is less than n, it indicates that there is a hash conflict, and at least two of the above-mentioned n paths are mapped to the same physical path.

[0112] In a specific embodiment, the above-mentioned tuple information can be the information of the quadruple, quintuple or septuple of the first data stream. The quadruple includes the source Internet Protocol (IP) address, destination IP address, source port and destination port; the quintuple includes the source IP address, destination IP address, protocol number, source port and destination port; the septuple includes the source IP address, destination IP address, protocol number, source port, destination port, service type and interface index.

[0113] Specifically, before sending the packets of the first data stream, the first network device can change the tuple information in the packets of the first data stream with a preset tuple adjustment granularity. The preset tuple adjustment granularity can be a packet or a continuous packet sequence (Flowlet) of a certain length. Changing the tuple information in the packets of the first data stream mainly means changing the source port number in the packets of the first data stream. That is, in order to distribute the first data stream to the above-mentioned n paths for transmission, the source port number in the packets of the first data stream can be mapped to n different virtual source port numbers, so as to achieve mapping the tuple information in the packets of the first data stream into n different tuple information, and hashing the packets of the first data stream to multiple different paths for transmission based on the n different tuple information.

[0114] Optionally, in a layer 2 forwarding network without passing through a gateway, multiple different tuple information can also be obtained by mapping the source IP address of the packets of the first data stream to multiple different virtual IP addresses, so as to realize hashing the first data stream to multiple different paths for transmission based on the multiple different tuple information, so as to achieve fine-grained load sharing. Hereinafter, mainly taking the example of changing the source port number of the packets of the first data stream for introduction, the specific implementation of changing the source IP address can refer to the specific implementation of changing the source port number.

[0115] In a possible implementation manner, the above n different virtual source port numbers can be n port numbers generated within a certain offset range with a port number as a reference value; then, the first network device can poll the n different virtual port numbers to obtain the changed source port number in the packet. For the convenience of understanding, an example is given below, as shown in Table 1.

[0116] Table 1

[0117]

[0118] In Table 1, assume that the above reference value source port number is 200, and the above certain offset range is the range within 2 offsets on both sides of the reference value, that is, the port numbers 198, 199, 200, 201, and 202 in Table 1 above are the above n different virtual source port numbers. Additionally, assume that the above first data stream includes 8 packets, and the original source port number in these 8 packets is 100. Then, before the first network device sends these 8 packets, it first changes the source port numbers of these 8 packets based on Table 1 above. Specifically, assume that the packet sequence numbers of these 8 packets are from 1 to 8, and one packet is the above preset tuple adjustment granularity. Then, based on Table 1 above, the first network device can change the source port numbers in packets 1 to 8 to 200, 201, 202, 198, 199, 200, 201, and 202 in sequence. It can be seen that the first network device uses a polling method to obtain port numbers from Table 1 to adjust the source port numbers in the packets.

[0119] Optionally, assume that the packet sequence numbers of the above 8 packets are from 1 to 8, and the continuous 2-packet sequence is the above preset tuple adjustment granularity. Then, based on Table 1 above, the first network device can change the source port numbers in packets 1 to 8 to 200, 200, 201, 201, 202, 202, 198, and 198 in sequence. Similarly, it can be seen that the first network device uses a polling method to obtain port numbers from Table 1 to adjust the source port numbers in the packets, but this time it polls the port numbers in Table 1 every two consecutive packet sequences.

[0120] It should be noted that the above is only polling starting from the reference value port number, but the polling can start from any one of the above n virtual port numbers, and the present application does not limit this. Moreover, the n port numbers generated within a certain offset range with one port number as the reference value may not be symmetric about the reference value on both sides of the reference value. Or, the offset range of the port number can only be offset to the left, or the offset range of the port number can only be offset to the right, etc. The embodiments of the present application do not limit the specific offset direction of the port number.

[0121] Since the source port numbers of the above 8 packets have changed, their respective corresponding tuple information has also changed, and the output ports hashed from the tuple information of the 8 packets have also changed. Thus, the 8 packets can be forwarded through multiple output ports, that is, through multiple paths.

[0122] In another possible implementation manner, the above n different virtual port numbers may be randomly generated source port numbers. Then, similarly, the first network device can poll the n different virtual port numbers to obtain the changed source port numbers in the packets. The specific polling method can refer to the relevant description in Table 1 above and will not be elaborated here.

[0123] Generally, different tuple information is hashed to different output ports (or paths), but in the case of hash conflicts, different tuple information is hashed to the same output port (or path). Optionally, in order to map to different output ports (or paths), a source port number can be regenerated to ensure that the packets corresponding to the above n different virtual port numbers are mapped to n different output ports (or paths), so as to better achieve load sharing.

[0124] For the two generation methods of the above n different virtual port numbers, the specific generation method adopted can be determined according to actual needs. For example, if the first data stream is a remote direct memory access (RDMA) stream, any one of the above two generation methods can be adopted.

[0125] If the first data stream is a stream of Transmission Control Protocol (TCP), since the second network device needs to uniquely identify the first data stream through tuple information for context lookup, therefore, to ensure that each virtual port number can be accurately mapped to the first data stream, the above-described method of generating n port numbers with a port number as a reference value within a certain offset range can be used to determine n different virtual port numbers of the first data stream. And, in the second network device, there is a mechanism to map the n different virtual port numbers of the first data stream to the original port number of the first data stream.

[0126] Based on the above description, the first network device can load-balance the packet payloads of the above first data stream and send them on n paths, achieving fine-grained load sharing. However, when network congestion occurs, how to regulate the load on these n paths to effectively alleviate network congestion is also an urgent problem to be solved. The load on a path refers to the amount of packets sent through the path or the sending speed of the stream passing through the path, etc.

[0127] In the above introduction, the mechanism of evenly using each path is adopted, without considering the load state differences of multiple paths. In a communication network, the traffic state is complex, and the actual load of each path cannot be completely balanced, which easily leads to congestion on the paths with heavy load. In addition, there is no congestion control mechanism adapted to multi-path in the existing solutions. There may be several relatively congested paths among multiple paths. If the congestion control algorithm for a single path is directly used, it will form a "barrel short board" of path bandwidth, affecting the overall throughput and violating the original intention of fully utilizing the link bandwidth of multi-path. Specifically, there are existing congestion algorithms and their implementations for a single path. If directly used in a multi-path scenario, since it cannot distinguish which path the congestion comes from, it can only uniformly increase or decrease the speed of all paths. The most congested path will continuously feedback congestion information, slowing down all paths, and thus becoming the short board path. For this reason, a multi-path congestion control mechanism is proposed in the embodiments of the present application.

[0128] This multi-path congestion control mechanism maintains congestion thresholds at the granularity of a stream and manages the credit amount at the granularity of the transmission path of the stream. The credit amount can indicate the size of the capacity of sending data, and the capacity of this data can be the number of bytes occupied by the sent packets. For the sake of easy understanding, taking the above first data stream as an example, it can be seen Figure 4 .

[0129] In Figure 4As can be seen, the first data stream has its own congestion threshold CW_flow, that is, the congestion threshold CW_flow of the first data stream is maintained in the first network device. And since the first data stream transmits packets through the above n paths, the first network device can allocate the amount of credit included in the CW_flow to the n paths, and the sum of the amounts of credit allocated to the n paths can be equal to the CW_flow, that is

[0130] Optionally, The N_fraction is the remaining amount of credit. Exemplarily, the remaining amount of credit is less than the target amount of credit, and the target amount of credit indicates the data volume size of a packet in the first data stream. When the first network device allocates the amount of credit to the n paths, generally the amount of credit for sending an integer number of packets is allocated to each path. If there are still bytes left that are insufficient to send one packet after the allocation, then the number of bytes can be represented as N_fraction, that is, the above remaining amount of credit.

[0131] Based on the above description, in a possible implementation, at the initial configuration, the first network device can evenly distribute the amount of credit included in the congestion threshold CW_flow of the first data stream to the n paths, that is, the amount of credit initially configured for each of the n paths can be CW_flow / n. Optionally, at this time, the above remaining amount of credit N_fraction is zero. Or, optionally, the first network device can also allocate the amount of credit included in the congestion threshold CW_flow of the first data stream to the n paths according to other allocation rules, and the embodiments of the present application do not limit this. The other allocation rules can be, for example, allocating less credit to relatively congested paths, or allocating more credit to paths with higher priorities, and so on.

[0132] In a possible implementation, the initial value of the congestion threshold CW_flow of the first data stream can be set to the bandwidth delay product (BDP) or a value close to the BDP to obtain a higher starting rate. Optionally, the value of the BDP can be estimated based on the static delay of a single path between the first network device and the second network device in the network. The bandwidth delay product is a network performance metric. In data communication, the bandwidth delay product is also known as the bandwidth delay product or the bandwidth delay product, etc., and refers to the product of the capacity of a data link (bits per second) and the round-trip communication delay (in seconds).

[0133] In a specific embodiment, after the first network device configures the initial credit amount for n paths of the first data stream, it can map the source port number of the packets of the first data stream to the above-mentioned n different virtual port numbers in the polling manner described above to split the packets of the first data stream and send them on these n paths. For each packet sent on a certain path among the n paths, one target credit amount is consumed, that is, the credit amount for sending one packet is consumed.

[0134] 302. The second network device receives multiple packets of the first data stream.

[0135] After the first network device sends multiple packets of the first data stream to the second network device through the n paths, the second network device can receive the multiple packets of the first data stream from these n paths.

[0136] In a possible implementation, if the first data stream is a TCP stream, then, since the first data stream needs to be uniquely identified by the original tuple information of the first data stream for context lookup, after the second network device receives the packets of the first data stream, it can, based on the pre-configured mapping relationship, find the original source port number of the packet based on the virtual source port number in the packet, so as to restore the original tuple information of the packet. The second network device continues to perform subsequent operations such as context lookup based on the original tuple information of the packet. Optionally, the pre-configured mapping relationship is the mapping relationship between the above-mentioned n different virtual port numbers and the original source port number of the first data stream.

[0137] 303. The second network device sends feedback information on the multiple packets of the first data stream to the first network device.

[0138] In a specific embodiment, after the second network device receives the packets of the first data stream, it can reply a feedback information to the first network device for each received packet. The feedback information may include information indicating whether congestion occurs in the transmission path of each packet.

[0139] Taking the first packet as an example, the first packet is any one of the packets of the first data stream received by the second network device, and the first packet is transmitted to the second network device through the first path among the n paths. Then, after the second network device receives the first packet, it can determine that the packet is sent through the first path by the port number for receiving the packet or by the source port number in the packet. Then, the second network device can send the feedback information of the first packet (referred to as the first feedback information for the convenience of subsequent description) to the first device, and the first feedback information may include information indicating whether the first path is congested.

[0140] Optionally, the above feedback information can be sent to the first network device by carrying it in an acknowledge character (ACK) packet or a Selective ACK (SACK) packet.

[0141] In a possible implementation, when the second network device sends the feedback information of a packet, it can directly use the source port number of the packet as the destination port number of the feedback information. Alternatively, optionally, a pre-set port number can also be used as the destination port number of the packet. For example, when a feedback packet (such as the above ACK packet) includes the feedback information of m packets, the pre-set port number can be used as the destination port number of the feedback packet. Whether using the source port number of the received packet or the pre-set port number as the destination port number of the feedback packet, the feedback packet can be sent to the first network device. The m packets are some or all of the multiple packets of the first data stream sent by the first network device to the first network device, and m is an integer greater than 1.

[0142] 304. The first network device receives the above feedback information, re-determines the credit amount of the above n paths based on the feedback information, and then determines the load amount of each of the n paths based on the credit amount of each re-determined path.

[0143] After the first network device receives the feedback information of the packets of the first data stream sent back by the second network device, it can update the congestion threshold of the first data stream based on each feedback information, and then re-determine the credit amount of the transmission path of the packet corresponding to each feedback information based on the new congestion threshold. Optionally, the first network device can determine the corresponding path according to the destination port number in the packet where the feedback information is located. Optionally, calculating the new congestion threshold based on the feedback information of the packet can be calculated using an existing congestion control algorithm without increasing the computational complexity. For the sake of easy understanding, the following further describes with the first packet as an example.

[0144] In a possible implementation manner, when the above first feedback information includes information indicating congestion on the above first path, for example, the first feedback information includes a congestion notification-echo (CE) flag. When the CE flag is at position 1, it indicates that congestion has occurred during the transmission of the above first packet on the above first path. Then, the first network device calculates a new congestion threshold for the first data stream based on the first feedback information. The method for calculating the congestion threshold can adopt existing congestion control calculation methods, and the embodiments of the present application do not limit this. The new congestion threshold is smaller than the congestion threshold of the first data stream before calculation. Then, in order to relieve the congestion on the above first path, the credit amount of the first path can be reduced, thereby reducing the number of packets of the first data stream sent from the first path. Specifically, the credit amount of the first path is reduced by a first credit amount, and the first credit amount is the absolute value of the sum of a first difference and the above one target credit amount. The first difference is the difference between the above new congestion threshold and the congestion threshold of the first data stream before calculation, and this difference is less than zero. This is because after the first network device receives the first feedback information, it needs to backfill, that is, add a target credit amount to the credit amount of the first path again for continuing to send packets. And since congestion has occurred on the first path and the calculated new congestion threshold becomes smaller, the first path also correspondingly reduces the reduction amount of the congestion threshold. Then, the absolute value of the sum of the reduction amount of the congestion threshold and the added one target credit amount is the actual adjusted value of the first path.

[0145] For ease of understanding, an example is given. For example, assume that the length of a packet, that is, one target credit amount, is 10 bytes, and the first difference is -5 bytes. Then, the first credit amount is 5 bytes, that is, the credit amount of the first path is reduced by 5 bytes. Alternatively, the credit amount can be normalized, with one target credit amount represented by 1. Then, the first difference is -0.5, and the first credit amount is 0.5, that is, the credit amount of the first path is reduced by 0.5.

[0146] In a possible implementation, when the above first feedback information includes indication information indicating that there is no congestion on the above first path, for example, when the first feedback information includes a congestion notification CE flag and the CE flag bit is not set to 1, it indicates that there is no congestion in the transmission process of the above first packet on the above first path. Then, the first network device calculates a new congestion threshold for the first data stream based on the first feedback information. The method for calculating the congestion threshold can adopt existing congestion control calculation methods, and the embodiments of the present application do not limit this. The new congestion threshold is larger than the congestion threshold of the first data stream before calculation. Then, in order to improve the transmission speed of the first data stream, the credit amount of the first path can be increased, so as to increase the number of packets of the first data stream sent from the first path. Specifically, the difference between the above new congestion threshold and the congestion threshold of the first data stream before calculation is greater than zero. When the sum of this difference and the above remaining credit amount N_fraction is greater than a target credit amount, the credit amount of the first path is increased by two target credit amounts, that is, the data amount of two packets is increased. This is because after the first network device receives the first feedback information, it needs to backfill, that is, add a target credit amount to the credit amount of the first path again for continuing to send packets; in addition, since there is no congestion on the first path and the calculated new congestion threshold becomes larger, the first path can also adaptively increase the credit amount. When the sum of this difference and the above remaining credit amount N_fraction is greater than a target credit amount, it indicates that there is enough credit amount available to send a packet, so the credit amount is allocated to the first path, so the first path can increase the credit amount of two packet lengths, that is, increase two target credit amounts.

[0147] In a possible implementation, the above first path can increase or decrease the credit amount with the same granularity. For example, assuming that the credit amount of the first path increases in integer multiples of a target credit amount, then the credit amount of the first path can decrease in integer multiples of a target credit amount; for another example, assuming that the credit amount of the first path increases in integer multiples of 0.5 packet lengths, then the credit amount of the first path can decrease in integer multiples of 0.5 packet lengths.

[0148] Based on the above description, after the first network device re-determines the credit amounts of the above n paths, that is, determines the load amount transmitted on each of the n paths, it continues to send packets of the first data stream on the n paths based on the re-determined load amount.

[0149] In the network congestion control method introduced above, each feedback message without congestion indication information received contributes to the increment of the remaining credit in the congestion threshold of the first data stream. Although only the feedback message that causes the remaining credit to exceed a target credit will increase the credit of the corresponding path, from a global perspective, the opportunity for each path to increase its credit is still equal. When the congestion threshold decreases, it directly reduces the credit of the corresponding path, so the association between the reduction of credit and the paths is precisely corresponding. Considering a scenario where there are local congestions in several of the multiple paths, the congested paths will continuously reduce their credit due to continuously receiving feedback messages carrying congestion indication information, and this part of the reduced credit will be obtained by other paths through increasing their credit, thereby realizing the shuffle of credit among different paths and making full use of the bandwidth of non-congested paths.

[0150] In addition, the packets sent on the faulty path will not be replied with feedback messages. Then, after a period of time, for example, after a round-trip time (RTT), the credit on the faulty path will drop to zero, and the faulty path will be kicked out of the available paths, thereby reducing packet loss and saving the transmission resources of the first network device. A fault convergence time of one RTT is also a very good fault convergence performance.

[0151] Optionally, in the above solution for transmitting the packets of the first data stream over multiple paths, usually using 8 paths can achieve good load balancing performance.

[0152] In a possible implementation manner, for the case where a feedback message includes feedback information of m packets, in order to facilitate the first network device to identify the transmission paths of the packets corresponding to the multiple feedback information, in the embodiments of the present application, the second network device may record whether the packets received from each of the above n paths encounter congestion in the transmission path, and carry the recorded situation in the feedback message to inform the first network device.

[0153] Exemplarily, the packet of the first data stream sent from the first network device to the second network device includes a congestion notification (CN) flag bit. In the case where the packet encounters congestion during transmission, the network device where the congestion occurs will set the CN flag bit in the packet to 1 and then continue to send it. Additionally, a CE flag bit may be included in the feedback packet to feedback the indication information of whether the path is congested to the first network device. If the CN flag bit in the packet received by the second network device is set to 1, it indicates that congestion has occurred on the transmission path corresponding to the received packet. Then, in the feedback packet, the CE flag bit corresponding to the received packet is also set to 1 to inform the first network device that congestion has occurred on the transmission path of the received packet. On the contrary, if the CN flag in the packet received by the second network device is not set, it indicates that no congestion has occurred on the transmission path corresponding to the received packet. Then, in the feedback packet, the CE flag bit corresponding to the received packet is not set either, so as to inform the first network device that no congestion has occurred on the transmission path of the received packet.

[0154] There are many ways for the second network device to record whether the packets received from each of the above n paths encounter congestion in the transmission path. Three possible implementation methods are exemplarily introduced below:

[0155] The first implementation method is that the second network device records whether the packets received from each of the above n paths encounter congestion in the transmission path, including: recording the number of packets received from each of the above n paths, and recording the number of packets that encounter congestion among the received packets. In this way, when the first network device receives the feedback packet including this information, it can know the number of packets that encounter congestion and the number of packets that do not encounter congestion, so as to determine the calculation times of the congestion reduction threshold and the calculation times of the congestion increase threshold. The calculation times of the congestion reduction threshold refer to the calculation times of reducing the congestion threshold of the first data stream. Each time a feedback message indicating that a packet encounters congestion is received, a calculation of the congestion reduction threshold will be performed. The calculation times of the congestion increase threshold refer to the calculation times of increasing the congestion threshold of the first data stream. Each time a feedback message indicating that a packet does not encounter congestion is received, a calculation of the congestion increase threshold will be performed. For the convenience of understanding the information recorded by the second network device, exemplarily, Table 2 can be referred to.

[0156] Table 2

[0157] Path 0 Path 1 …… Path n - 1 Num_0 Num_1 …… Num_n - 1 CE_acc_0 CE_acc_1 …… CE_acc_n - 1

[0158] In Table 2, Num_i represents the number of packets of the first data stream received by the second network device on path i, where the value of i ranges from 0 to n-1; CE_acc_i includes the number of packets of the first data stream that have encountered congestion received by the second network device on path i. For example, for path 0, assume that the second network device receives 3 packets of the first data stream on this path 0. The CN flag bits of the first two packets among these 3 packets are not set to 1 but the default value 0, and the CN flag bit of the third packet is set to 1. Then, the information recorded by Num_0 is "3", and the information recorded by CE_accout_0 can be "1".

[0159] In the second implementation manner, the second network device records whether the packets received on each of the n paths encounter congestion in the transmission path as follows: For each packet received on each path, if a certain packet encounters congestion in the transmission path, record a preset flag indicating that the packet encounters congestion; if a certain packet does not encounter congestion in the transmission path, record a preset flag indicating that the packet does not encounter congestion. After the recording is completed, carry these flags together in the feedback packet and send them to the first network device. The first network device can determine the number of times to calculate the congestion reduction threshold based on the number of preset flags indicating that the packet encounters congestion in the feedback packet, and determine the number of times to calculate the congestion increase threshold based on the number of preset flags indicating that the packet does not encounter congestion in the feedback packet. The preset flag indicating that the packet encounters congestion can be, for example, "1", or it can be "Y", etc. The preset flag indicating that the packet does not encounter congestion can be, for example, "0", or it can be "N", etc. The embodiments of the present application do not limit this preset flag.

[0160] In the third implementation manner, the second network device records whether the packets received on each of the n paths encounter congestion in the transmission path as follows: respectively count and record the number of packets that encounter congestion and the number of packets that do not encounter congestion in the received packets. Then, carry the recorded number of packets that encounter congestion and the number of packets that do not encounter congestion together in the feedback packet and send them to the first network device. The first network device can determine the number of times to calculate the congestion reduction threshold based on the number of packets that encounter congestion in the feedback packet, and determine the number of times to calculate the congestion increase threshold based on the number of packets that do not encounter congestion in the feedback packet.

[0161] Optionally, the second network device can record whether the packets received on each of the n paths encounter congestion in the transmission path by recording in the context.

[0162] After recording whether the packets received on each of the above n paths have encountered congestion in the transmission path, the second network device can send the recorded information to the first network device through a feedback packet. Specifically, the second network device can add an extended packet header to the feedback packet and copy the recorded information into the extended packet header and then feedback it to the first network device. For example, reference can be made to Figure 5 。 Figure 5 Taking the first implementation mode above as an example, an exemplary format diagram of a feedback packet is shown. It can be seen that in addition to including the media access control (MAC) address, IP address, user datagram protocol (UDP) information, and base transport header (BTH) in the packet header of the feedback packet, it also includes an extended header, and the information shown in Table 2 above is included in the extended header. The payload in the feedback packet can be empty. It should be noted that the information recorded in the second implementation mode or the third implementation mode above can also be copied into the above extended packet header and sent to the above first network device.

[0163] After the first network device receives the feedback packet including multiple feedback messages, it can calculate the congestion threshold of the first data stream based on each feedback message in the feedback packet, and determine the credit amount and load amount of the corresponding path based on the new congestion threshold. The specific implementation operations can refer to the corresponding description in step 304 above and will not be elaborated here.

[0164] In a possible implementation mode, the multiple feedback messages included in the feedback packet can be the feedback messages of the packets sent through some of the above n paths, or can be the feedback messages of some of the packets sent through the n paths, and are not limited to the feedback messages of the packets sent through all of the above n paths. Optionally, in the feedback packet, if the feedback message of the packet sent on a certain path among the above n paths is not included, then in the feedback packet, the information about the number of packets corresponding to the certain path and the congestion situation is empty or 0.

[0165] Based on the above description, the first data stream is load-balanced and transmitted on the above n paths. However, due to the different delays of different paths, the packets received by the second network device are out of order. To alleviate the out-of-order impact caused by the multi-path delay difference, out-of-order receiving (OOR) technology can be used to process the received packets. When using OOR technology, a bitmap is usually introduced to record the situation of the received packets. For the convenience of understanding, it can be exemplarilyFigure 6 Illustrated by way of example.

[0166] In Figure 6 it is assumed that the bitmap includes 15 bits, that is, the length of the bitmap is 15. Generally, the sequence number of the packet recorded by the bit at position number 1 in the bitmap is preset, that is, the sequence number of the packet recorded by the bit at position number 1 is specified in advance. Then, the packets after the preset packet sequence number are sequentially recorded in the bits after the bit at position number 1 in the order of the sequence numbers. For example, assuming that the bit at position number 1 records the packet with the sequence number 100, then the bits at positions 2 to 15 record the packets with the sequence numbers 101 to 114 in sequence, that is, there is a one-to-one mapping relationship between the sequence numbers of the preset received packets and the bits of the bitmap. Specifically, the initial value of the bits in the bitmap can be 0. After receiving a packet, the bit position corresponding to the sequence number of the packet is set to 1. For example, if the sequence number of the received packet is 101, then the bit at position number 2 is set to 1, so as to record the sequence number of the received packet, and thus the out-of-order situation of the packets can be monitored through the bitmap.

[0167] The above Figure 6 The length of the bitmap and the sequence number of the packet shown above are only an example and do not constitute a limitation on the implementation of this application. The embodiments of this application do not limit the length of the bitmap and the specific sequence numbers of the received packets.

[0168] The length of the above bitmap is subject to certain limitations. When the delay differences of multiple paths are relatively large, for example, when some paths are congested, there may be a situation where the sequence number of the received packet exceeds the maximum sequence number that the bitmap can record (this situation can be called bitmap overflow). And the packets whose sequence numbers exceed the maximum sequence number that the bitmap can record can only be discarded and then wait for retransmission, and this additional retransmission will cause bandwidth loss. In order to reduce the bandwidth loss caused by retransmission, the embodiments of this application provide a solution for controlling the out-of-order degree between multiple paths, which can actively kick out the path with the largest delay when the out-of-order exceeds a certain degree, ensure that the delay differences between multiple paths are controlled within a certain range, and avoid bitmap overflow.

[0169] Specifically, an out-of-order degree threshold (oor_degree_max) can be set in the second network device. This threshold is the maximum acceptable absolute value of the difference between two packet sequence numbers, and the two packet sequence numbers are the sequence number of the currently received packet and the maximum sequence number of the already received packets. The maximum acceptable value can be set according to the actual situation, and the embodiments of this application do not limit this.

[0170] Then, after the second network device receives a packet, taking this packet as the aforementioned first packet for example, the second network device obtains the sequence number in the first packet, and obtains the maximum sequence number of the packets received before receiving the first packet, and calculates the absolute value of the difference between the sequence number of the first packet and the maximum sequence number. If this absolute value is greater than the aforementioned out-of-order degree threshold, it indicates that serious out-of-order has occurred in the received packets. The second network device sends feedback information of the first packet (i.e., the aforementioned first feedback information) to the first network device based on this out-of-order situation. In addition to including information indicating whether congestion has occurred in the aforementioned first path, the first feedback information also includes indication information that the first packet is out-of-order at the destination end, i.e., in this second network device. The indication information of out-of-order can be, for example, setting the bitmap overflow flag in the packet of the first feedback information to 1.

[0171] After the first network device receives the first feedback information including the aforementioned out-of-order indication information, since the first feedback information is the feedback information for the first packet and the first packet is transmitted in the first path, the first network device can keep the current credit amount of the first path unchanged, and subtract a target credit amount from the current congestion threshold of the first data stream to maintain the balance between the path credit amount and the congestion threshold. That is, for the received first feedback information including the aforementioned out-of-order indication information, the operations of backfilling a transmission packet length into the credit amount of the first path and calculating a new congestion threshold to re-determine the credit amount of the first path are no longer performed.

[0172] Based on the above description, since each transmitted packet consumes a target credit amount, if the packets received by the second network device from the first path continuously show out-of-order situations exceeding the threshold, then the credit amount of the first path in the first network device will gradually decrease until it becomes zero. After the credit amount of the first path becomes zero, the first network device will no longer send packets of the first data stream through the first path. Thus, the situation of bitmap overflow in the second network device can be alleviated or even avoided, thereby reducing the retransmitted packets and saving bandwidth resources. In addition, the continuous out-of-order situations of the packets transmitted in the first path indicate that the delay of the first path is large and congestion has occurred. In the embodiment of the present application, by reducing or even not sending packets on the first path, the congestion situation of the first path is further effectively alleviated.

[0173] In a possible implementation manner, after the credit amount of the first path drops to 0, the packet of the first data stream can be sent on the first path through a path recovery mechanism.

[0174] In a specific embodiment, the first network device may monitor the credit amounts of the above n paths. When the credit amount of a certain path drops to 0, a path recovery mechanism is triggered. In a possible implementation manner, the first network device may set a credit deficit flag field for the above n paths. The credit deficit flag field may include n bits, and each bit is used to mark whether the credit amount of one of the n paths has a deficit, that is, whether the credit amount drops to zero. Taking the first path as an example above, if the credit amount in the first path drops to zero, the bit corresponding to the first path in the credit deficit flag field is set to 1. When the first network device detects that the bit corresponding to the first path is set to 1, it starts the recovery mechanism for the first path.

[0175] Specifically, for the recovery mechanism of the first path, the first network device may send a probe message to the second network device through the first path. Sending a probe message through the first path is the same as the processing of sending the message of the first data stream through the first path, that is, the source port number of the message is mapped to the virtual port number corresponding to the first path and then sent out. If the probe message encounters congestion during the transmission process of the first path, the probe message may also be marked with congestion by the network devices it passes through. If there is no congestion, the probe message received by the second network device has no congestion mark.

[0176] After receiving the probe message sent by the first network device, the second network device, similarly, sends feedback information of the probe message to the first network device. The feedback information also includes information indicating whether the first path is congested. After receiving the feedback information, if the feedback information indicates that the first path is congested, that is, the CE flag bit in the message where the feedback information is located is set to 1, then the first network device does not perform other processing on the feedback information.

[0177] Optionally, in the case where the feedback information indicates that the first path is congested, the first network device no longer continues to send probe messages to the second network device through the first path, but sends probe messages to the second network device through the first path after an interval of time. During this period, the first network device may send probe messages to the second network device through other paths whose credit amounts have dropped to zero.

[0178] If the above feedback information indicates that there is no congestion on the first path, that is, the CE flag bit in the packet where the feedback information is located is not set to 1, then the first network device recalculates the congestion threshold of the first data stream based on the feedback information. When the sum of the difference between the new congestion threshold and the congestion threshold of the first data stream before calculation plus the above remaining credit amount is greater than a target credit amount, the first network device can allocate a target credit amount to the first path, thereby resuming sending packets of the first data stream on the first path. At the same time, the corresponding bit in the credit deficit identification field for the first path is restored to the default value, for example, set to 0.

[0179] In addition, optionally, when the sum of the difference between the new congestion threshold and the congestion threshold of the first data stream before calculation plus the above remaining credit amount is still less than a target credit amount, the first network device continues to send probe packets to the second device through the first path. For subsequent operations, refer to the foregoing description and will not be elaborated here.

[0180] Optionally, the payload in the above probe packet can be empty or 0, so that the sending of the probe packet has little impact on the network bandwidth, and the sending frequency of the probe packet is low and will not aggravate the congestion state of the path.

[0181] Through the above path recovery mechanism, available paths can be restored in a timely manner, the load sharing ability can be improved, and the situation of finally degenerating into sending packets of the first data stream on a single path due to congestion can be avoided.

[0182] In a possible implementation, packets of the first data stream may be lost during transmission. Packet loss will cause the value of the congestion threshold of the first data stream to be unequal to the sum of the credits of the above n paths and the remaining credit amount. That is, when the number of in - flight packets of the first data stream is 0, the value of the congestion threshold of the first data stream will be greater than the sum of the available credits of all paths of the n paths and the remaining credit amount. Long - term cumulative packet loss will cause this difference to become larger and larger, resulting in inaccurate congestion control management. The in - flight packets of the first data stream being 0 means that the first network device receives feedback information of a packet with a target sequence number, and the target sequence number includes all sequence numbers of packets sent through the above n paths. Although the first network device receives the feedback information of the packet with the target sequence number, the packets re - transmitted due to packet loss also consume the corresponding path's credit amount, and due to packet loss, the first network device does not receive the feedback information of these lost packets, so the corresponding path's credit amount will not be backfilled, resulting in the aforementioned inequality.

[0183] Currently, the main scenarios where packet loss occurs include the above bitmap overflow, link failure, and no feedback information when pseudo - retransmission (duplicate packets) occurs. The following will introduce them case by case.

[0184] For the case of packet loss caused by bitmap overflow, when the second network device receives a packet with a sequence number overflowing the bitmap, it can send a feedback message of this packet to the first network device. This feedback message is used to indicate that the packet has not been correctly received, that is, packet loss has occurred. For example, this feedback message can be a non-acknowledgment character NACK. After receiving this feedback message, the first network device can fill back a target credit amount in the credit amount of the transmission path of this packet, that is, increase the credit amount of this transmission path by a target credit amount.

[0185] For the case of packet loss caused by the above link failure, what may be lost due to the link failure is either a packet or the feedback message of this packet. In one possible implementation for this case, when the inflight packets of the first data stream are 0, the first network device can check the credit amounts of the above n paths. If the sum of the credit amounts of these n paths and the remaining credit amount is not equal to the value of the congestion threshold of the first data stream, and the value of the congestion threshold of the first data stream is larger, then the first network device can reduce the congestion threshold of the first data stream so that the value of the congestion threshold of the first data stream is equal to the sum of the credit amounts of these n paths and the remaining credit amount.

[0186] For the case of packet loss caused by the above link failure, in another possible implementation, a scheme similar to the link layer fault recovery of Fibre Channel (FC) can be adopted. Specifically, the first network device sends a specific synchronization message to the second network device every M packets sent. The second network device can check whether the number of packets received before receiving this synchronization message is M based on this synchronization message. If it is M, then there is no packet loss. If it is less than M, then packet loss has occurred. Similarly, the second network device also replies with a synchronization message for every M packets' feedback messages. The first network device can check whether the number of feedback messages received before receiving this synchronization message is M based on this synchronization message. If it is M, then there is no packet loss. If it is less than M, then packet loss has occurred. Based on this method, it can also be checked whether the credit amount of a path is lost. If so, the congestion threshold of the first data stream is reduced to be equal to the sum of the credit amounts of these n paths and the remaining credit amount.

[0187] For the case of the above pseudo-retransmission, after the second network device receives a duplicate packet, it also sends a feedback message of this packet to the first network device, so that the first network device can determine the credit amount of the corresponding transmission path based on this feedback message. For the specific implementation method, refer to the corresponding description in step 304 above, which will not be elaborated here.

[0188] To facilitate the understanding of the above-introduced network congestion control method, the following combines Figure 7A and Figure 7BExemplary introduction.

[0189] Figure 7A Shown is a topological schematic diagram of a three - layer fat - tree network. Figure 7A It is assumed that there are 4096 end - nodes, which are connected to the network devices in the access layer. The end - node can be, for example, the server described above Figure 2 in the above. These 4096 end - nodes are distributed within 16 pods. Inside each pod, there is a two - layer (access layer and aggregation layer) fat - tree network topology. It is assumed that the networking convergence ratio of the access layer and the aggregation layer within each pod is 2:1. For example, there are 16 network devices in the access layer and 8 network devices in the aggregation layer within each pod. The network devices in the access layer are also called top - of - rack (TOR) nodes, and the network devices in the aggregation layer are also called leaf nodes. Each network device included in the access layer and the aggregation layer has a corresponding number. For the specific numbering, reference can be made to Figure 7A . In addition, the leaf nodes in the aggregation layer are connected to the spine nodes in the core layer. Among them, it is assumed that the core layer includes 8 spine - node sets, and each spine - node set includes 8 spine nodes. It is assumed that the static BDP when forwarding packets through 5 - hop nodes is 20 maximum transmission units (MTU), that is, the static BDP when forwarding packets through 5 - hop nodes is 20 * MTU. Figure 7A The representation of some links and nodes in is omitted.

[0190] For the sake of convenience of explanation, it is assumed that the above - mentioned n virtual port numbers used by each flow of each node in this embodiment are 4 virtual port numbers. Here, the working mode and effect of the multi - path technology are illustrated through the following three steady - state data flows (Flow 1, Flow 2, and Flow 3).

[0191] Flow 1: Sent from node A to node E; node A is node 0 under TOR node 4096 in pod 0, and node E is node 15 under TOR node 4456 in pod 15.

[0192] Flow 2: Sent from node B to node E; node B is node 15 under TOR node 4096 in pod 0.

[0193] Flow 3: Sent from node C to node D; node C is node 255 under TOR node 4471 in pod 15, and node D is node 0 under TOR node 4456 in pod 15.

[0194] Figure 7AIn the figure, packets of flow 1 are exemplarily shown on the left side of node A. Each small square represents a packet, and the "1" in the square is used to indicate that it is a packet of flow 1. Packets of flow 2 are exemplarily shown on the right side of node B. Each small square represents a packet, and the "2" in the square is used to indicate that it is a packet of flow 2. Packets of flow 3 are exemplarily shown on the right side of node C. Each small square represents a packet, and the "3" in the square is used to indicate that it is a packet of flow 3. Figure 7A The packets on the link indicate that the packets are being transmitted on the link.

[0195] Figure 7B The figure shows the credit amounts of the above three flows on each path, and the process of how the credit amounts of each path change with the network state.

[0196] Initial flow establishment: Taking flow 1 as an example. For flow 1, the multi-path software logic unit in node A will select a reference source port number (base_port). The virtual source port numbers used by multi-path are generated by offsetting the reference port number. For example, assume that the available source ports for flow 1 are base_port to base_port + 3. In addition, flow 1 is a flow from node A to node E and needs to be forwarded through 5-hop nodes. Since the BDP value when forwarding packets through 5-hop nodes is 20 * MTU based on the above assumption, the initial congestion threshold of flow 1 can take a value close to this BDP value, for example, it can be 16 * MTU, and the total credit amount equal to the initial congestion threshold is evenly distributed to the 4 transmission paths of flow 1, that is, each transmission path obtains an initial credit amount of 4 * MTU. The initial flow establishment of flow 2 can refer to the description of the initial flow establishment of flow 1. For flow 3, since flow 3 is a flow from node C to node D and only needs to be forwarded through 3-hop nodes, however, to reserve some extra credit amounts, the initial congestion threshold of flow 3 can also be set to a value close to the BDP value when forwarding packets through 5-hop nodes, for example, it can also be set to 16 * MTU. In addition, the process of the initial flow establishment of flow 3 can refer to the description of flow 1 above and will not be elaborated here.

[0197] In a specific embodiment, assume that packets of flow 1 with different virtual source port numbers are evenly hashed to 4 upstream paths on TOR node 4096. For flow 2, due to hash conflicts, packets corresponding to multiple virtual source port numbers of flow 2 are hashed to the 2nd upstream path of TOR node 4096. For flow 3, more traffic is also allocated to the 2nd path of TOR node 4471 for flow 3.

[0198] On leaf node 4119, all data of flow 1 and flow 2 are hashed onto the path leading to spine node 4536, forming a slight congestion; on the path from leaf node 4479 to TOR node 4456, partial traffic of flow 1, flow 2, and flow 3 converges, causing a relatively serious congestion. At the same time, since the destination nodes of flow 1 and flow 2 are both node E, a convergence also occurs on the downstream path from TOR node 4456 to node E, causing congestion.

[0199] Based on the above description, after multiple RTTs of multi-path congestion control and bandwidth rotation, finally, the above three flows all achieve convergence. For the specific congestion control and bandwidth rotation, reference can be made to the corresponding descriptions in the network congestion control method and its possible implementation manners shown above. Figure 3 It will not be elaborated here.

[0200] After the above three flows all achieve convergence, reference can be made to Figure 7B the path credit volume situation corresponding to time point 1 therein. Among them, the total congestion thresholds of flow 1 on node A and flow 2 on node B both drop to 1 / 2 of the BDP at 5-hop forwarding, that is, 10 * MTU. And it is assumed that path 4 in both flow 1 and flow 2 is the path with more sent packets or congestion. Therefore, it can be seen that the credit volume of path 4 in both flow 1 and flow 2 drops to a relatively small value. Affected by the path from leaf node 4479 to TOR node 4456 (this path has serious congestion), the delay on the corresponding path 4 of flow 3 on node C is too long, resulting in out-of-order exceeding the threshold, and this path 4 is kicked out, and the credit volume drops to 0. However, the credit volumes of other non-congested paths increase, and the total congestion threshold still reaches the BDP value at 3-hop forwarding (this BDP value at 3-hop forwarding can be, for example, 12 * MTU), that is, full-bandwidth transmission. The specific implementation of this out-of-order control can be referred to the specific introduction in the above-mentioned scheme for controlling the out-of-order degree between multiple paths. It will not be elaborated here.

[0201] In addition, reference can be made to Figure 7B the path credit volume situation corresponding to time point 2 therein. Specifically, after the above time point 1, it is assumed that the path from leaf node 4112 to spine node 4487 fails. The corresponding paths of flow 1 and flow 2 (assumed to be both path 2) are kicked out after one RTT due to credit volume exhaustion, and the credit volumes of the corresponding paths drop to 0. The specific implementation of the credit volume of the corresponding path dropping to 0 due to the failure can also be referred to the above description. It will not be elaborated here. Flow 3 detects its corresponding path 4 through probe packets, re-enables path 4, and obtains a more balanced multi-path load balance while maintaining full bandwidth. The specific implementation of the mechanism for path recovery through probe packets can also be referred to the above description. It will not be elaborated here.

[0202] In summary, for the problems of insufficient network load balancing in the current data center, low overall link utilization, and high dynamic latency caused by network congestion, the existing ECMP technology performs hash-based routing based on quintuples at the flow granularity, and it is difficult to solve the problem of uneven load caused by hash conflicts and multi-node independent routing in large-scale network deployments. The existing adaptive routing technologies require the cooperation of components such as the control plane and switches, and need private hop-by-hop status notification and adaptive routing switching mechanisms, and are complex to implement and cannot interface with commercial switching devices. In contrast, the embodiments of the present application implement multi-path forwarding at the packet level or flowlet level through the network congestion control method shown in Figure 3 to reduce bandwidth competition at the flow level, and through multi-path congestion control and dynamic load balancing between paths, achieve a significant improvement in the overall network link utilization and greatly reduce network congestion hotspots.

[0203] Regarding the problem of relying on the control plane to achieve a more comprehensive perception of network state information and implement more balanced congestion control, flow control, and dynamic balancing, existing congestion control algorithms only cover single-flow single-path forwarding, and can only control the congestion degree of a single path based on the status of a specific path, and cannot achieve multi-path congestion control and load balancing. The embodiments of the present application achieve the perception of the congestion status of multiple paths by the source end, and adjust the bandwidth allocation between paths to distribute traffic proportionally and on demand, thereby achieving more balanced congestion control, flow control, and dynamic balancing.

[0204] Regarding the problem of how to quickly avoid and discover faulty paths in large-scale network deployments and reduce the impact of network failures, existing technologies identify faulty paths based on in-band or out-of-band status monitoring and avoid faulty paths through the intervention of the management plane. Among them, the out-of-band monitoring method has low real-time performance and slow network fault convergence. Although the in-band monitoring method based on technologies such as in-band network telemetry (INT) has high real-time performance, it requires the introduction of complex in-band telemetry rules, the deployment of information collectors, analyzers and other components, and the fault path convergence performance depends on the performance of the control plane, and the performance is limited in large-scale deployments. The embodiments of the present application identify and feedback severely out-of-order packets, and cooperate with the credit management at the source end to kick out severely congested or faulty paths within one RTT, achieving quick avoidance, greatly reducing the fault convergence time, reducing the impact of network failures, and realizing the screening of the optimal latency path and excellent fault convergence performance. In addition, the embodiments of the present application actively explore potential available paths with little impact on the existing forwarded traffic, obtain the bandwidth of congestion mitigation and fault recovery paths, achieve quick recovery, and ensure that more paths are used as much as possible to achieve better load balancing performance.

[0205] To better understand the method embodiments introduced above, the following is combined with Figure 8A examples for illustration.Figure 8A An exemplary flowchart showing a possible process of an embodiment of the present application is presented. Specifically, the process may include, but is not limited to, the following steps:

[0206] S1: The source end sends multiple packets of the first data stream through the above n paths;

[0207] S2: The source end receives the first feedback information of the first packet among the multiple packets;

[0208] S3: The source end determines whether the first feedback information includes an indication that the first packet is out of order at the destination end; if the first feedback information does not include an indication that the first packet is out of order at the destination end, S4 is executed; if the first feedback information includes an indication that the first packet is out of order at the destination end, S10 is executed;

[0209] S4: The source end recalculates the congestion threshold of the first data stream based on the first feedback information;

[0210] S5: The source end determines whether the difference between the recalculated congestion threshold and the congestion threshold of the first data stream before calculation is greater than zero; if it is greater than zero, S6 is executed, and if it is less than zero, S9 is executed;

[0211] S6: The source end determines whether the sum of the difference and the remaining credit amount is greater than the target credit amount; if it is less than the target credit amount, S7 is executed; if it is greater than the target credit amount, S8 is executed;

[0212] S7: The source end maintains the credit amount of the first path unchanged;

[0213] S8: The source end increases the credit amount of the first path by two target credit amounts;

[0214] S9: The source end reduces the credit amount of the first path by the first credit amount, where the first credit amount is the absolute value of the sum of the first difference and the target credit amount;

[0215] S10: The source end maintains the current credit amount of the first path based on the indication that the first packet is out of order at the destination end;

[0216] S11: After executing S9 or S10, the source end determines whether the credit amount of the first path has dropped to zero;

[0217] S12: When the credit amount of the first path drops to zero, the source end sends a probe packet through the first path and receives the second feedback information of the probe packet;

[0218] S13: The source end determines whether the second feedback information indicates that the first path is congested; if it indicates congestion, S14 is executed, and if it indicates no congestion, S15 is executed;

[0219] S14: The source end keeps the credit volume of the first path unchanged, i.e., keeps it at zero;

[0220] S15: The source end recalculates the congestion threshold of the first data stream based on the second feedback information;

[0221] S16: The source end determines whether the sum of the difference between the recalculated congestion threshold and the congestion threshold of the first data stream before calculation and the remaining credit volume is greater than the target credit volume. If it is greater, execute S17; if it is less, execute S14;

[0222] S17: The source end increases the credit volume of the first path by a target credit volume.

[0223] The above Figure 8A For the specific implementation of each step above, reference can be made to the description in the above Figure 3 and its possible implementation manners, which will not be elaborated here.

[0224] Based on the above-introduced network congestion control method and its possible implementation manners, the following exemplarily gives a virtual component architecture of the first network device. The first network device can implement the operations performed by the first network device in the above network congestion control method and its possible implementation manners through these components. Exemplarily, reference can be made to Figure 8B .

[0225] As Figure 8B shown, the virtual component structure of the first network device includes a data plane channel and a multi-path control unit. Among them, the data plane channel includes a data plane forwarding unit, a multi-path Scatter&Gather unit, a sending channel, and a receiving channel; the multi-path control unit includes a congestion control unit, a credit volume management unit, an out-of-order control unit, and a path detection unit.

[0226] The main functions of the above component units are as follows:

[0227] Multi-path Scatter&Gather unit: In the sending direction, modify the source port number of the packet according to the preset tuple adjustment granularity in step 301 above, and then realize the multi-path (such as the above n paths) distribution of the packet through the sending channel; in the receiving direction, receive the packet through the receiving channel, and parse the multi-path packet to identify the data stream and context associated with the packet;

[0228] Data plane forwarding unit: Realize the basic forwarding and actions of the data plane of the first network device, including searching the routing table, packet editing, and quality of service (Qos) control, etc.;

[0229] Congestion control unit: Implement the multi-path congestion control algorithm, and realize the adjustment of the sending window or rate in terms of flow granularity according to the network state feedback;

[0230] Credit volume management unit: Cooperating with the congestion control unit, it manages the real-time credit volume of each path (for example, each of the above n paths), and realizes multi-path load balancing through credit volume allocation; in addition, it also assists in kicking out faulty and congested paths.

[0231] Out-of-order control unit: The second network device also includes this out-of-order control unit. When the second network device is the destination end of the packet, its out-of-order control unit tracks the out-of-order degree of the packet in terms of flow granularity, gives feedback to the packets whose out-of-order exceeds the threshold, and assists the out-of-order control unit of the source end, that is, the first network device above, to kick out the paths with severely large forwarding delays.

[0232] Path detection unit: According to the multi-path congestion control information, it actively initiates the detection of potential available paths to realize the recovery of available paths.

[0233] Figure 8B For the specific operations implemented by each of the units shown, reference can be made to the corresponding descriptions in the foregoing Figure 3 and its possible implementation manners, which will not be elaborated here.

[0234] It can be understood that in order for each device to implement the corresponding functions above, it includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of this application.

[0235] The embodiments of this application can divide the functions of the device according to the above method examples. For example, each function can correspond to a divided functional module, or two or more functions can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0236] In the case of dividing each functional module corresponding to each function, Figure 9 A possible schematic diagram of the logical structure of the device is shown. This device can be the first network device above, etc. This device 900 includes a sending unit 901, a receiving unit 902, and a processing unit 903. Among them:

[0237] A sending unit 901 is configured to send multiple packets of a first data stream through n paths. Each of the n paths is configured with a credit amount, which indicates the size of the capacity for sending data of each path. The sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, and the first congestion threshold is the congestion threshold of the first data stream. Here, n is an integer greater than 1.

[0238] A receiving unit 902 is configured to receive first feedback information, which is the feedback information of a first packet. The first feedback information includes indication information on whether the first path is congested. The first path is the path used to send the first packet among the n paths. The first packet is any one of the multiple packets of the first data stream.

[0239] A processing unit 903 is configured to re-determine the credit amount of the first path based on the indication information.

[0240] The processing unit 903 is further configured to re-determine the load amount of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re-determined load amount of the first path decreases.

[0241] In a possible implementation, the first feedback information is the information in a feedback packet received by the network device. The feedback packet includes the feedback information of m packets among the multiple packets, and m is an integer greater than 1.

[0242] The feedback information of the m packets includes indication information on whether the transmission path of each packet is congested during the transmission of the m packets.

[0243] In a possible implementation, the processing unit 903 is specifically configured to:

[0244] Calculate a second congestion threshold based on the first feedback information. The second congestion threshold is the new congestion threshold of the first data stream.

[0245] Adjust the credit amount of the first path based on a first difference. The first difference is the difference between the second congestion threshold and the congestion threshold of the first data stream before calculating the second congestion threshold.

[0246] In a possible implementation, the sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold.

[0247] The processing unit 903 is specifically configured to:

[0248] In the case that the first difference is greater than zero and the sum of the first difference and the remaining credit amount is greater than the target credit amount, increase the credit amount of the first path by two times the target credit amount, where the target credit amount indicates the data volume size of a packet in the first data stream; or,

[0249] In the case that the first difference is less than zero, reduce the credit amount of the first path by the first credit amount, where the first credit amount is the absolute value of the sum of the first difference and the target credit amount.

[0250] In a possible implementation, the first feedback information further includes indication information that the first packet is out-of-order at the destination; each time the first path sends a packet, it consumes the target credit amount of the credit amount in the first path, where the target credit amount indicates the data volume size of a packet in the first data stream;

[0251] The processing unit 903 is specifically configured to: maintain the current credit amount of the first path based on the indication information that the first packet is out-of-order at the destination.

[0252] In a possible implementation, the sum of the credit amounts and the remaining credit amount of the n paths is equal to the first congestion threshold;

[0253] The sending unit 901 is further configured to, when the credit amount of the first path drops to zero, send a probe packet through the first path;

[0254] The receiving unit 902 is further configured to receive second feedback information of the probe packet;

[0255] The processing unit 903 is further configured to, when the second feedback information indicates that the first path is not congested, calculate a third congestion threshold based on the second feedback information, where the third congestion threshold is the new congestion threshold of the first data stream;

[0256] The processing unit 903 is further configured to, when the sum of the second difference and the remaining credit amount is greater than the target credit amount, increase the credit amount of the first path by the target credit amount; the second difference is the difference between the third congestion threshold and the congestion threshold of the first data stream before calculating the third congestion threshold, and the target credit amount indicates the data volume size of a packet in the first data stream.

[0257] In a possible implementation, the receiving unit 902 is further configured to receive third feedback information, where the third feedback information is the feedback information of the second packet among the multiple packets, and the third feedback information includes information indicating that the second packet is lost, and the second packet is sent through the second path among the n paths;

[0258] The processing unit 903 is further configured to increase the credit amount of the second path by a target credit amount, where the target credit amount indicates the data volume size of a packet in the first data stream.

[0259] In a possible implementation, when the network device receives feedback information of a packet with a target sequence number, where the target sequence number includes all sequence numbers of the packets sent through the n paths, the processing unit 903 is further configured to:

[0260] When the congestion threshold of the first data stream is greater than the actual credit amount, adjust the congestion threshold of the first data stream to a value equal to the actual credit amount, where the actual credit amount is the sum of the credit amounts of the n paths and the remaining credit amount.

[0261] In a possible implementation, the processing unit 903 is further configured to:

[0262] Map the source port numbers of the packets in the first data stream to n virtual port numbers, where the n virtual port numbers correspond to the n paths one by one.

[0263] Figure 9 For the specific operations and beneficial effects of each unit in the device 900 shown, reference can be made to the corresponding descriptions in the above Figure 3 and its possible method embodiments, which will not be elaborated here.

[0264] Figure 10 Shown is a possible schematic hardware structure of the device according to an embodiment of the present application. The device may be the first network device described in the above embodiment. The device 1000 includes: a processor 1001, a memory 1002, and a communication interface 1003. The processor 1001, the communication interface 1003, and the memory 1002 may be connected to each other or connected to each other through a bus 1004.

[0265] Exemplarily, the memory 1002 is used to store the computer programs and data of the device 1000. The memory 1002 may include but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM), etc.

[0266] The communication interface 1003 includes a sending interface and a receiving interface. The number of communication interfaces 1003 may be multiple, and is used to support the device 1000 to communicate, such as receiving or sending data or messages, etc.

[0267] Exemplarily, the processor 1001 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, transistor logic device, hardware component, or any combination thereof. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The processor 1001 may be used to read the program stored in the above-mentioned memory 1002, so that the device 1000 performs the operations performed by the first network device in any one of the methods described in the above Figure 3 and its possible embodiments.

[0268] In a possible implementation manner, the processor 1001 may be used to read the program stored in the above-mentioned memory 1002 and perform the following operations:

[0269] Sending multiple packets of a first data stream through an n number of paths via a sending interface, where each of the n paths is configured with a credit amount, the credit amount indicating the size of the data sending capacity of each path, the sum of the credit amounts of the n paths being less than or equal to a first congestion threshold, the first congestion threshold being the congestion threshold of the first data stream, and n being an integer greater than 1;

[0270] Receiving first feedback information through a receiving interface, the first feedback information being the feedback information of a first packet, the first feedback information including indication information on whether the first path is congested, the first path being the path used to send the first packet among the n paths; the first packet being any one of the multiple packets of the first data stream;

[0271] Re-determining the credit amount of the first path based on the indication information; and re-determining the load amount of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re-determined load amount of the first path is reduced.

[0272] Figure 10 For the specific operations and beneficial effects of each unit in the device 1000 shown, reference may be made to the corresponding descriptions in the above Figure 3 and its possible method embodiments, which will not be elaborated here.

[0273] The embodiments of the present application further provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the operations performed by the first network device in any one of the methods described in the above Figure 3 and its possible method embodiments.

[0274] The embodiments of the present application further provide a computer program product. When the computer program product is read and executed by a computer, the aboveFigure 3 The operations performed by the first network device in the method described in any one of the embodiments and their possible method embodiments will be executed.

[0275] In summary, in the embodiments of the present application, by splitting a single flow load across multiple paths for transmission, fine-grained load balancing is achieved, bandwidth competition at the flow level is reduced, and the load balancing performance of the entire network is improved. Additionally, by maintaining a congestion threshold for a single flow, allocating the amount of credit included in the congestion threshold to multiple transmission paths of the flow, and adaptively adjusting the credit amount of the corresponding path based on the feedback information of the packets transmitted on each path, and then adjusting the load amount, that is, by combining flow-level congestion control and path-level credit management, the congestion status of multiple paths is perceived, and the dynamic load of the paths is adjusted according to the congestion status of the paths, achieving the optimal load ratio among multiple paths, eliminating the shortcoming effect caused by congestion on a specific path, obtaining an increase in the overall throughput rate, and reducing network congestion.

[0276] In the present application, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same function and role. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples, the first node can be referred to as the second node, and similarly, the second node can be referred to as the first node. The first node and the second node can both be nodes, and in some cases, they can be separate and different nodes.

[0277] It should also be understood that in each of the embodiments of the embodiments of the present application, the magnitude of the serial number of each process does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0278] It should also be understood that the term "comprising" (also referred to as "includes", "including", "comprises", and / or "comprising") when used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.

[0279] It should also be understood that the "one embodiment", "an embodiment", and "a possible implementation" mentioned throughout the specification mean that the specific features, structures, or characteristics related to the embodiment or implementation are included in at least one embodiment of the embodiments of the present application. Therefore, the "in one embodiment" or "in an embodiment", "a possible implementation" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.

[0280] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A network congestion control method, characterized in that, the method includes: The source end sends multiple packets of the first data stream through n paths, and each path in the n paths is configured with a credit amount, the credit amount indicating the size of the data sending capacity of each path. The sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, and the first congestion threshold is the congestion threshold of the first data stream. n is an integer greater than 1; The source end receives first feedback information, which is the feedback information of the first packet. The first feedback information includes indication information on whether the first path is congested. The first path is the path used to send the first packet among the n paths; the first packet is any one of the multiple packets; The source end re-determines the credit amount of the first path based on the indication information; The source end re-determines the load amount of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re-determined load amount of the first path decreases; The source end re-determines the credit amount of the first path based on the indication information, including: The source end calculates a second congestion threshold based on the first feedback information, and the second congestion threshold is the new congestion threshold of the first data stream; The source end adjusts the credit amount of the first path based on a first difference, and the first difference is the difference between the second congestion threshold and the congestion threshold of the first data stream before calculating the second congestion threshold; The sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold; The source end adjusts the credit amount of the first path based on the first difference, including: When the first difference is greater than zero and the sum of the first difference and the remaining credit amount is greater than the target credit amount, the source end increases the credit amount of the first path by two target credit amounts, and the target credit amount indicates the data amount size of a packet in the first data stream; or, When the first difference is less than zero, the source end reduces the credit amount of the first path by a first credit amount, and the first credit amount is the absolute value of the sum of the first difference and the target credit amount.

2. The method according to claim 1, characterized in that, The first feedback information is the information in a feedback packet received by the source end, and the feedback packet includes the feedback information of m packets among the multiple packets, and m is an integer greater than 1; The feedback information of the m packets includes indication information on whether the transmission path of each packet is congested during the transmission of the m packets.

3. The method according to claim 1, characterized in that, The first feedback information further includes indication information on out-of-order of the first packet at the destination end; each time the first path sends a packet, it consumes a target credit amount of the credit amount in the first path, and the target credit amount indicates the data amount size of a packet in the first data stream; The source end re-determines the credit amount of the first path based on the first feedback information, including: The source end maintains the current credit amount of the first path based on the indication information that the first packet is out-of-order at the destination end.

4. The method according to any one of claims 1-3, wherein, the sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold; the method further includes: when the credit amount of the first path drops to zero, the source end sends a probe packet through the first path; the source end receives the second feedback information of the probe packet; when the second feedback information indicates that the first path is not congested, the source end calculates a third congestion threshold based on the second feedback information, and the third congestion threshold is the new congestion threshold of the first data stream; when the sum of the second difference and the remaining credit amount is greater than the target credit amount, the source end increases the credit amount of the first path by the target credit amount; the second difference is the difference between the third congestion threshold and the congestion threshold of the first data stream before calculating the third congestion threshold, and the target credit amount indicates the data volume size of a packet in the first data stream.

5. The method according to any one of claims 1-3, wherein, the method further includes: the source end receives third feedback information, and the third feedback information is the feedback information of the second packet among the multiple packets, and the third feedback information includes information indicating that the second packet is lost, and the second packet is sent through the second path among the n paths; the source end increases the credit amount of the second path by the target credit amount, and the target credit amount indicates the data volume size of a packet in the first data stream.

6. The method according to any one of claims 1-3, wherein, the source end receives the feedback information of the packet with the target sequence number, and the target sequence number includes all the sequence numbers of the packets sent through the n paths, and the method further includes: when the congestion threshold of the first data stream is greater than the actual credit amount, the source end adjusts the congestion threshold of the first data stream to a value equal to the actual credit amount, and the actual credit amount is the sum of the credit amounts and the remaining credit amounts of the n paths.

7. The method according to any one of claims 1-3, wherein, the method further includes: the source end maps the source port numbers of the packets of the first data stream to n virtual port numbers, and the n virtual port numbers correspond to the n paths one by one.

8. A network device, wherein, the device includes: a sending unit, configured to send multiple packets of a first data stream through n paths, and each path among the n paths is configured with a credit amount, and the credit amount indicates the size of the data sending capacity of each path, and the sum of the credit amounts of the n paths is less than or equal to a first congestion threshold, and the first congestion threshold is the congestion threshold of the first data stream, and n is an integer greater than 1; A receiving unit, configured to receive first feedback information, where the first feedback information is feedback information of a first message, and the first feedback information includes indication information on whether a first path is congested. The first path is a path used to send the first message among the n paths; the first message is any one of the multiple messages; A processing unit, configured to re-determine the credit amount of the first path based on the indication information; The processing unit is further configured to re-determine the load amount of the first path based on the credit amount of the first path. When the indication information indicates that the first path is congested, the re-determined load amount of the first path is reduced; Specifically, the processing unit is configured to: Calculate a second congestion threshold based on the first feedback information, where the second congestion threshold is a new congestion threshold of the first data stream; Adjust the credit amount of the first path based on a first difference, where the first difference is the difference between the second congestion threshold and the congestion threshold of the first data stream before calculating the second congestion threshold; The sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold; Specifically, the processing unit is configured to: When the first difference is greater than zero and the sum of the first difference and the remaining credit amount is greater than a target credit amount, increase the credit amount of the first path by two target credit amounts, where the target credit amount indicates the data volume size of a message in the first data stream; or, When the first difference is less than zero, reduce the credit amount of the first path by a first credit amount, where the first credit amount is the absolute value of the sum of the first difference and the target credit amount; 9. The apparatus according to claim 8, wherein, The first feedback information is information in a feedback message received by the network apparatus, and the feedback message includes feedback information of m messages among the multiple messages, where m is an integer greater than 1; The feedback information of the m messages includes indication information on whether the transmission path of each message is congested during the transmission of the m messages.

10. The apparatus according to claim 8, wherein, The first feedback information further includes indication information on out-of-order of the first message at the destination end; each time the first path sends a message, it consumes a target credit amount of the credit amount in the first path, and the target credit amount indicates the data volume size of a message in the first data stream; Specifically, the processing unit is configured to maintain the current credit amount of the first path based on the indication information on out-of-order of the first message at the destination end.

11. The apparatus according to any one of claims 8-10, wherein, The sum of the credit amounts and the remaining credit amounts of the n paths is equal to the first congestion threshold; The sending unit is further configured to send a probe message through the first path when the credit amount of the first path drops to zero; The receiving unit is further configured to receive second feedback information of the probe message; The processing unit is further configured to calculate a third congestion threshold based on the second feedback information when the second feedback information indicates that there is no congestion on the first path, where the third congestion threshold is the new congestion threshold of the first data stream; The processing unit is further configured to increase the credit of the first path by a target credit amount when the sum of the second difference and the remaining credit amount is greater than the target credit amount; The second difference is the difference between the third congestion threshold and the congestion threshold of the first data stream before calculating the third congestion threshold, and the target credit amount indicates the data volume size of a packet in the first data stream.

12. The apparatus according to any one of claims 8-10, wherein, The receiving unit is further configured to receive third feedback information, where the third feedback information is the feedback information of the second packet among the multiple packets, and the third feedback information includes information indicating packet loss of the second packet, and the second packet is sent through the second path among the n paths; The processing unit is further configured to increase the credit of the second path by a target credit amount, where the target credit amount indicates the data volume size of a packet in the first data stream.

13. The apparatus according to any one of claims 8-10, wherein, The network device receives feedback information of a packet with a target sequence number, where the target sequence number includes all sequence numbers of the packets sent through the n paths, and the processing unit is further configured to: When the congestion threshold of the first data stream is greater than the actual credit amount, adjust the congestion threshold of the first data stream to a value equal to the actual credit amount, where the actual credit amount is the sum of the credits of the n paths and the remaining credit amount.

14. The apparatus according to any one of claims 8-10, wherein, The processing unit is further configured to: Map the source port number of the packets of the first data stream to n virtual port numbers, where the n virtual port numbers correspond to the n paths one by one.

15. A network device, wherein, The network device includes a processor, a sending interface, a receiving interface, and a memory; wherein, the memory is used to store a computer program, the sending interface is used to send information, the receiving interface is used to receive information, and the processor is used to execute the computer program stored in the memory, so that the network device executes the method according to any one of claims 1-7.

16. A data transmission system, wherein, The system includes a first network device and a second network device, the first network device is the network device according to any one of claims 8-14, or the network device according to claim 15, and the second network device is used to receive the packets sent by the first network device.

17. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Routing method and equipment based on load balancing

    CN105634973A

  • Message transmission method and terminal, network equipment and communication system

    CN108270682A