Congestion control method and corresponding device
By detecting the available bandwidth first, and then sending data packets, the lag and resource waste problems of RDMA network congestion control are solved, and efficient congestion control and resource optimization are achieved.
Patent Information
- Application Number
- PCT/CN2024/110313
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-24
AI Technical Summary
The existing RDMA network congestion control methods have lag and resource waste problems, especially when the queue logarithm increases, detection bandwidth resource expansion, resulting in limited hardware resources and frequent network congestion.
The method of detecting messages first detecting the available bandwidth and then sending data messages is adopted. A set of detection resources provides available bandwidth for multiple queues, reducing resource waste and reducing the probability of congestion.
Effectively control congestion, improve service stability, reduce detection resource use, reduce IOPS requirements, and optimize bandwidth utilization.
Smart Images

Figure CN2024110313_24072025_PF_FP_ABST
Abstract
Description
A congestion control method and corresponding device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 17, 2024, with application number 202410071325.2 and application name “A method and corresponding device for congestion control”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communication technology, and in particular to a congestion control method and corresponding device. Background Art
[0003] Traditional remote direct memory access (RDMA) congestion control uses the switch's explicit congestion notification (ECN) mark as a congestion signal. The sender speeds up or slows down data transmission based on whether the data packet's acknowledgment (ACK) packet carries the ECN mark. This congestion control method has a certain degree of hysteresis.
[0004] To address the lag in RDMA network congestion control, we can pre-send probe packets to detect network transmission resources and then send data packets based on the detection results. This "probe first, send later" scheduling mode reduces the probability of congestion from the source.
[0005] Currently, when using probe packets to detect bandwidth available for data packet transmission, the bandwidth allocated for probes is allocated based on queue pairs (QPs). This causes the bandwidth used for probes to scale linearly with the number of QPs. Due to limited hardware resources, this congestion control method wastes a significant amount of bandwidth when the number of QPs increases.
[0006] Summary of the Invention
[0007] The present application provides a congestion control method for detecting available bandwidth for data packets in multiple queues using fewer detection resources, thereby effectively controlling congestion. The present application also provides a corresponding device, a computer-readable storage medium, and a computer program product.
[0008] In a first aspect, the present application provides a congestion control method, which is applied to a source device, the source device including multiple queues, and the method including: sending multiple probe messages, the multiple probe messages being used to detect the allocated bandwidth on a target path between the source device and a destination device for the multiple queues; receiving multiple responses from the destination device, the multiple responses being used to indicate the available bandwidth detected by the multiple probe messages, the available bandwidth being the bandwidth in the allocated bandwidth that can be used to transmit data messages in the multiple queues; and sending the data messages in the multiple queues to the destination device using the available bandwidth.
[0009] In this application, the source device may be a network card, a data processing unit (DPU), a just a bunch of flash drives (JBOF), a storage box, a computing box, or other devices with network communication capabilities. The source device may also be a computer device equipped with the above-mentioned network card, DPU, etc.
[0010] In the present application, a queue is used to store data packets, and data packets sent from a source device are dispatched from the queue.
[0011] In this application, the detection message may also be referred to as a "probe" or a "control message".
[0012] In this application, on a path, the bandwidth used to transmit data packets and the bandwidth used to transmit probe packets are generally set proportionally. This ratio can be the ratio of the length of the data packet to the length of the probe packet. For example, if the ratio of the length of the data packet to the length of the probe packet is 95:5, and the allocated bandwidth of the target path is 100 megabits (M), then 95M bandwidth can be used to transmit data packets, and 5M bandwidth can be used to transmit probe packets.
[0013] In this application, when a destination device receives a probe message via a target path, it will send an acknowledgment (ACK) on the target path. Thus, over a period of time, by counting the number of probe messages sent and the number of acknowledgments received, the available bandwidth for probe messages on the target path can be determined. Then, based on the ratio of the bandwidth of data messages to the bandwidth of probe messages, the available bandwidth for data messages on the target path can be determined. Data messages can then be sent to the destination device based on this available bandwidth, preventing congestion and achieving active congestion control.
[0014] Because the congestion situation of the same target path is the same when transmitting probe messages and data messages, and even if the probe message is discarded by the device during transmission, it will not affect the service. Therefore, the available bandwidth on the target path is first detected by the probe message, and then the data message is sent, which replaces the direct competition for bandwidth by the data message, thus improving the stability of the service. Moreover, by using a set of probe message resources for detection in the first aspect above, data messages in multiple queues can be sent, and there is no need to configure a set of resources for detection for each queue, which reduces the resources used for detection and reduces the probability of congestion on the target path from the source. In addition, there is no need to send a large number of probe messages, which also reduces the requirements for (input / output operations per second, IOPS).
[0015] In a possible implementation, the multiple detection messages do not include a queue identifier of any queue among the multiple queues.
[0016] In this possible implementation, because a set of detection resources is used to detect available bandwidth for multiple queues, the destination device does not need to distinguish the queues corresponding to the detection messages. Therefore, the detection messages do not need to carry queue identifiers, shortening the length of the detection messages and thus reducing the bandwidth used by the detection messages.
[0017] In one possible implementation, before sending data packets in multiple queues to a destination device using available bandwidth, the method further includes: generating multiple tokens based on a first relationship and multiple responses, and the multiple tokens are used to schedule data packets in multiple queues, where the first relationship is a ratio between the amount of responses received and the amount of data packets sent.
[0018] In this possible implementation, if the ratio of the number of detection messages received to the number of data messages sent is 1:2, the source device will generate two tokens for each response it receives. By scheduling the data messages in the queue through tokens, the sending of data messages in multiple queues can be effectively controlled.
[0019] In a possible implementation, the increment of tokens corresponds to the number of received responses, and the consumption of tokens corresponds to the number of data packets dispatched from the multiple queues.
[0020] In a possible implementation, the method further includes: determining a packet loss rate of the probe messages according to the number of sent probe messages and the number of received responses, where the packet loss rate is used to adjust a sending rate of the probe messages.
[0021] In this possible implementation, the sending rate of the probe message is adjusted according to the packet loss rate, so as to improve the detection quality of the probe message on the target path.
[0022] In a possible implementation, the packet loss rate of the probe message may be determined by a sequence number in a received response, where the sequence number in the response is the same as the sequence number of the probe message corresponding to the response.
[0023] In one possible implementation, there are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include a target path; the method also includes: counting the delay of each transmission path in the multiple transmission paths; and adjusting the allocated bandwidth of the target path according to the delay of each transmission path.
[0024] In this possible implementation, when there are multiple transmission paths between the source device and the destination device, the allocated bandwidth of each transmission path can be adjusted based on the total bandwidth of the multiple transmission paths according to the delay of each transmission path. Such flexible adjustment can improve the utilization of bandwidth resources.
[0025] In a possible implementation, the source device may further determine the order of the multiple transmission paths according to the transmission delays of the multiple transmission paths, and then determine the bandwidth adjustment amount of each transmission path according to the order.
[0026] In a possible implementation, the source device may further determine the transmission path that needs to be optimized according to the order of the multiple transmission paths.
[0027] In a possible implementation, the source device may further determine a transmission path for reselection according to the order of the multiple transmission paths.
[0028] In one possible implementation, the source device can be configured with at least one probe management block (PMB), each corresponding to a different destination device. Each PMB can manage steps such as sending probe messages, receiving responses, and scheduling data messages from the source device to the corresponding destination device.
[0029] A second aspect of the present application provides a network device, comprising: a transceiver unit and a processing unit; wherein,
[0030] The transceiver unit is used to:
[0031] sending a plurality of probe messages, the plurality of probe messages being used to detect allocated bandwidth on a target path between a source device and a destination device for a plurality of queues;
[0032] receiving a plurality of responses from the destination device, the plurality of responses being used to indicate available bandwidth detected by the plurality of probe messages, the available bandwidth being bandwidth available in the allocated bandwidth for transmitting the data messages in the plurality of queues;
[0033] The data packets in the multiple queues are sent to the destination device using the available bandwidth.
[0034] In a possible implementation, the multiple detection messages do not include a queue identifier of any queue among the multiple queues.
[0035] In one possible implementation, the processing unit is used to generate multiple tokens based on a first relationship and multiple responses, and the multiple tokens are used to schedule data packets in multiple queues. The first relationship is a ratio between the amount of responses received and the amount of data packets sent.
[0036] In a possible implementation, the increment of tokens corresponds to the number of received responses, and the consumption of tokens corresponds to the number of data packets dispatched from the multiple queues.
[0037] In a possible implementation, the processing unit is further configured to determine a packet loss rate of the probe messages according to the number of sent probe messages and the number of received responses, and the packet loss rate is used to adjust a sending rate of the probe messages.
[0038] In a possible implementation, the packet loss rate of the probe message may be determined by a sequence number in a received response, where the sequence number in the response is the same as the sequence number of the probe message corresponding to the response.
[0039] In one possible implementation, the processing unit is further configured to, when there are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include a target path, calculate the delay of each of the multiple transmission paths; and adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0040] In a possible implementation, the processing unit is further configured to determine an order of the multiple transmission paths according to transmission delays of the multiple transmission paths, and determine a bandwidth adjustment amount for each transmission path according to the order.
[0041] In a possible implementation, the processing unit is further configured to determine a transmission path that needs to be optimized according to the order of the multiple transmission paths.
[0042] In a possible implementation manner, the processing unit is further configured to determine a transmission path for reselection according to an order of the multiple transmission paths.
[0043] The third aspect of the present application provides a network device, a communication interface, a processor and a memory, wherein the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the feature screening device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0044] The fourth aspect of the present application provides a chip system, which includes one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected by lines; the interface circuit is used to receive signals from the memory of the cloud device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the feature screening device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0045] A fifth aspect of the present application provides a computer device, comprising the network device in the above-mentioned second aspect, third aspect or any possible implementation of the second aspect.
[0046] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction runs on a computer device, the computer device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0047] In a seventh aspect, the present application provides a computer device program product, which includes a computer device program code. When the computer device program code is executed on a computer device, the computer device executes the method in the aforementioned first aspect or any possible implementation of the first aspect.
[0048] In an eighth aspect, the present application provides a congestion control system, comprising a source device and at least one destination device, wherein the source device is configured to execute the method of the first aspect or any possible implementation of the first aspect. The destination device is configured to send a response to the source device upon receiving a probe message.
[0049] Among them, the technical effects brought about by the second to eighth aspects or any possible implementation methods thereof can refer to the technical effects brought about by the first aspect or different possible implementation methods of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] FIG1A is a schematic diagram of an architecture of a data center provided in an embodiment of the present application;
[0051] FIG1B is another schematic diagram of the architecture of a data center provided in an embodiment of the present application;
[0052] FIG2 is a schematic diagram of an embodiment of a congestion control method provided in an embodiment of the present application;
[0053] FIG3 is a schematic diagram of another embodiment of a congestion control method provided in an embodiment of the present application;
[0054] FIG4A is a schematic diagram showing a relationship between a probe management block PMB and a destination device according to an embodiment of the present application;
[0055] FIG4B is another schematic diagram of the relationship between the PMB and the destination device provided in an embodiment of the present application;
[0056] FIG5 is a schematic diagram of an example of token management provided in an embodiment of the present application;
[0057] FIG6 is a schematic diagram of an example of multiple transmission paths provided in an embodiment of the present application;
[0058] FIG7 is a schematic diagram of an example of a cloud computing scenario provided by an embodiment of the present application;
[0059] FIG8 is a schematic diagram of an example of an artificial intelligence scenario provided by an embodiment of the present application;
[0060] FIG9 is a schematic structural diagram of a network device provided in an embodiment of the present application;
[0061] FIG10 is a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0063] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0064] Embodiments of the present application provide a congestion control method for detecting available bandwidth for data packets in multiple queues using fewer detection resources, thereby effectively controlling congestion. This application also provides corresponding apparatus, computer-readable storage media, and computer program products. These are described in detail below.
[0065] To facilitate understanding, the following briefly introduces the technical terms involved in the embodiments of this application:
[0066] 1. Remote direct memory access (RDMA): This is a new network communication technology developed to address server-side data processing delays during network transmission. RDMA transfers data directly to a computer's storage area over the network, rapidly moving data from one system to a remote system's memory without disrupting the operating system. This reduces the need for computer processing power, eliminates the overhead of external memory copying and context switching, and offers low latency and high bandwidth.
[0067] 2. Proactive congestion control algorithm (PCC): This is a common signaling-based network congestion control algorithm for the RDMA protocol. It is often used for network congestion control on the RDMA protocol. It can first detect available bandwidth by sending probe packets and then sending data packets based on the available bandwidth, reducing the probability of congestion at the source.
[0068] 3. Delay: refers to the time required for a message or packet to be transmitted from the source end to the destination end of the network.
[0069] 4. Packet loss rate (loss tolerance or packet loss rate): refers to the ratio of the number of lost data packets in the test to the number of sent data packets.
[0070] 5. Queue Pair (QP): It is the basic communication unit of the RDMA protocol reliable connection mode.
[0071] 6. Probe message: also known as a probe or control message, used to detect network load or round-trip time (RTT).
[0072] The congestion control method provided in the embodiment of the present application can be applied to a data center, and the architecture of the data center can be understood by referring to Figure 1A. As shown in Figure 1A, the data center may include multiple data center distribution points (PODs). POD is the smallest business unit of the data center and is a resource collection of switches, firewalls, and servers. Each POD can include a computer device cluster (such as a server cluster), and different computer devices can communicate through switches. In order to meet the connection requirements between a large number of computer devices in a data center, a multi-layer switch is usually set up.
[0073] As shown in Figure 1A, computer devices can first connect through access switches, which are then connected to different aggregation switches. The aggregation switches are then connected to different core switches. This allows communication between different computer devices in the data center. The access switches and aggregation switches in different PODs do not need to be shared; computer devices in different PODs communicate through the core switch.
[0074] The data center architecture can also be understood by referring to Figure 1B. As shown in Figure 1B, a data center may include multiple PODs, which may include storage PODs and compute PODs. The computing devices in the storage PODs are primarily storage servers, used to store data, while the computing devices in the compute PODs are primarily compute servers, used to complete computing tasks.
[0075] Compute servers may need to read data from storage servers while performing computational tasks, and after the computation is complete, they may also need to store the results in storage servers. Furthermore, when multiple compute servers jointly perform a computational task, data may need to be transferred between them, requiring inter-server communication. Data transfers may also occur between storage servers, requiring inter-server communication. Therefore, in a data center, whether a compute server or a storage server, each server acting as a source can have multiple destinations, and the source and destination can be located in the same pod or in different pods.
[0076] If the source and destination ends in the same POD are connected through the same access switch, they can communicate directly through the access switch. If the source and destination ends in the same POD are not connected to the same access switch, they can communicate through the aggregation switch that is higher than the access switch.
[0077] The source and destination ends in different PODs can communicate through the access switches, aggregation switches in their respective PODs, and the core switches located above the PODs.
[0078] Access switches can be divided into three types based on their integration with computer equipment: end-of-row (EOR) switches, middle-of-row (MOR) switches, and top-of-rack (TOR) switches. These types of switches can also be referred to as cabinet switches.
[0079] Among them, the integrated architecture of the EOR switch is to centrally install the access switches in the cabinet (switch cabinet) at the end of a row of cabinets. The hosts, servers and minicomputers in the equipment cabinet are connected by permanent links through horizontal cables.
[0080] The integrated architecture of the MOR switch allows cables to be routed from the cabinet in the middle to both ends, reducing cable congestion at the entrances and exits of the routing channel and shortening the average cable length.
[0081] The integrated architecture of TOR switches deploys one or two access switches at the top of each cabinet in a POD (Point of View) (POD). Servers are connected to the cabinets via patch cables. The uplink ports on the access switches are connected to aggregation switches or core switches via copper or fiber cables.
[0082] The access switch shown in Figure 1B is a TOR switch. The TOR switch is first connected to the aggregation switch, which is then connected to the core switch. If the data center is small, the TOR switch can also be directly connected to the core switch without configuring an aggregation switch.
[0083] In addition, it should be noted that Figures 1A and 1B above are based on a single data center as an example. In reality, devices in different data centers can also communicate with each other. That is, the source and destination can be located in the same region or in different regions. Furthermore, the source and destination can be located in the same availability zone (AZ) or in different AZs, with each AZ including one data center or multiple geographically close data centers. Typically, a region can include multiple AZs.
[0084] The source and destination can be located in the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is located within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC. Interconnection between VPCs is achieved through the communication gateway.
[0085] In the embodiment of the present application, the source and the destination can communicate using RDMA, but because the bandwidth resources between the source and the destination are limited by hardware capabilities, network congestion may occur when transmitting data packets between the source and the destination.
[0086] In order to control network congestion between the source and destination, active congestion control can be used. Probe messages are sent on the transmission path between the source and destination to first detect the available bandwidth for sending data messages, and then send data messages based on the available bandwidth. In this way, the probability of network congestion can be reduced from the source.
[0087] The following uses FIG2 as an example to introduce the process of active congestion control.
[0088] As shown in Figure 2, the transmission path from the source to the destination can involve one or more switches. Because different switches may have different loads, some switches may carry large amounts of data. Alternatively, some switches may have limited hardware capabilities, creating transmission bottlenecks between the source and the destination. The available bandwidth along a transmission path is often determined by the bottleneck. Therefore, the source can first send a probe message to determine available bandwidth before sending data packets.
[0089] On a transmission path, the bandwidth used to transmit data packets and the bandwidth used to transmit probe packets are typically set proportionally. This ratio can be the ratio of the length of the data packet to the length of the probe packet. For example, if the ratio of the length of the data packet to the length of the probe packet is 95:5, and the allocated bandwidth of the target path is 100 megabits (M), 95M of bandwidth can be used to transmit data packets, and 5M of bandwidth can be used to transmit probe packets.
[0090] The source can first send a probe message over the 5M bandwidth in the example above to detect the available bandwidth for transmitting data messages on the transmission path. During this detection process, the source can send probe messages at a relatively high rate. However, the probe messages may be rate-limited during transmission in the network, and some probe messages may be discarded. However, the destination will send an acknowledgment (ACK) to the source for each probe message it receives.
[0091] The source can count the number of probe messages sent and the number of responses received, and then calculate the available bandwidth within the 5 Mbps allocated for probe messages. This is then used to calculate the available bandwidth within the 95 Mbps allocated bandwidth. For example, if the calculated available bandwidth for probe messages is 2.5 Mbps, the ratio of the available bandwidth for probe messages (2.5 Mbps) to the allocated bandwidth (5 Mbps) is 1:2. Using 95*(1 / 2)=47.5, we can determine that the available bandwidth for transmitting data messages is 47.5 Mbps. Data messages can then be sent using this 47.5 Mbps bandwidth, thereby reducing the probability of network congestion between the source and destination.
[0092] The source and destination in Figure 2 can both be computer devices (servers), virtual machines, or containers. Of course, the source and destination can also be devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The PLDs can be complex programmable logical devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof.
[0093] The congestion control method provided in the embodiment of the present application can also be understood by referring to FIG3 .
[0094] As shown in FIG3 , an embodiment of a congestion control method provided in an embodiment of the present application includes:
[0095] 301. The source device sends multiple probe messages, where the multiple probe messages are used to detect the allocated bandwidth on the target path between the source device and the destination device for multiple queues.
[0096] In an embodiment of the present application, the source device includes multiple queues, which are used to store data packets, and the data packets sent from the source device are scheduled from the queues.
[0097] In this application, the source device may be a network card, a data processing unit (DPU), a just a bunch of flash drives (JBOF), a storage box, a computing box, or other devices with network communication capabilities. The source device may also be a computer device equipped with the above-mentioned network card, DPU, etc.
[0098] 302. The source device receives multiple responses from the destination device, where the multiple responses are used to indicate available bandwidth detected by the multiple probe messages, where the available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data messages in the multiple queues.
[0099] The determination of available bandwidth can be understood by referring to the introduction in FIG. 2 above.
[0100] 303. The source device sends data packets in multiple queues to the destination device using available bandwidth.
[0101] In this application, when a destination device receives a probe message via a target path, it sends an ACK on the target path. This allows, over a period of time, to determine the available bandwidth for probe messages on the target path by counting the number of probe messages sent and the number of ACKs received. The available bandwidth for data messages on the target path can then be determined based on the ratio of the bandwidth of data messages to the bandwidth of probe messages. Data messages can then be sent to the destination device based on this available bandwidth, preventing congestion and achieving proactive congestion control.
[0102] Because the congestion situation of the same target path is the same when transmitting probe messages and data messages, and even if the probe message is discarded by the device during transmission, it will not affect the service. Therefore, the available bandwidth on the target path is first detected by the probe message, and then the data message is sent, which replaces the direct competition for bandwidth by the data message, thus improving the stability of the service. Moreover, the solution provided by the embodiment of the present application uses a set of probe message resources for detection, so that data messages in multiple queues can be sent. There is no need to configure a set of resources for detection for each queue, which reduces the resources used for detection and reduces the probability of congestion on the target path from the source. In addition, there is no need to send a large number of probe messages, which also reduces the requirements for (input / output operations per second, IOPS).
[0103] Optionally, in an embodiment of the present application, at least one probe management block (PMB) may be configured in the source device, with each probe management block corresponding to a different destination device. Each probe management block may manage the sending of probe messages, the receiving of responses, and the scheduling of data messages from the source device to the corresponding destination device.
[0104] The relationship between PMBs and queues can be understood by referring to Figure 4A . As shown in Figure 4A , a source device includes multiple queues. For example, three queues are represented by QP1, QP2, and QP3. These three queues are used to store data packets destined for the same destination device 1. The source device includes PMB 1, which can be created when any queue, QP1, QP2, or QP3, is established. PMB 1 serves as a unified exit for probe packets destined for destination device 1 and a unified entry for responses from destination device 1.
[0105] Because PMB 1 centrally manages probe messages and responses for destination device 1, the multiple probe messages sent by PMB 1 do not include the queue identifiers of any of the multiple queues. That is, the multiple probe messages sent by PMB 1 do not need to carry the identifiers QP1, QP2, or QP3. The destination device also does not need to distinguish the queues corresponding to probe messages from the source device. It only needs to reply to probe messages from the source device, without distinguishing which queue the response is intended for. This not only shortens the length of probe messages sent by the source device, but also shortens the length of responses sent by the destination device, thereby reducing bandwidth resources consumed by probes.
[0106] The source device illustrated in Figure 4A above only shows one PMB 1. Because data centers have many nodes, each source device communicates with multiple destination devices, and therefore, multiple PMBs are present in the source device. As shown in Figure 4B , a source device communicating with N destination devices communicates with destination device 1 through PMB 1, destination device 2 through PMB 2, ..., and destination device N through PMB N. Each PMB corresponds to a different destination device. Regardless of the number of queues for data packets destined for the same destination device, the corresponding PMB manages the sending of probe messages, the receiving of acknowledgments, and the scheduling of data packets. For example, the sending of probe messages, the receiving of acknowledgments, and the scheduling of data packets for queues QP4 and QP5 are all managed by PMB 2, ..., while the sending of probe messages, the receiving of acknowledgments, and the scheduling of data packets for queues QPx and QPy are all managed by PMB N.
[0107] In addition, it should be noted that the above only uses the source device as an example to illustrate the relationship between the PMB and the queue. In fact, each computer device in the data center or the network card in the computer device can be used as a source device to communicate with other computer devices or network cards.
[0108] In the embodiment of the present application, the PMB may schedule data packets in a queue by distributing tokens to the queue. Typically, one token may schedule one data packet.
[0109] For more information about token management, refer to Figure 5. As shown in Figure 5, PMB 1 manages three queues, QP1, QP2, and QP3. After a source device sends a probe message, it receives a response. Upon receiving the response, PMB 1 generates tokens based on a first relationship between the response and the response. The first relationship is the ratio of received responses to sent data messages. This ratio can be pre-configured or dynamically determined. For example, if the ratio is 1:2, PMB 1 generates two tokens for each response received by the source device.
[0110] PMB 1 then distributes tokens to the queues to send data packets. PMB 1 can distribute tokens to the queues one by one. For example, after all the data packets in queue QP1 are scheduled, tokens are distributed to queue QP2. After all the data packets in queue QP2 are scheduled, tokens are distributed to queue QP3.
[0111] Of course, PMB 1 may distribute tokens to the queues in a round-robin manner, such as distributing one token to each of the QP1 queue, the QP2 queue, and the QP3 queue, and then starting the next round of distribution. This application does not limit this.
[0112] Tokens are maintained in a credit pool. The number of tokens in the credit pool increases with the number of responses received, and the number of tokens consumed corresponds to the number of data packets dispatched from multiple queues. For example, if the ratio is 1:2, receiving a response will increase the credit pool by two tokens. If a token is used to dispatch a data packet, sending a data packet will consume one token from the credit pool.
[0113] Optionally, in an embodiment of the present application, a packet loss rate of the probe message can be determined based on the number of probe messages sent and the number of responses received. The packet loss rate is used to adjust the sending rate of the probe message. By adjusting the sending rate of the probe message based on the packet loss rate, the quality of the probe message's detection of the target path can be improved.
[0114] Statistics on packet loss rate can be performed on a periodic basis. For example, if one period is 30 microseconds (us), the source device counts the number of probe messages sent within 30us and the number of responses received. If the number of messages sent is 100 and the number of responses received is 95, then 5 packets are lost and the packet loss rate is 5%.
[0115] Of course, the packet loss rate can also be counted in other ways, such as determining the packet loss rate by the sequence number in the received response, where the sequence number in the response is the same as the sequence number of the probe message corresponding to the response.
[0116] For example: the source device sends 100 probe messages with sequence numbers from 0 to 99. In the received response with a sequence number less than 100, the sequence numbers do not include 3, 11, 62, 78 and 95. It can be determined that during the transmission of these 100 probe messages from the source device to the destination device, the probe messages with sequence numbers 3, 11, 62, 78 and 95 were lost, and the packet loss rate can be determined to be 5%.
[0117] The above description process is based on an example of a transmission path between the source device and the destination device. In fact, there may be multiple transmission paths between the source device and the destination device. When there are multiple transmission paths, the PMB's management process of multiple transmission paths can be understood by referring to Figure 6.
[0118] As shown in Figure 6, there are four transmission paths between the source device and the destination device: path 1, path 2, path 3, and path 4. The source device can refer to the previous method for detecting the target path and send a probe message to each transmission path. Then, it receives a response from each transmission path and determines the available bandwidth of each transmission path based on the received response. If only one transmission path is needed to transmit the data message, the source device can select the path with the largest available bandwidth from the four transmission paths to send the data message.
[0119] In addition, the source device may also calculate the delay of each transmission path among the multiple transmission paths; and adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0120] After calculating the latency of the four transmission paths, the source device can sort them in ascending order of latency. For example, if the order of latency is Path 1, Path 2, Path 3, and Path 4, then Path 1 can be determined to have the lowest latency and the fastest data transmission speed. Therefore, the bandwidth of at least one of the other three paths can be reduced, and the excess bandwidth can be allocated to Path 1. For example, the bandwidth of Path 4 can be reduced by 10 Mbps and allocated to Path 1, resulting in a bandwidth of 110 Mbps for Path 1. This improves the transmission quality of data packets transmitted on Path 1.
[0121] The statistical transmission path latency can be determined using the probe message's sequence number and transmission timestamp, as well as the received response's sequence number and reception timestamp. For example, for a probe message and response with the same sequence number, the latency can be determined by subtracting the probe message's transmission time from the response's reception time. Alternatively, the transmission path latency can be determined by averaging the latency calculated from multiple probe messages and their corresponding responses along the transmission path.
[0122] In the embodiment of the present application, the paths may also be managed based on the path sorting, such as path reselection and path aging processing.
[0123] For example, if paths 1, 2, 3, and 4 are sorted in ascending order by latency, the order remains 1, 2, 3, and 4. Weighted credit can then be assigned to these four transmission paths. If each transmission path starts with a weight of 1, then based on the sorting results, the weight of path 4 with the highest latency can be lowered by one level, for example, by 0.05, resulting in an adjusted weight of 0.95. The weight of path 1 with the lowest latency can be increased by one level, for example, by 0.05, resulting in an adjusted weight of 1.05. Paths 2 and 3 can remain unchanged. Of course, the adjustment amount for each credit level can be preconfigured, and the weighting of different paths can be configured as needed.
[0124] If multiple transmission paths exist between the source and destination devices and path reselection occurs, the reselected path can be determined based on each path's transmission delay or weighted credit. For example, path 1 can be selected during reselection. If aged paths need to be processed, the path with the longest transmission delay or the smallest weighted credit can be selected for aging, such as path 4 in Figure 6.
[0125] The congestion control solution provided in the embodiments of the present application can be applied to a variety of scenarios, such as cloud computing scenarios or artificial intelligence (AI) scenarios, which are introduced below.
[0126] In a cloud computing scenario, a network center with 10K nodes may be built according to its scale, and the servers may be deployed in PODs and AZs. As shown in Figure 7, in the cloud system, a terminal device 01, a scheduling node 02, a working node cluster 03, and a working node cluster 04 may be included. Communication connections can be established between the terminal device 01 and the scheduling node 02, between the scheduling node 02 and the working node cluster 03, between the scheduling node 02 and the working node cluster 04, and between the working node cluster 03 and the working node cluster 04. For example, a communication connection can be established between the terminal device 01 and the scheduling node 02 through the network. Optionally, the network can be a local area network, the Internet, or other networks. In a cloud computing scenario, the network can also be a virtual private cloud (VPC), which is not limited in the embodiments of the present application.
[0127] Among them, worker node cluster 03 and worker node cluster 04 can be deployed in different PODs or different AZs. Although not shown in Figure 7, it is understandable that communication between terminal device 01 and scheduling node 02, between scheduling node 02 and worker node cluster 03, between scheduling node 02 and worker node cluster 04, and between worker node cluster 03 and worker node cluster 04 can be achieved through switches. For the connection between the switch and the server, please refer to Figure 1B for an understanding.
[0128] In this implementation, cloud system operators can interact with scheduling node 02 via terminal device 01. For example, operators can send cloud system deployment instructions to scheduling node 02 via terminal device 01, instructing scheduling node 02 to deploy the cloud system based on the resources of scheduling node 02, worker node cluster 03, and worker node cluster 04. The cloud system manages the resources of worker node cluster 03 and worker node cluster 04 and provides cloud services to tenants based on the resources of worker node cluster 03 and worker node cluster 04.
[0129] Optionally, the terminal device 01, also referred to as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., is a device that includes wireless communication functionality (providing voice / data connectivity to users), for example, a handheld device with wireless connection functionality. Currently, some examples of terminal devices include: mobile phones, tablet computers, laptops, PDAs, notebook computers, wireless routers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving cars, wireless terminals in the Internet of Vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes.
[0130] The scheduling node 02 can be a cloud server or a cloud physical machine, or a cloud server cluster or a physical machine cluster composed of several cloud servers, or a cloud computing service center. The functions of the scheduling node can be implemented by software or hardware.
[0131] Worker node cluster 03 and worker node cluster 04 can be a server cluster consisting of several servers, or a cloud computing service center. A large number of basic resources owned by the cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources, and network resources are deployed in the cloud computing service center.
[0132] When the cloud system shown in Figure 7 above runs RDMA services, the link bandwidth is generally between 100Gbps, 200Gbps, and 400Gbps, and may subsequently increase to 800Gbps. The single queue pair (QP) capacity of the product network card chip is around 50Gbps-80Gbps, and business applications will experience a 10%-20% bandwidth loss. Therefore, if a 200Gbps link is used for networking, two working nodes need to establish at least 3-5 QPs to ensure bandwidth capacity, and 4 QPs are required. At this time, if 10,000 working nodes are fully interconnected, each working node needs to maintain 4 (QP) * 10K = 40,000 QPs (not all of which are active). For ease of description, a fixed number of active QPs is used (taking 25 working nodes and a total of 100 QPs as an example), but the actual number of active QPs will be much higher.
[0133] If the bandwidth requirement of an active QP exceeds the 200Gbps link bandwidth, assuming a maximum transmission unit (MTU) of 1KB, the PMB adopts a 2:1 data packet:probe packet management model, with probe packets managed uniformly by destination. Therefore, each working node needs to maintain 24 probe message flows, each destined for the other 24 working nodes. Each probe message flow sends probe messages at an initial rate of 200Gbps * 2.5%. After rate limiting and fair competition on the network switch, a portion of the probe messages reaches the destination. Upon receiving the probe message, the destination returns an ACK for that probe message. Upon receiving the probe message from the destination, the source node's PMB generates tokens based on the number of ACKs received, in a 1:2 ratio. The PMB periodically distributes available tokens to the active QPs at that destination. The PMB also records the timestamp of the ACK reception and calculates the probe message loss rate based on the ACK sequence number. Then, the sending rate is adjusted according to the packet loss rate, and the probe message is continued to be sent at the adjusted rate.
[0134] When a QP completes sending its data packets, it has no new data to transmit. However, it continues sending probe packets to the same destination until the last QP completes its service. Therefore, some tokens need to be reallocated to other QPs with the same destination.
[0135] The solution provided in the embodiments of the present application can also be applied to AI scenarios. In order to complete large model training faster, AI scenarios require higher bandwidth and lower latency. In this case, the multipath mode is a more efficient transmission mode.
[0136] As shown in Figure 8, in an AI scenario, an AI model is trained jointly by Server A and Server B. Server A performs the first phase of training on the AI model, and the resulting data needs to be transmitted to Server B. Server B then uses the data transmitted from Server A to perform the second phase of training on the AI model. To increase the transmission rate, four transmission paths can be configured between Server A and Server B. These four transmission paths can be understood by referring to Figure 6. Server A or Server B can manage these four transmission paths. The specific management process can be understood by referring to the previous description.
[0137] The above describes the congestion control method provided by the embodiment of the present application. The following describes the network device in the embodiment of the present application with reference to the accompanying drawings.
[0138] As shown in FIG9 , the network device 90 provided in the embodiment of the present application includes: a transceiver unit 901 and a processing unit 902; wherein,
[0139] The transceiver unit 901 is used for:
[0140] sending a plurality of probe messages, the plurality of probe messages being used to detect allocated bandwidth on a target path between a source device and a destination device for a plurality of queues;
[0141] receiving a plurality of responses from the destination device, the plurality of responses being used to indicate available bandwidth detected by the plurality of probe messages, the available bandwidth being bandwidth available in the allocated bandwidth for transmitting the data messages in the plurality of queues;
[0142] The data packets in the multiple queues are sent to the destination device using the available bandwidth.
[0143] The network device provided in the embodiment of the present application uses a set of probe message resources for detection, and can send data messages in multiple queues. There is no need to configure a set of resources for detection for each queue, which reduces the resources used for detection and reduces the probability of congestion on the target path from the source. There is no need to send a large number of probe messages, which also reduces the requirements for IOPS.
[0144] Optionally, the multiple detection messages do not include a queue identifier of any queue among the multiple queues.
[0145] Optionally, the processing unit 902 is used to generate multiple tokens based on the first relationship and multiple responses, and the multiple tokens are used to schedule data packets in multiple queues. The first relationship is a ratio between the amount of responses received and the amount of data packets sent.
[0146] Optionally, the increment of the token corresponds to the number of received responses, and the consumption of the token corresponds to the number of data packets scheduled from the multiple queues. Usually, one token is used to schedule one data packet.
[0147] Optionally, the processing unit 902 is further configured to determine a packet loss rate of the probe messages according to the number of sent probe messages and the number of received responses, and the packet loss rate is used to adjust a sending rate of the probe messages.
[0148] Optionally, the packet loss rate of the probe message may be determined by a sequence number in a received response, where the sequence number in the response is the same as the sequence number of the probe message corresponding to the response.
[0149] Optionally, the processing unit 902 is further configured to, when there are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include a target path, count the delay of each transmission path in the multiple transmission paths; and adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0150] Optionally, the processing unit 902 is further configured to determine an order of the multiple transmission paths according to the transmission delays of the multiple transmission paths, and determine a bandwidth adjustment amount for each transmission path according to the order.
[0151] Optionally, the processing unit 902 is further configured to determine a transmission path that needs to be optimized according to the order of the multiple transmission paths.
[0152] Optionally, the processing unit 902 is further configured to determine a transmission path for reselection according to the order of the multiple transmission paths.
[0153] In the embodiment of the present application, the operations performed by each unit in the network device are similar to the description of the steps performed by the source device or source end in the embodiments shown in Figures 2 to 8 above, and will not be repeated here.
[0154] In addition, a structure of a computer device provided in an embodiment of the present application can also be understood by referring to Figure 10. The computer device can be a server. As shown in Figure 10, the computer device 100 provided in an embodiment of the present application includes: a processor 1001, a communication interface 1002, a memory 1003, and a bus 1004. The processor 1001, the communication interface 1002, and the memory 1003 are interconnected via the bus 1004. In an embodiment of the present application, the processor 1001 is used to control and manage the actions of the computer device 100. For example, the processor 1001 is used to determine the available bandwidth for transmitting data packets. The communication interface 1002 is used to support the computer device 100 in communicating. For example, the communication interface 1002 can perform the steps of sending a probe message, receiving a response, and sending a data message. The memory 1003 is used to store the program code and data of the computer device 100.
[0155] The processor 1001 may be a central processing unit (CPU), a general-purpose processor (GPOR), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device (PLD), a transistor logic device (TLD), a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. A processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The bus 1004 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG. 10 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0156] In another embodiment of the present application, a computer-readable storage medium is further provided, in which computer-executable instructions are stored. When the processor of a computer device executes the computer-executable instructions, the computer device executes the steps performed by the source device or source end in Figures 2 to 8 above.
[0157] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer program code. When the computer program code is executed on a computer, the computer device executes the steps executed by the source device or source end in Figures 2 to 8 above.
[0158] In another embodiment of the present application, a chip system is also provided, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected through lines; the interface circuits are used to receive signals from the memory of a computer device and send signals to the processor, and the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the computer device executes the steps performed by the source device or source end in the aforementioned Figures 2 to 8.
[0159] In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the source device or source end. The chip system may be composed of a chip or may include a chip and other discrete devices.
[0160] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0161] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0162] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in whole or in part through software, hardware, firmware, or any combination thereof.
[0163] When software is used to implement the integrated unit, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
Claims
1. A congestion control method, characterized in that, The method is applied to a source device, which includes a plurality of queues. The method includes: Sending a plurality of probe messages, which are used to probe the allocated bandwidth on the target path between the source device and the destination device for the plurality of queues; Receiving a plurality of responses from the destination device, which are used to indicate the available bandwidth detected by the plurality of probe messages. The available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data messages in the plurality of queues; Sending the data messages in the plurality of queues to the destination device at the available bandwidth.
2. The method according to claim 1, wherein None of the plurality of probe messages includes the queue identifier of any one of the plurality of queues.
3. The method according to claim 1 or 2, characterized in that, Before sending the data messages in the plurality of queues to the destination device at the available bandwidth, the method further includes: Generating a plurality of tokens according to a first relationship and the plurality of responses, which are used to schedule the data messages in the plurality of queues. The first relationship is the ratio relationship between the received amount of responses and the sent amount of data messages.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Determining the packet loss rate of the probe messages according to the number of sent probe messages and the number of received responses, where the packet loss rate is used to adjust the sending rate of the probe messages.
5. The method according to any one of claims 1 to 4, characterized in that There are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include the target path. The method further includes: Statistical the delay of each transmission path in the multiple transmission paths; Adjusting the allocated bandwidth of the target path according to the delay of each transmission path.
6. A network device, characterized in that, Including: A transceiver unit and a processing unit; wherein, the transceiver unit is used to execute the steps related to sending or receiving in any one of claims 1-5 above, and the processing unit is used to execute the steps other than sending and receiving in any one of claims 1-5 above.
7. A network device, characterized in that, Including: A communication interface, a processor and a memory, where the communication interface and the processor are coupled to the memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the network device is caused to execute the method as described in any one of claims 1 to 5.
8. A computer device, characterized in that, Including: A network device, which is used to execute the method as described in any one of claims 1-5 above.
9. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When the instructions run on a computer device, the computer device is caused to execute the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, The computer program product includes computer program code. When the computer program code runs on a computer device, the computer device is caused to execute the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Network quality detection method, device and server
CN110445666A
Congestion control method and related equipment
CN116260773A
Ethernet congestion control and prevention
US20180034740A1
Network congestion control method and related apparatus
WO2022022539A1
Cited By
Data forwarding method, intelligent computing center, node of intelligent computing center, medium and product
CN120567753A
RDMA traffic DDoS attack detection method and system in edge computing power network environment
CN120979753A