Congestion control method and corresponding device
By detecting the available bandwidth first, and then sending data packets, the problem of congestion control lag and resource waste in RDMA network is solved, and efficient congestion control and resource utilization are achieved.
Patent Information
- Application Number
- CN202410071325.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing RDMA network congestion control method has lag, and the bandwidth allocation method of detection packets leads to waste of hardware resources, especially when the number of queues increases, resource waste is serious.
The method of detecting messages first detecting the available bandwidth and then sending data messages is adopted. Bandwidth detection is carried out for multiple queues through a set of detection resources, reducing resource waste, and optimizing the detection process through token management and packet loss rate adjustment.
Effectively control congestion, reduce the use of detection resources, improve service stability, reduce IOPS requirements, and improve bandwidth resource utilization.
Smart Images

Figure CN120342954A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a method and corresponding device for congestion control. Background Art
[0002] Traditional remote direct memory access (RDMA) congestion control takes the explicit congestion notification (ECN) mark of the switch as the congestion signal. The sender speeds up or slows down according to whether the acknowledge (ACK) packet of the data packet carries the ECN mark. This congestion method has a certain lag.
[0003] To solve the problem of the lag in network congestion control of RDMA, a method of pre-sending probe packets can be used to detect the transmission resources of the network, and then data packets are sent according to the detection results. This scheduling mode of detecting first and then sending reduces the probability of congestion from the source.
[0004] Currently, when using probe packets to detect the bandwidth for transmitting data packets, the bandwidth for detection is allocated according to queue pairs (QPs). In this way, the bandwidth for detection will linearly expand with the number of QPs. Due to limited hardware resources, when the number of QPs increases, this congestion control method will waste a large amount of bandwidth. Summary of the Invention
[0005] This application provides a method for congestion control, which is used to detect the available bandwidth for data packets in multiple queues with fewer detection resources, thereby effectively controlling congestion. This application also provides a corresponding device, a computer-readable storage medium, a computer program product, etc.
[0006] In a first aspect of this application, a method for congestion control is provided. This method is applied to a source device, and the source device includes multiple queues. The method includes: sending multiple probe packets, where the multiple probe packets are used to detect the allocated bandwidth on the target path between the source device and the destination device for the multiple queues; receiving multiple responses from the destination device, where the multiple responses are used to indicate the available bandwidth detected by the multiple probe packets, and the available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data packets in the multiple queues; sending data packets in the multiple queues to the destination device with the available bandwidth.
[0007] In this application, the source device can be a network card, a data processing unit (DPU), a just a bunch of flash drives (JBOF), a storage box, a computing box, or other devices with network communication functions in other forms. The source device can also be a computer device installed with the above-mentioned network card, DPU, etc.
[0008] In this application, a queue is used to store data packets, and the data packets sent from the source device are scheduled from the queue.
[0009] In this application, the probe packet can also be referred to as a "probe" or a "control packet".
[0010] In this application, on a path, the bandwidth used to transmit data packets and the bandwidth used to transmit probe packets are usually set proportionally, and this proportion can be the ratio of the length of the data packet to the length of the probe packet. For example: the ratio of the length of the data packet to the length of the probe packet is 95:5. If the allocated bandwidth of the target path is 100 megabytes (M), then 95M bandwidth can be used to transmit data packets, and 5M bandwidth can be used to transmit probe packets.
[0011] In this application, when the destination device receives a probe packet through the target path, an acknowledge (ACK) will be sent on the target path. In this way, within a period of time, by counting the number of probe packets sent and the number of received acknowledgments, the available bandwidth of the probe packets on the target path can be determined, and then according to the ratio of the bandwidth of the data packets to the bandwidth of the probe packets, the available bandwidth of the data packets on the target path can be determined, and then the data packets can be sent to the destination device according to the available bandwidth, and congestion will not occur, thus realizing active control of congestion.
[0012] Because the congestion situation of the same target path is the same when transmitting probe packets and data packets, and even if the probe packets are discarded by the device during transmission, it will not affect the service. Therefore, by first detecting the available bandwidth on the target path through the probe packets and then sending the data packets, it replaces the direct competition of the data packets for bandwidth. In this way, the service stability can be improved. Moreover, in the above first aspect, using a set of resources of the probe packets for detection can send the data packets in multiple queues, without the need to configure a set of resources for detection for each queue, reducing the resources used for detection, and reducing the probability of congestion on the target path from the source. Moreover, without the need to send a large number of probe packets, the requirement for input / output operations per second (IOPS) is also reduced.
[0013] In a possible implementation manner, the queue identifiers of any of the multiple queues are not included in the multiple probe messages.
[0014] In this possible implementation manner, since a set of probe resources is used to probe the available bandwidth for multiple queues and the destination device does not need to distinguish the queue corresponding to the probe message, the queue identifier does not need to be carried in the probe message, shortening the length of the probe message and thus reducing the bandwidth used by the probe message.
[0015] In a possible implementation manner, before sending data messages in multiple queues to the destination device at the available bandwidth, the method further includes: generating multiple tokens according to a first relationship and multiple responses, where the multiple tokens are used to schedule the data messages in the multiple queues, and the first relationship is the ratio relationship between the received amount of responses and the sent amount of data messages.
[0016] In this possible implementation manner, if the ratio relationship between the received amount of probe messages and the sent amount of data messages is 1:2, the source device generates two tokens for each received response, and the data messages in the queue are scheduled through the tokens, which can effectively control the sending of data messages in the multiple queues.
[0017] In a possible implementation manner, the increment of the token corresponds to the number of received responses, and the consumption of the token corresponds to the number of data messages scheduled from the multiple queues.
[0018] In a possible implementation manner, the method further includes: determining the packet loss rate of the probe message according to the number of sent probe messages and the number of received responses, and the packet loss rate is used to adjust the sending rate of the probe message.
[0019] In this possible implementation manner, by adjusting the sending rate of the probe message through the packet loss rate, the probing quality of the probe message for the target path can be improved.
[0020] In a possible implementation manner, the packet loss rate of the probe message can be determined through the sequence number in the received response, and the sequence number in the response is the same as the sequence number of the probe message corresponding to the response.
[0021] In a possible implementation manner, there are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include the target path; the method further includes: statistically analyzing the delay of each transmission path in the multiple transmission paths; adjusting the allocated bandwidth of the target path according to the delay of each transmission path.
[0022] In this possible implementation, when there are multiple transmission paths between the source device and the destination device, the allocated bandwidth of each transmission path can be adjusted based on the total bandwidth of the multiple transmission paths according to the delay of each transmission path. Such flexible adjustment can improve the utilization rate of bandwidth resources.
[0023] In a possible implementation, the source device can also determine the sorting of the multiple transmission paths according to the transmission delays of the multiple transmission paths, and then determine the bandwidth adjustment amount of each transmission path according to the sorting.
[0024] In a possible implementation, the source device can also determine the transmission paths that need to be optimized according to the sorting of the multiple transmission paths.
[0025] In a possible implementation, the source device can also determine the transmission paths for reselection according to the sorting of the multiple transmission paths.
[0026] In a possible implementation, at least one probe management block (PMB) can be configured in the source device, and each probe management block corresponds to a different destination device. Each probe management block can manage steps such as sending probe messages from the source device to the corresponding destination device, receiving responses, and scheduling data messages.
[0027] A second aspect of this application provides a network device, including: a transceiver unit and a processing unit; wherein,
[0028] The transceiver unit is used for:
[0029] Sending multiple probe messages, where the multiple probe messages are used to probe the allocated bandwidth on the target path between the source device and the destination device for multiple queues;
[0030] Receiving multiple responses from the destination device, where the multiple responses are used to indicate the available bandwidth detected by the multiple probe messages, and the available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data messages in multiple queues;
[0031] Sending data messages in multiple queues to the destination device with the available bandwidth.
[0032] In a possible implementation, none of the multiple probe messages includes the queue identifier of any one of the multiple queues.
[0033] In a possible implementation, the processing unit is used to generate multiple tokens according to the first relationship and the multiple responses, where the multiple tokens are used to schedule data messages in multiple queues, and the first relationship is the ratio relationship between the received amount of responses and the sent amount of data messages.
[0034] In a possible implementation, the increment of the token corresponds to the number of received responses, and the consumption of the token corresponds to the number of data packets scheduled from multiple queues.
[0035] In a possible implementation, the processing unit is further configured to determine the packet loss rate of the probe packets according to the number of sent probe packets and the number of received responses, and the packet loss rate is used to adjust the sending rate of the probe packets.
[0036] In a possible implementation, the packet loss rate of the probe packets can be determined by the sequence numbers in the received responses, and the sequence numbers in the responses are the same as the sequence numbers of the probe packets corresponding to the responses.
[0037] In a possible implementation, when there are multiple transmission paths between the source device and the destination device and the multiple transmission paths include a target path, the processing unit is further configured to statistically calculate the delay of each transmission path in the multiple transmission paths; adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0038] In a possible implementation, the processing unit is further configured to determine the sorting of the multiple transmission paths according to the transmission delays of the multiple transmission paths, and determine the bandwidth adjustment amount of each transmission path according to the sorting.
[0039] In a possible implementation, the processing unit is further configured to determine the transmission paths that need to be optimized according to the sorting of the multiple transmission paths.
[0040] In a possible implementation, the processing unit is further configured to determine the transmission paths for reselection according to the sorting of the multiple transmission paths.
[0041] The third aspect of the present application provides a network device, a communication interface, a processor, and a memory. The communication interface and the processor are coupled to the memory. The memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the feature screening device executes the methods in the foregoing first aspect or any possible implementation manner of the first aspect.
[0042] The fourth aspect of the present application provides a chip system, which includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are used to receive signals from the memory of the cloud device and send the signals to the processors, and the signals include computer instructions stored in the memory; when the processors execute the computer instructions, the feature screening device executes the methods in the foregoing first aspect or any possible implementation manner of the first aspect.
[0043] The fifth aspect of the present application provides a computer device, including the network device in the foregoing second aspect, third aspect, or any possible implementation manner of the second aspect.
[0044] The sixth aspect of the present application provides a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction runs on a computer device, the computer device is caused to execute the method in the foregoing first aspect or any possible implementation manner of the first aspect.
[0045] The seventh aspect of the present application provides a computer device program product, which includes computer device program code. When the computer device program code is executed on a computer device, the computer device is caused to execute the method in the foregoing first aspect or any possible implementation manner of the first aspect.
[0046] The eighth aspect of the present application provides a congestion control system, which includes a source device and at least one destination device. The source device is used to execute the method in the foregoing first aspect or any possible implementation manner of the first aspect. The destination device is used to send a response to the source device after receiving a probe message.
[0047] Among them, for the technical effects brought by the second aspect to the eighth aspect or any possible implementation manner thereof, reference may be made to the technical effects brought by the first aspect or different possible implementation manners of the first aspect, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1A It is a schematic diagram of an architecture of a data center provided by an embodiment of the present application;
[0049] Figure 1B It is another schematic diagram of an architecture of a data center provided by an embodiment of the present application;
[0050] Figure 2 It is a schematic diagram of an embodiment of a congestion control method provided by an embodiment of the present application;
[0051] Figure 3 It is another schematic diagram of an embodiment of a congestion control method provided by an embodiment of the present application;
[0052] Figure 4A It is a schematic diagram of a relationship between a probe management block PMB and a destination device provided by an embodiment of the present application;
[0053] Figure 4B It is another schematic diagram of a relationship between a PMB and a destination device provided by an embodiment of the present application;
[0054] Figure 5 It is a schematic diagram of an example of token management provided by an embodiment of the present application;
[0055] Figure 6 It is a schematic diagram of an example of multiple transmission paths provided by an embodiment of the present application;
[0056] Figure 7 An exemplary schematic diagram of the cloud computing scenario provided by the embodiment of the present application;
[0057] Figure 8 An exemplary schematic diagram of the artificial intelligence scenario provided by the embodiment of the present application;
[0058] Figure 9 A schematic structural diagram of the network device provided by the embodiment of the present application;
[0059] Figure 10 It is a schematic structural diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0060] Next, in conjunction with the accompanying drawings, the embodiments of the present application will be described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0061] Terms such as "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0062] The embodiment of the present application provides a congestion control method for detecting the available bandwidth of data packets in multiple queues with fewer detection resources, so as to effectively control congestion. The present application also provides corresponding devices, computer-readable storage media, computer program products, etc. The following will be described in detail respectively.
[0063] For ease of understanding, the following briefly introduces the technical terms related to the embodiments of the present application:
[0064] 1. Remote Direct Memory Access (RDMA): A new type of network communication technology, which is developed to address the latency in server - side data processing during network transmission. RDMA directly transfers data into the storage area of a computer through the network, quickly moving data from one system to the remote system's memory without any impact on the operating system. This way, it doesn't require much processing power of the computer, eliminating the overhead of external memory copying and context switching, and featuring low latency and high bandwidth.
[0065] 2. Proactive Congestion Control Algorithm (PCC): A type of signaling - based network congestion control algorithm commonly used in RDMA protocols. It can be used for network congestion control in RDMA protocols. By first sending probe packets to detect the available bandwidth and then sending data packets based on the available bandwidth, it reduces the probability of congestion at the source.
[0066] 3. Delay: The time required for a message or packet to be transmitted from the source end to the destination end of a network.
[0067] 4. Loss Tolerance or Packet Loss Rate: The ratio of the number of lost data packets to the total number of transmitted data packets during a test.
[0068] 5. Queue Pair (QP): The basic communication unit in the reliable connection mode of the RDMA protocol.
[0069] 6. Probe Packet: Also known as a probe or control packet, used to detect network load or round - trip time (RTT).
[0070] The congestion control method provided in the embodiments of this application can be applied to a data center. The architecture of the data center can be understood by referring to Figure 1A As shown in Figure 1A A data center may include multiple Points of Delivery (PODs). A POD is the smallest business unit in a data center and is a resource collection of switches, firewalls, and servers. Each POD can include a computer device cluster (such as a server cluster), and different computer devices can communicate through switches. To meet the connection requirements among a large number of computer devices in the data center, multi - layer switches are usually set up.
[0071] As Figure 1AIn the data center, computer devices can first be connected through access switches, and then the access switches are connected to different aggregation switches, and the aggregation switches are connected to different core switches, so that communication between different computer devices in the data center can be achieved. Among them, the access switches and aggregation switches in different PODs may not be shared, and communication between computer devices in different PODs is achieved through core switches.
[0072] The architecture of the data center can also be referred to Figure 1B for understanding. As Figure 1B shown, the data center may include multiple PODs, and the multiple PODs may include storage PODs and computing PODs. Among them, the computer devices in the storage POD are mainly storage servers for storing data, and the computer devices in the computing POD are mainly computing servers for completing computing tasks.
[0073] During the process of the computing server executing the computing task, it may need to read data from the storage server. After the computing is completed, it may also need to store the computing result in the storage server. In addition, when multiple computing servers jointly execute a computing task, data may also need to be transferred between different computing servers. Therefore, communication is also required between different computing servers. Data transfer may also occur between different storage servers. Therefore, communication is also required between different storage servers. Therefore, in the data center, whether it is a computing server or a storage server, when each server is the source end, there can be multiple destination ends. The source end and the destination end can be in the same POD or in different PODs.
[0074] For the source end and the destination end in the same POD, if they are connected through the same access switch, they can communicate directly through the access switch. If the source end and the destination end in the same POD are not connected to the same access switch, communication can be achieved through the aggregation switch at a higher level that connects their respective access switches.
[0075] For the source end and the destination end in different PODs, communication can be achieved through the access switches, aggregation switches in their respective PODs, and the core switch located above the PODs.
[0076] According to the integration mechanism with computer devices, access switches can be divided into three types, namely end of row (EOR) switches, middle of row (MOR) switches, and top of rack (TOR) switches. These types of switches can also be collectively referred to as cabinet switches.
[0077] Among them, the integrated architecture of the EOR switch is to centrally install the access switches in the cabinet (switch cabinet) at the end of a row of cabinets. The hosts, servers, and minicomputers in the equipment cabinet are connected through the horizontal cables via the permanent links.
[0078] The integrated architecture of the MOR switch allows the cables to run from the cabinets in the middle position to both ends, reducing the congestion of the cables at the cable routing channel entrances and exits and reducing the average length of the cables.
[0079] The integrated architecture of the TOR switch is to deploy 1-2 access switches at the upper end of each cabinet in the POD. The servers are connected to the cabinets through jumpers. On the access switch, the uplink ports of the access switch are connected to the aggregation switch or the core switch through copper cables or optical fibers.
[0080] Figure 1B The access switch shown is a TOR switch. This TOR switch first connects to the aggregation switch, and the aggregation switch then connects to the core switch. If the data center is small, the TOR switch can also be directly connected to the core switch without the need to configure an aggregation switch.
[0081] In addition, it should be noted that the above Figure 1A and Figure 1B are all exemplified by a single data center. In fact, the devices between different data centers can also communicate with each other. That is to say, the source end and the destination end can be distributed in the same region, or in different regions. Further, the source end and the destination end can be distributed in the same availability zone (AZ), or in different AZs. Each AZ includes one data center or multiple data centers with close geographical locations. Usually, one region can include multiple AZs.
[0082] The source end and the destination end can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, communication gateways need to be set in each VPC, and the interconnection between VPCs is realized through the communication gateways.
[0083] In the embodiments of this application, the source end and the destination end can communicate in the RDMA manner. However, because the bandwidth resources between the source end and the destination end are limited by the hardware capabilities, network congestion may occur when transmitting data packets between the source end and the destination end.
[0084] To control network congestion between the source end and the destination end, an active congestion control method can be adopted. Probe messages are sent on the transmission path between the source end and the destination end to first detect the available bandwidth for sending data messages, and then data messages are sent according to the available bandwidth. In this way, the probability of network congestion can be reduced from the source.
[0085] Next, taking Figure 2 as an example, the process of this active congestion control will be introduced.
[0086] As Figure 2 shown, there can be one or more switches on the transmission path from the source end to the destination end. Since the loads connected to different switches may be different, the amount of data on some switches may be larger. Or, the hardware capabilities of some switches are limited, which will cause transmission bottlenecks between the source end and the destination end. And the available bandwidth on a transmission path is usually determined by the transmission bottleneck. Therefore, the source end can first send probe messages for detection, and then send data messages after determining the available bandwidth.
[0087] On a transmission path, the bandwidth for transmitting data messages and the bandwidth for transmitting probe messages are usually set proportionally, and this proportion can be the ratio of the length of the data message to the length of the probe message. For example: the ratio of the length of the data message to the length of the probe message is 95:5. If the allocated bandwidth of this target path is 100 megabytes (M), then 95M bandwidth can be used to transmit data messages, and 5M bandwidth can be used to transmit probe messages.
[0088] The source end can first send probe messages on the 5M bandwidth in the above example to detect the available bandwidth for transmitting data messages on this transmission path through this probe message. This detection process can be that the source end can send probe messages at a relatively high rate. When the probe messages are transmitted in the network, they may be rate-limited, and some probe messages may also be discarded. But every time the destination end receives a probe message, it will send an acknowledgment (ACK) to the source end.
[0089] The source end can count the data of the sent probe messages and the number of received acknowledgments, then calculate the available bandwidth in the 5M bandwidth allocated to the probe messages, and then calculate the available bandwidth in the 95M allocated bandwidth according to the ratio. For example: if the calculated available bandwidth of the probe message is 2.5M, then the ratio of the available bandwidth (2.5M) of the probe message to the allocated bandwidth (5M) is 1:2. Then, through 95*(1 / 2) = 47.5, it can be obtained that the available bandwidth for transmitting data messages is 47.5M. Then data messages can be sent according to this 47.5M bandwidth, thereby reducing the probability of network congestion between the source end and the destination end from the source.
[0090] The above-mentioned Figure 2 Both the source end and the destination end in the above can be computer devices (servers), virtual machines or containers. Of course, the source end and the destination end can also be devices implemented by using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. Among them, the above-mentioned PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0091] For the congestion control method provided by the embodiments of the present application, reference can also be made to Figure 3 for understanding.
[0092] As Figure 3 shown, an embodiment of the congestion control method provided by the embodiments of the present application includes:
[0093] 301. The source device sends multiple probe packets, and the multiple probe packets are used to probe the allocated bandwidth on the target path between the source device and the destination device for multiple queues.
[0094] In the embodiments of the present application, the source device includes multiple queues, and the queues are used to store data packets. The data packets sent from the source device are scheduled from the queues.
[0095] In the present application, the source device can be a network card, a data processing unit (DPU), a just a bunch of flash drives (JBOF), a storage frame, a computing frame or other devices with network communication functions in other forms. The source device can also be a computer device installed with the above-mentioned network card, DPU, etc.
[0096] 302. The source device receives multiple responses from the destination device, and the multiple responses are used to indicate the available bandwidth detected by the multiple probe packets. The available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit the data packets in the multiple queues.
[0097] For the determination of the available bandwidth, reference can be made to the introduction in the above Figure 2 part for understanding.
[0098] 303. The source device sends the data packets in the multiple queues to the destination device at the available bandwidth.
[0099] In this application, when the destination device receives a probe message through a target path, an ACK will be sent on that target path. In this way, within a certain period of time, by counting the number of probe messages sent and the number of responses received, the available bandwidth of the probe messages on the target path can be determined. Then, according to the ratio of the bandwidth of the data message to the bandwidth of the probe message, the available bandwidth of the data message on the target path can be determined. Then, data messages are sent to the destination device according to this available bandwidth, and congestion will not occur, thus realizing active control of congestion.
[0100] Because the congestion situation of the same target path is the same when transmitting probe messages and data messages, and even if the probe message is discarded by the device during transmission, it will not affect the service. Therefore, by first detecting the available bandwidth on the target path through the probe message and then sending the data message, it replaces the direct competition of the data message for bandwidth. In this way, the service stability can be improved. Moreover, the solution provided in the embodiment of this application uses a set of resources of the probe message for detection, and data messages in multiple queues can be sent. There is no need to configure a set of resources for detection for each queue, reducing the resources used for detection, and reducing the probability of congestion on the target path from the source. Moreover, there is no need to send a large number of probe messages, which also reduces the requirement for (input / output operations per second, IOPS).
[0101] Optionally, in the embodiment of this application, at least one probe management block (PMB) can be configured in the source device, and each probe management block corresponds to a different destination device. Each probe management block can manage the sending of probe messages, the reception of responses, and the scheduling of data messages from the source device to the corresponding destination device.
[0102] Regarding the relationship between the PMB and the queue, reference can be made to Figure 4A for understanding. As Figure 4A shown, the source device includes multiple queues. Taking three queues as an example, the three queues are represented by QP1, QP2, and QP3 respectively. These three queues are used to store data messages destined for the same destination device 1. The source device includes PMB 1, and this PMB 1 can be created with the establishment of any one of QP1, QP2, and QP3. PMB 1 can be the unified outlet for probe messages destined for destination device 1 and the unified inlet for responses from destination device 1.
[0103] Since PMB 1 uniformly manages the probe messages and responses for destination device 1, among the multiple probe messages sent through PMB 1, the queue identifier of any one of the multiple queues is not included. That is: the multiple probe messages sent through PMB 1 do not need to carry the identifiers of QP1, QP2, or QP3. The destination device also does not need to distinguish the queue corresponding to the probe message from the source device. For the probe message from the source device, it only needs to reply with a response, and it does not need to distinguish which queue the response is for. In this way, not only the length of the probe message sent by the source device is shortened, but also the length of the response replied by the destination device is shortened, thus reducing the bandwidth resources occupied by the probe.
[0104] The above-mentioned Figure 4A In the source device shown above, only one PMB 1 is shown. Because there are many nodes in the data center and each source device communicates with multiple destination devices, there will be multiple PMBs in the source device. As Figure 4B shown, if the source device communicates with N destination devices, then it is through PMB 1 for destination device 1, through PMB 2 for destination device 2, …, through PMB N for destination device N. Each PMB corresponds to a different destination device. For the data messages of the same destination device, regardless of how many queues there are, the corresponding probe message sending, response receiving, and data message scheduling are all managed through the corresponding PMB. For example: the probe message sending, response receiving, and data message scheduling of the two queues QP4 and QP5 are all managed through PMB 2, …, the probe message sending, response receiving, and data message scheduling of the two queues QPx and QPy are all managed through PMB N.
[0105] In addition, it should be noted that the above only takes the source device as an example to illustrate the relationship between PMB and the queue. In fact, each computer device in the data center or the network card in the computer device can be used as the source device to communicate with other computer devices or network cards.
[0106] In the embodiment of the present application, the scheduling of the data messages in the queue by PMB can be achieved by distributing tokens (tokens) to the queue. Generally, one token can schedule one data message.
[0107] For the management of tokens, reference can be made to Figure 5 for understanding. As Figure 5As shown, PMB 1 manages three queues, namely QP1, QP2, and QP3. After the source device sends a probe message, it will receive a response. After receiving the response, PMB 1 can generate tokens based on the first relationship and the response. The first relationship is the ratio relationship between the received amount of responses and the sent amount of data messages. This ratio relationship can be pre-configured or dynamically determined. Taking the ratio relationship of 1:2 as an example, for every response received by the source device, PMB 1 will generate two tokens.
[0108] Then PMB 1 distributes tokens to the queues to send data messages. PMB 1 can distribute tokens to the queues one by one. For example, after the data messages in the QP1 queue are scheduled, tokens are distributed to the QP2 queue. After the data messages in the QP2 queue are scheduled, tokens are distributed to the QP3 queue.
[0109] Of course, PMB 1 can also distribute tokens in a round-robin manner. For example, one token is distributed to the QP1 queue, the QP2 queue, and the QP3 queue respectively, and then the next round of distribution starts. This application does not make any restrictions on this.
[0110] Tokens can be maintained in a credit pool. The increment of tokens in this credit pool corresponds to the number of received responses, and the consumption of tokens corresponds to the number of data messages scheduled from multiple queues. For example, if the ratio relationship is 1:2, then for every response received, two tokens will be added to the credit pool. If one token is used to schedule one data message, then for sending one data message, one token will be consumed from the credit pool.
[0111] Optionally, in the embodiments of this application, the packet loss rate of the probe messages can also be determined according to the number of sent probe messages and the number of received responses. The packet loss rate is used to adjust the sending rate of the probe messages. By adjusting the sending rate of the probe messages through the packet loss rate, the detection quality of the probe messages for the target path can be improved.
[0112] Regarding the statistics of the packet loss rate, it can be carried out periodically. For example, taking 30 microseconds (us) as a period, the source device counts the number of sent probe messages and the number of received responses within 30 us. If the number of sent messages is 100 and the number of received responses is 95, then 5 packets are lost and the packet loss rate is 5%.
[0113] Of course, the packet loss rate can also be statistically analyzed by other methods. For example, the packet loss rate can be determined through the sequence numbers in the received responses. The sequence number in the response is the same as the sequence number of the probe message corresponding to this response.
[0114] For example: The source device sends 100 probe messages with sequence numbers from 0 to 99. Among the received responses with sequence numbers less than 100, the sequence numbers do not include 3, 11, 62, 78, and 95. Then it can be determined that during the transmission of these 100 probe messages from the source device to the destination device, the probe messages with sequence numbers 3, 11, 62, 78, and 95 are lost, and the packet loss rate can be determined to be 5%.
[0115] The above-described process is introduced by taking the example of having one transmission path between the source device and the destination device. In fact, there may be multiple transmission paths between the source device and the destination device. When there are multiple transmission paths, the management process of the PMB for multiple transmission paths can refer to Figure 6 for understanding.
[0116] As Figure 6 shown, there are four transmission paths between the source device and the destination device, namely Path 1, Path 2, Path 3, and Path 4. The source device can refer to the previous detection method for the target path, send probe messages on each transmission path, then receive responses from each transmission path, and then determine the available bandwidth of each transmission path according to the received responses. If only one transmission path is needed to transmit data messages, the source device can select the one with the largest available bandwidth from the four transmission paths to send data messages.
[0117] In addition, the source device can also count the delay of each transmission path among multiple transmission paths; adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0118] After the source device counts the delays of the above four transmission paths, it can sort these four transmission paths in ascending order of delay. For example: if the sorting of Path 1, Path 2, Path 3, and Path 4 in ascending order of delay is also Path 1, Path 2, Path 3, and Path 4, then it can be determined that the delay of Path 1 is the smallest and the data is transmitted fastest. Then, the bandwidth of at least one of the other three paths can be reduced, and the extra bandwidth can be allocated to Path 1. For example: reduce the bandwidth of Path 4 by 10M and allocate it to Path 1, so that Path 1 will have a bandwidth of 110M. In this way, transmitting data messages on Path 1 is beneficial to improving the transmission quality of data messages.
[0119] The delay of the transmission path can be determined by the sequence number of the probe message, the sending timestamp, as well as the sequence number of the received response and the receiving timestamp. For example: for a probe message and a response with the same sequence number, the receiving time of the response minus the sending time of the probe message can be used to determine the delay. Of course, the delay of the transmission path can be determined by the average value of the delays calculated from multiple probe messages and the corresponding responses on this transmission path.
[0120] In the embodiments of the present application, path management can also be performed based on path sorting, such as path reselection and path aging processing.
[0121] For example, if the above paths 1, 2, 3, and 4 are sorted in ascending order of transmission delay, the order remains paths 1, 2, 3, and 4. Then weighted credit allocation can be performed for these four transmission paths. If the initial weight of each transmission path is 1, according to the sorting result, the weight of path 4 with the largest transmission delay can be reduced by one grade, for example, reduced by 0.05, so the adjusted weight of path 4 is 0.95. The weight of path 1 with the smallest transmission delay can be increased by one grade, for example, increased by 0.05, so the adjusted weight of path 1 is 1.05. Paths 2 and 3 can remain unchanged. Of course, the adjustment amount for each credit grade can be pre-configured, and how the weights of different paths are adjusted can also be configured according to requirements.
[0122] When there are multiple transmission paths between the source device and the destination device, if path reselection occurs, the reselected path can be determined according to the transmission delay or weighted credit of each transmission path. For example, path 1 can be selected during reselection. When aging paths need to be processed, the path with the largest transmission delay or the smallest weighted credit can be selected for aging processing. For example, Figure 6 path 4 in
[0123] The above congestion control solutions provided by the embodiments of the present application can be applied to various scenarios, such as cloud computing scenarios or artificial intelligence (AI) scenarios, which will be introduced separately below.
[0124] In a cloud computing scenario, a network center with 10K - order nodes may be built according to its scale, and servers are deployed by POD and by AZ. As Figure 7 shown, in this cloud system, it may include terminal device 01, scheduling node 02, working node cluster 03, and working node cluster 04. Communication connections can be established between terminal device 01 and scheduling node 02, between scheduling node 02 and working node cluster 03, between scheduling node 02 and working node cluster 04, and between working node cluster 03 and working node cluster 04. For example, a communication connection can be established between terminal device 01 and scheduling node 02 through a network. Optionally, the network can be a local area network, the Internet, or other networks. In a cloud computing scenario, the network can also be a virtual private cloud (VPC), which is not limited in the embodiments of the present application.
[0125] Among them, working node cluster 03 and working node cluster 04 can be deployed in different PODs or different AZs.Figure 7 Although not drawn in the figure, it can be understood that communication can be carried out between the terminal device 01 and the scheduling node 02, between the scheduling node 02 and the working node cluster 03, between the scheduling node 02 and the working node cluster 04, and between the working node cluster 03 and the working node cluster 04 through a switch. For the connection form between the switch and the server, reference can be made to Figure 1B for understanding.
[0126] In this implementation environment, the operation and maintenance personnel of the cloud system can interact with the scheduling node 02 through the terminal device 01. For example, the operation and maintenance personnel can send an instruction to deploy the cloud system to the scheduling node 02 through the terminal device 01, so as to instruct the scheduling node 02 to deploy the cloud system based on the resources of the scheduling node 02, the working node cluster 03, and the working node cluster 04. The cloud system is used to manage the resources of the working node cluster 03 and the working node cluster 04, and provide cloud services for tenants based on the resources of the working node cluster 03 and the working node cluster 04.
[0127] Optionally, the terminal device 01, also known as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., is a device including wireless communication functions (providing voice / data connectivity to users). For example, it is a handheld device with wireless connection functions. Currently, some examples of terminal devices are: mobile phone, tablet computer, laptop computer, palmtop computer, laptop computer, wireless router, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in vehicle networking, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, or wireless terminal in smart home, etc.
[0128] The scheduling node 02 can be a cloud server or a cloud physical machine, or a cloud server cluster or a physical machine cluster composed of several cloud servers, or a cloud computing service center. The functions of the scheduling node can be implemented by software or hardware.
[0129] The working node cluster 03 and the working node cluster 04 can be a server cluster composed of several servers or a cloud computing service center. Among them, a large number of basic resources owned by the cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources, network resources, etc. are deployed in the cloud computing service center.
[0130] When the cloud system shown above Figure 7 runs the RDMA service, the link bandwidth is generally 100Gbps / 200Gbps / 400Gbps, and it may develop to 800Gbps in the future. The single queue pair (QP) capacity of the product network card chip is about 50Gbps - 80Gbps. At the same time, there will be a bandwidth loss of 10% - 20% in the business application. Therefore, if a 200Gbps link is used for networking, at least 3 - 5 QPs need to be established between two working nodes to ensure the bandwidth capacity. If 4 QPs are needed. At this time, for a full - interconnected network of 10,000 working nodes, each working node needs to maintain 4 (QPs) * 10K = 40K QPs (not all active). For the convenience of description, taking a fixed number of active QPs (taking 100 QPs of 25 working nodes as an example), the actual number of active QPs will be more.
[0131] If the bandwidth requirement of the active QP is higher than the 200Gbps link bandwidth, taking the Maximum Transmission Unit (MTU) as 1KB as an example, the PMB adopts a management mode of data packet: probe packet = 2:1, and the probe packets are uniformly managed according to the destination. Then, for each working node, 24 probe packet streams to other 24 working nodes need to be maintained. Each probe packet stream sends probe packets at an initial rate of 200Gbps * 2.5%. After the probe packets are rate - limited and fairly competed on the network switch side, a part of the probe packets reach the destination. After the destination receives the probe packet, it will return an ACK for the probe packet. After the source end receives the ACK returned by the destination end, its PMB will generate tokens according to the ratio of 1:2 based on the number of received ACKs and regularly distribute the available tokens to the active QP of this destination. The PMB also records the timestamp of ACK reception and calculates the packet loss rate of the probe packet according to the sequence number of the ACK. Then, the packet loss rate is used to adjust the sending rate, and the probe packets are continued to be sent at the adjusted rate.
[0132] When the data packet transmission of a QP is completed, at this time, there is no new data to be transmitted for this QP, but the probe packets sent to the same destination will not stop being sent until the business of the last QP is completed. Therefore, some tokens need to be re - allocated to other QPs with the same destination.
[0133] The solution provided by the embodiments of the present application can also be applied to the AI scenario. In order to complete the large model training faster, the AI scenario requires higher bandwidth and lower latency. At this time, the multi-path mode is a more efficient transmission mode.
[0134] As Figure 8 shown, in the AI scenario, an AI model is jointly trained by server A and server B. Server A conducts the first stage of training on the AI model, and the data generated during training needs to be transmitted to server B. Server B uses the data transmitted by server A to conduct the second stage of training on the AI model. In order to improve the transmission rate, four transmission paths can be configured between server A and server B. These four transmission paths can be understood with reference to Figure 6 . Server A or server B can manage these four transmission paths. The specific management process can be understood with reference to the previous introduction.
[0135] The above introduced the congestion control method provided by the embodiments of the present application. Next, the network device in the embodiments of the present application will be introduced with reference to the accompanying drawings.
[0136] As Figure 9 shown, the network device 90 provided by the embodiments of the present application includes: a transceiver unit 901 and a processing unit 902; wherein,
[0137] The transceiver unit 901 is used for:
[0138] Sending multiple probe packets, where the multiple probe packets are used to probe the allocated bandwidth on the target path between the source device and the destination device for multiple queues;
[0139] Receiving multiple responses from the destination device, where the multiple responses are used to indicate the available bandwidth detected by the multiple probe packets, and the available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data packets in multiple queues;
[0140] Sending data packets in multiple queues to the destination device with the available bandwidth.
[0141] The network device provided by the embodiments of the present application can use a set of resources of probe packets for detection, and then send data packets in multiple queues, without the need to configure a set of resources for detection for each queue, reducing the resources used for detection, and reducing the probability of congestion on the target path from the source. Moreover, there is no need to send a large number of probe packets, which also reduces the requirements for IOPS.
[0142] Optionally, none of the multiple probe packets include the queue identifier of any one of the multiple queues.
[0143] Optionally, a processing unit 902 is configured to generate a plurality of tokens according to a first relationship and a plurality of responses, where the plurality of tokens are used to schedule data packets in a plurality of queues, and the first relationship is a ratio relationship between the received amount of responses and the sent amount of data packets.
[0144] Optionally, the increment of the token corresponds to the number of received responses, and the consumption of the token corresponds to the number of data packets scheduled from the plurality of queues. Generally, one token is used to schedule one data packet.
[0145] Optionally, the processing unit 902 is further configured to determine a packet loss rate of the probe packets according to the number of sent probe packets and the number of received responses, and the packet loss rate is used to adjust the sending rate of the probe packets.
[0146] Optionally, the packet loss rate of the probe packets can be determined by the sequence number in the received response, and the sequence number in the response is the same as the sequence number of the probe packet corresponding to the response.
[0147] Optionally, when there are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include a target path, the processing unit 902 is further configured to statistically calculate the delay of each transmission path in the multiple transmission paths; adjust the allocated bandwidth of the target path according to the delay of each transmission path.
[0148] Optionally, the processing unit 902 is further configured to determine the sorting of the multiple transmission paths according to the transmission delays of the multiple transmission paths, and determine the bandwidth adjustment amount of each transmission path according to the sorting.
[0149] Optionally, the processing unit 902 is further configured to determine the transmission path that needs to be optimized according to the sorting of the multiple transmission paths.
[0150] Optionally, the processing unit 902 is further configured to determine the transmission path for reselection according to the sorting of the multiple transmission paths.
[0151] In the embodiments of the present application, the operations performed by each unit in the network device are similar to the descriptions of the steps performed by the source device or the source end in the foregoing Figures 2 to 8 illustrated embodiments, and will not be described in detail here.
[0152] In addition, for the structure of a computer device provided in the embodiments of the present application, reference can also be made to Figure 10 for understanding. The computer device may be a server, such as Figure 10As shown in the figure, the computer device 100 provided by the embodiment of the present application includes: a processor 1001, a communication interface 1002, a memory 1003, and a bus 1004. The processor 1001, the communication interface 1002, and the memory 1003 are interconnected through the bus 1004. In the embodiment of the present application, the processor 1001 is used to control and manage the actions of the computer device 100. For example, the processor 1001 is used to determine the available bandwidth for transmitting data packets. The communication interface 1002 is used to support the computer device 100 to communicate. For example, the communication interface 1002 can perform the steps of sending a probe packet, receiving a response, and sending a data packet. The memory 1003 is used to store the program code and data of the computer device 100.
[0153] Among them, the processor 1001 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 1004 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0154] In another embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores computer-executable instructions. When the processor of the computer device executes the computer-executable instructions, the computer device executes the above Figures 2 to 8 steps performed by the source device or the source end.
[0155] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer program code. When the computer program code is executed on a computer, the computer device executes the above Figures 2 to 8 steps performed by the source device or the source end.
[0156] In another embodiment of the present application, a chip system is further provided. The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are configured to receive signals from the memory of a computer device and send the signals to the processors, and the signals include computer instructions stored in the memory; when the processors execute the computer instructions, the computer device executes the steps performed by the source device or the source end in the foregoing Figures 2 to 8 by the source device or the source end.
[0157] In a possible design, the chip system may further include a memory, which is used to store the necessary program instructions and data of the source device or the source end. The chip system may be composed of chips or may include chips and other discrete devices.
[0158] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be in electrical, mechanical, or other forms.
[0159] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated units may be implemented in whole or in part by software, hardware, firmware, or any combination thereof.
[0161] When the integrated unit is implemented using software, it can be implemented in the form of a computer program product, wholly or partially. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A congestion control method, characterized in that, The method is applied to a source device, which includes a plurality of queues. The method includes: Sending a plurality of probe packets, which are used to probe the allocated bandwidth on the target path between the source device and the destination device for the plurality of queues; Receiving a plurality of responses from the destination device, which are used to indicate the available bandwidth detected by the plurality of probe packets. The available bandwidth is the bandwidth in the allocated bandwidth that can be used to transmit data packets in the plurality of queues; Sending the data packets in the plurality of queues to the destination device at the available bandwidth.
2. The method according to claim 1, wherein The plurality of probe packets do not include the queue identifier of any one of the plurality of queues.
3. The method according to claim 1 or 2, characterized in that, Before sending the data packets in the plurality of queues to the destination device at the available bandwidth, the method further includes: Generating a plurality of tokens according to a first relationship and the plurality of responses, which are used to schedule the data packets in the plurality of queues. The first relationship is the ratio relationship between the received amount of responses and the sent amount of data packets.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Determining the packet loss rate of the probe packets according to the number of sent probe packets and the number of received responses, where the packet loss rate is used to adjust the sending rate of the probe packets.
5. The method according to any one of claims 1-4, characterized in that There are multiple transmission paths between the source device and the destination device, and the multiple transmission paths include the target path. The method further includes: Statistical the delay of each transmission path in the multiple transmission paths; Adjusting the allocated bandwidth of the target path according to the delay of each transmission path.
6. A network device, characterized in that, Including: A transceiver unit and a processing unit; wherein, the transceiver unit is used to execute the steps related to sending or receiving in any one of the above claims 1-5, and the processing unit is used to execute the steps other than sending and receiving in any one of the above claims 1-5.
7. A network device, characterized in that, Including: A communication interface, a processor and a memory. The communication interface and the processor are coupled to the memory. The memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the network device is enabled to execute the method as described in any one of claims 1 to 5.
8. A computer device, characterized in that, Including: A network device, which is used to execute the method as described in any one of the above claims 1-5.
9. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer device, the computer device is enabled to execute the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, The computer program product includes computer program code. When the computer program code is run on a computer device, the computer device is enabled to execute the method as described in any one of claims 1 to 5.