Congestion control method and device for data center network, equipment and medium

The congestion control method implemented on the network card through the LLM-CC algorithm, adjusting the transmission rate based on the explicit congestion notification packet, solving the problem of insufficient network performance in the cross-domain multi-intelligent computing cluster network, and improving network performance and large model training efficiency.

CN120378361AActive Publication Date: 2025-07-25BEIJING JILIU TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510712008.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-25
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The existing network congestion control algorithms are insufficient in cross-domain multi-intelligent computing cluster networks, and fail to effectively optimize the long-term delay and insufficient bandwidth of links, resulting in the inability to fully utilize network performance, affecting the efficiency of tasks such as large-scale model training.

Method used

Using the LLM-CC algorithm, through the congestion control method implemented on the network card, the traffic type and congestion position of the data flow are determined according to the explicit congestion notification packet, the transmission rate is adjusted according to the congestion degree and transmission progress, and the rate recovery is performed when a speed-up event is detected, and the congestion position is subdivided for targeted control.

Benefits of technology

Effectively respond to the long-term link delay and bandwidth shortage caused by cross-domain connections, adapt to the traffic mode of smart computing cluster networks, and improve network performance and efficiency of large-model training tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378361A_ABST
    Figure CN120378361A_ABST
Patent Text Reader

Abstract

The invention provides a congestion control method and device for a data center network, equipment and a medium. The method is applied to a first node in a first intelligent computing cluster in a multi-intelligent computing cluster network, the first node is in communication connection with a second node, the second node is located in the first intelligent computing cluster or a second intelligent computing cluster different from the first intelligent computing cluster, and the first node is used for transmitting target data to the second node. When an explicit congestion notification packet returned by the second node is detected, determining the flow type of the target data flow which is congested currently according to the explicit congestion notification packet, and when the flow type is cross-intelligent cluster flow, determining the congestion position of the target data flow; determining the congestion degree and the transmission progress of the target data flow according to the information of the historical explicit congestion notification packet; and reducing the transmission rate of the target data stream according to the traffic type, the congestion position, the congestion degree and the transmission progress. According to the invention, problems in related technologies are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a congestion control method, apparatus, device, and medium for a data center network. Background Art

[0002] Intelligent computing refers to computing specifically used for artificial intelligence tasks, such as deep learning model training, large-scale data analysis, natural language processing, computer vision, etc. A cluster refers to connecting multiple independent computers (called nodes) through a high-speed network and working together as a whole to provide more powerful computing capabilities, storage capabilities, or higher availability than a single computer. An intelligent computing cluster refers to a computer cluster specifically built to meet the needs of large-scale artificial intelligence computing. It usually has the following characteristics: (1) Powerful parallel computing capabilities: extensive use of graphics processors, tensor processors, or other AI acceleration chips; (2) High-speed interconnected network: nodes require a network connection with extremely high bandwidth and extremely low latency to efficiently transmit massive amounts of training data and model parameters. (3) High-performance storage: a storage system that can quickly read and write large files to match the data consumption speed of computing nodes. (4) Optimized software stack: running specialized AI frameworks, libraries, and cluster management software. A cross-domain multi-intelligent computing cluster is an overall system composed of multiple intelligent computing clusters distributed in different geographical locations or network domains, and communication and collaborative work are required between these clusters.

[0003] Data Center Quantized Congestion Notification (CQCN) is a congestion control mechanism specifically designed for networks based on Remote Direct Memory Access (RDMA), with the goal of preventing congestion and maximizing throughput in the high-speed, low-latency networks of modern data centers. High-Precision Congestion Control (HPCC) is a congestion control algorithm aimed at achieving extremely low latency (near-zero queuing), high bandwidth utilization, and fast convergence, usually for high-performance data center networks. RTT-based congestion control is a general term for a class of congestion control algorithms that mainly infer network congestion conditions by measuring and analyzing the Round Trip Time (RTT) of data packets, rather than relying on explicit marking (Explicit Congestion Notification, ECN) or in-band telemetry of switches.

[0004] Existing network congestion control algorithms (such as the above-mentioned DCQCN, HPCC, RTT-based) have deficiencies when applied to cross-domain multi-intelligent computing cluster networks. The main problems include: (1) Adaptability issue: Although DCQCN has better hardware adaptability than HPCC and RTT-based algorithms (without special function modification of network cards and switches), the existing DCQCN implementation (such as on Nvidia network cards) cannot well adapt to the traffic patterns of intelligent computing cluster networks, resulting in the inability to fully utilize network performance; (2) Cross-domain characteristics not considered: Existing network congestion control algorithms do not optimize the configuration for the characteristics of long link delay and insufficient link bandwidth brought by cross-domain connections; (3) Low network utilization: In the cross-domain multi-intelligent computing cluster network environment, the network utilization of existing congestion control algorithms is insufficient, affecting the efficiency of tasks such as large model training. Summary of the Invention

[0005] This application provides a congestion control method, device, equipment and medium for a data center network, which can effectively solve the problems existing in the prior art.

[0006] This application provides a congestion control method for a data center network, which is applied to a first node in a first intelligent computing cluster in a multi-intelligent computing cluster network. The first node is communicatively connected to a second node, and the second node is located within the first intelligent computing cluster or in a second intelligent computing cluster different from the first intelligent computing cluster. The intelligent computing clusters are located in a data center. The first node is used to transmit target data to the second node. The network congestion control method includes: When detecting an explicit congestion notification packet returned by the second node, according to the explicit congestion notification packet, determine the traffic type of the target data stream where congestion currently occurs, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine the congestion location of the target data stream. The target data stream is any one of at least one data stream included in the target data; Determine the congestion degree and transmission progress of the target data stream according to the information of the received historical explicit congestion notification packets; Reduce the transmission rate of the target data stream according to the traffic type, congestion location, congestion degree and transmission progress of the target data stream.

[0007] According to the congestion control method provided by this application, the traffic type of the target data stream includes cross-intelligent computing cluster traffic and intra-intelligent computing cluster traffic, and the congestion location includes the link connecting the first intelligent computing cluster and the second intelligent computing cluster and the internal network of the first intelligent computing cluster. The reducing the transmission rate of the target data stream according to the traffic type, congestion location, congestion degree and transmission progress of the target data stream includes: If the traffic type of the target data stream is cross-intelligent computing cluster traffic, and the congestion location is the link connecting the first intelligent computing cluster and the second intelligent computing cluster, determine a first multiplicative slowdown factor according to the congestion degree and the transmission progress, determine a first rate according to the first multiplicative slowdown factor, and control the transmission rate of the target data stream according to the first rate; If the traffic type of the target data stream is cross-intelligent computing cluster traffic, and the congestion location is the internal network of the intelligent computing cluster, determine a second multiplicative slowdown factor according to the congestion degree and the transmission progress, determine a second rate according to the second multiplicative slowdown factor, and control the transmission rate of the target data stream according to the second rate; If the type of the communication flow is intra-intelligent computing cluster traffic, determine a third multiplicative slowdown factor according to the congestion degree and the transmission progress, determine a third rate according to the third multiplicative slowdown factor, and control the transmission rate of the target data stream according to the third rate.

[0008] According to the congestion control method provided by the present application, after reducing the transmission rate of the target data stream, the method further includes: When a speed-up event is first detected, control the transmission rate of the target data stream to enter the fast recovery mode. In the fast recovery mode, the transmission rate approaches the target rate in a logarithmic form, and the target rate is the transmission rate of the target data before reducing the transmission rate of the target data stream; When the number of detected speed-up events meets a first preset condition, control the transmission rate of the target data stream to switch from the fast recovery mode to the aggressive increase mode. In the aggressive increase mode, the transmission rate approaches the target rate with a first preset constant value as the growth amplitude; When the number of detected speed-up events meets a second preset condition, control the transmission rate of the target data stream to switch from the aggressive increase mode to the super-aggressive increase mode. In the super-aggressive increase mode, the transmission rate approaches the target rate with a second preset constant value as the growth amplitude, and the second preset constant value is greater than the first preset constant value.

[0009] According to the congestion control method provided by the present application, the method further includes: In any one of the fast recovery mode, the aggressive increase mode, and the super-aggressive increase mode, if an explicit congestion notification packet returned by the second node is detected, re-execute the step of reducing the transmission rate of the target data stream.

[0010] According to the congestion control method provided by the present application, the speed-up event is generated when the following rules are met: Within a preset time interval, no explicit congestion notification packet is detected and a preset amount of data bytes in the target data stream is successfully sent.

[0011] According to the network congestion control method provided by this application, the first node is communicatively connected to the second node through a switch, and the congestion control strategy adopted in the switch is as follows: When receiving any data packet in the target data stream, determine the length of the queue at the egress port of the switch; If the length of the queue is less than a preset minimum threshold, do not explicitly mark the data packet and send the data packet to the second node; If the length of the queue is greater than or equal to the preset minimum threshold and less than or equal to a preset maximum threshold, explicitly mark the data packet according to a marking probability matching the length of the queue, and send the processed data packet to the second node; If the length of the queue at the egress port of the switch is greater than the preset minimum threshold, explicitly mark the data packet according to the case where the marking probability is 1, and send the marked data packet to the second node.

[0012] According to the congestion control method provided by this application, the congestion control strategy adopted in the second node is as follows: After detecting the data packet carrying the explicit mark for the first time, generate an explicit congestion notification packet for the target data stream where the data packet is located, and return the explicit congestion notification packet to the first node; After an interval of a preset duration, determine whether at least one data packet carrying the explicit mark of the target data stream is received within the preset duration. If the result is yes, generate an explicit congestion notification packet for the target data stream, and return the explicit congestion notification packet to the first node.

[0013] This application also provides a congestion control device for a data center network, configured in a first node in a first intelligent computing cluster in a multi-intelligent computing cluster network. The first node is communicatively connected to a second node, and the second node is located within the first intelligent computing cluster or in a second intelligent computing cluster different from the first intelligent computing cluster. The intelligent computing cluster is located in a data center. The first node is used to transmit target data to the second node. The network congestion control device includes: A first determination module, configured to, when detecting an explicit congestion notification packet returned by the second node, determine the traffic type of the target data stream where congestion currently occurs according to the explicit congestion notification packet, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine the congestion location of the target data stream. The target data stream is any one of at least one data stream included in the target data; A second determination module, configured to determine the congestion degree and transmission progress of the target data stream according to the information of the received historical explicit congestion notification packets; A speed reduction module, configured to reduce the transmission rate of the target data stream according to the traffic type of the target data stream, the congestion location, the congestion degree, and the transmission progress.

[0014] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a congestion control method for a data center network as described in any one of the above.

[0015] This application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a congestion control method for a data center network as described in any one of the above.

[0016] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements a congestion control method for a data center network as described in any one of the above.

[0017] When implementing the congestion control method for the data center network of this application, when detecting the explicit congestion notification packet returned by the second node, according to the explicit congestion notification packet, determine the traffic type of the target data stream where congestion currently occurs, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine the congestion location of the target data stream. The target data stream is any one of at least one data stream included in the target data; then, determine the congestion degree and transmission progress of the target data stream according to the information of the received historical explicit congestion notification packets; finally, reduce the transmission rate of the target data stream according to the traffic type of the target data stream, the congestion location, the congestion degree, and the transmission progress. When the traffic type of the target data stream is cross-intelligent computing cluster traffic, this application further subdivides the congestion location of the target data stream and performs targeted control, which can effectively cope with the long link delay and insufficient link bandwidth brought by cross-domain connections, can better adapt to the traffic pattern of the intelligent computing cluster network, give full play to the network performance, and then overcome the problem of insufficient network utilization of the existing congestion control algorithm in the cross-domain multi-intelligent computing cluster network environment, and can significantly improve the efficiency of tasks such as large model training based on the intelligent computing cluster. Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic diagram of a communication architecture shown in an embodiment of the present application; Figure 2 It is a flowchart of a congestion control method for a data center network shown in an embodiment of the present application; Figure 3 It is a schematic diagram of a switch data packet marking algorithm shown in an embodiment of the present application; Figure 4 It is a schematic diagram of a state machine at the receiving end shown in an embodiment of the present application; Figure 5 It is a schematic diagram of a state machine at the sending end shown in an embodiment of the present application; Figure 6 It is a schematic diagram of a speed reduction control principle shown in an embodiment of the present application; Figure 7 It is a schematic diagram of an acceleration process shown in an embodiment of the present application; Figure 8 It is a schematic diagram of the parameters of the key mathematical formulas in the LLM-CC algorithm shown in an embodiment of the present application; Figure 9 It is a schematic diagram of the control logic of the LLM-CC algorithm shown in an embodiment of the present application; Figure 10 It is a structural block diagram of a congestion control device for a data center network shown in an embodiment of the present application; Figure 11 It is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application with reference to the drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] To solve the problems in the prior art, this application proposes a new congestion control algorithm, which is an improvement based on the DCQCN algorithm and is named LLM-CC. The LLM-CC protocol is fully implemented on the network card, mainly by modifying the protocol stacks of the network card and the switch. The LLM-CC algorithm consists of the algorithms corresponding to the sender (Sender NIC), the switch (Switch), and the receiver (Receiver NIC) respectively. Among them, the sender corresponds to the Reaction Point (RP), the Congestion Point (CP), and the receiver corresponds to the Notification Point (NP), as Figure 1 shown. Figure 1 It is a schematic diagram of a communication architecture shown in an embodiment of this application.

[0022] In this application, the sender refers to a specific computing node within the intelligent computing cluster (specifically, the network interface card NIC responsible for data transmission on this node and its running protocol stack). An intelligent computing cluster is a collection composed of many computing nodes (usually including CPUs, GPUs, etc.), high-speed interconnection network devices (switches, network card NICs), and storage systems. It is a computing environment or infrastructure. The main body that executes the sender logic (RP function) is the network interface card (NIC) of a single node and its associated drivers and software. For example, in the communication scenario of intra-DC (Intra-Data Center) within the intelligent computing cluster, the first node A (as the sender RP) in the first intelligent computing cluster sends data to the second node B (as the receiver NP) in the same cluster. Another example is in the communication scenario of Inter-DC (Inter-Data Center) across intelligent computing clusters, the first node A (as the sender RP) in the first intelligent computing cluster sends data to the second node C (as the receiver NP) in the second intelligent computing cluster 2.

[0023] The network congestion control method of this application is applied to the first node in the first intelligent computing cluster in a multi-intelligent computing cluster network. The first node is communicatively connected to the second node. The first node is used to transmit target data to the second node, and the second node is located within the first intelligent computing cluster or in a second intelligent computing cluster different from the first intelligent computing cluster. The first node serves as the sender RP, and the second node serves as the receiver NP.

[0024] Figure 2 It is a flowchart of a congestion control method for a data center network shown in an embodiment of this application. Referring to Figure 2 this, the method of this application includes: Step 210: When an Explicit Congestion Notification Packet (CNP) returned by the second node is detected, determine the traffic type of the target data flow where congestion currently occurs according to the explicit congestion notification packet. When the traffic type of the target data flow is cross-intelligent computing cluster traffic, determine the congestion location of the target data flow. The target data flow is any one of at least one data flow included in the target data.

[0025] Step 202: Determine the congestion degree and transmission progress of the target data flow according to the information of the received historical explicit congestion notification packets.

[0026] Step 203: Reduce the transmission rate of the target data flow according to the traffic type, congestion location, congestion degree, and transmission progress of the target data flow.

[0027] Before describing the control strategy of the sender of the present application, the congestion control strategy (CP algorithm) adopted in the switch is first introduced in detail. Specifically, the congestion control strategy adopted in the switch is as follows: When any data packet in the target data flow is received, determine the length of the queue at the egress port of the switch; If the length of the queue is less than the preset minimum threshold value, do not explicitly mark the data packet, and send the data packet to the second node; If the length of the queue is greater than or equal to the preset minimum threshold value and less than or equal to the preset maximum threshold value, explicitly mark the data packet according to the marking probability matching the length of the queue, and send the processed data packet to the second node; If the length of the queue at the egress port of the switch is greater than the preset minimum threshold value, explicitly mark the data packet according to the case where the marking probability is 1, and send the marked data packet to the second node.

[0028] In the present application, the CP algorithm is the same as the DCQCN algorithm. In the egress queue, if the queue length (QueueSize) exceeds the threshold, the arriving packet is ECN marked. This is done through the Random Early Detection (RED) function supported by all modern switches. Marking congestion is a probability function of the queue length, as Figure 3 shown. Two threshold values of the queue length define the marking probability. When the queue length is below the lower threshold value, the ECN bit is not marked. When the queue length exceeds the upper threshold value, all data packets transmitted from this queue are ECN marked. When the queue length is between the two threshold values, the data packets are ECN marked with a probability that linearly increases with the queue length. Figure 3 is a schematic diagram of a switch data packet marking algorithm shown in an embodiment of the present application.

[0029] Specifically, the core of the CP algorithm is to utilize the RED mechanism commonly supported by switches. Before the queue is completely filled and packet loss occurs, packets are probabilistically marked to warn of congestion.

[0030] The implementation process of the CP algorithm includes the following aspects: (1) Monitoring object: The switch monitors the real-time length of the egress queue at its outlet port (usually referring to the number of packets or bytes cached in the queue).

[0031] (2) Key thresholds: The algorithm pre-sets two key queue length thresholds, including and . is the minimum threshold. When the queue length is below this value, the network is considered unblocked. is the maximum threshold. When the queue length exceeds this value, the network is considered to have significant congestion.

[0032] (3) Marking logic (judging according to queue length): When a packet arrives at the switch and is about to enter the egress queue, the switch checks the current length of the queue and decides whether to perform an ECN mark on the packet according to the following rules: A: , that is, the queue length is less than the minimum threshold, indicating that the network is unblocked. The switch does not perform an ECN mark on the packet, and the marking probability is 0, as shown by the middle line segment from 0 to Figure 3 in .

[0033] B: , that is, the queue length is between the minimum threshold and the maximum threshold, indicating the early warning stage of congestion. The switch probabilistically performs an ECN mark on the packet, and this probability increases linearly with the increase of the queue length. The marking probability increases linearly from 0 (when ) to (when ), as shown by the slanted line part from point Figure 3 to point to point in is usually a value close to 1, or can be 1 according to the specific configuration of RED.

[0034] C: , that is, when the queue length exceeds the maximum threshold value, it indicates that significant network congestion has occurred. The switch deterministically (i.e., with a 100% probability) marks all packets arriving at this queue with ECN. The marking probability is 1. As Figure 3 shown by the horizontal line after

[0035] In summary, through the RED mechanism, the CP algorithm uses and these two thresholds to divide the queue state into three regions: the unobstructed region (no marking), the warning region (probabilistic marking), and the congestion region (certain marking). Thus, when the network just starts to show signs of congestion (when the queue length is between and ), it can notify the receiving end through ECN marking and affect the sending strategy of the sending end through the receiving end, rather than waiting until the queue is completely full before dropping packets, thereby achieving earlier congestion control.

[0036] Next, the control strategy of the receiving end NP of this application is introduced. Specifically, the control strategy of the receiving end (NP algorithm) is as follows: After detecting a packet with an explicit mark for the first time, generate an explicit congestion notification packet for the target data stream where the packet is located and return the explicit congestion notification packet to the first node; After an interval of a preset duration, determine whether at least one packet with an explicit mark of the target data stream is received within the preset duration. If the result is yes, generate an explicit congestion notification packet for the target data stream and return the explicit congestion notification packet to the first node.

[0037] In this application, in the NP algorithm, after the receiving end NP receives a packet with an ECN mark, it determines that network congestion has occurred, and then transmits this information back to the sending end RP. The RoCEv2 standard defines an explicit congestion notification packet (Congestion Notification Packet, CNP) for this. The NP algorithm is used to specify how and when to generate a CNP.

[0038] For each data stream, this algorithm follows the Figure 4 shown state machine. If a marked packet of a data stream arrives at the receiving end and no CNP has been sent for this data stream in the most recent N microseconds, immediately send a CNP. Then, if any packet arriving within this time window is marked, the network card generates at most one CNP packet for this data stream every N microseconds. This application uses in the deployment. Processing marked packets and generating CNP is costly, so this application minimizes the activity of each marked packet. Figure 4 is a schematic diagram of the state machine of a receiving end shown in an embodiment of this application.Figure 4 In 。

[0039] In this application, the NP algorithm is a key component deployed at the receiving end in the LLM-CC (and its underlying DCQCN) congestion control framework. Its core objective is to monitor the ECN markings set by network switches (CP) in the incoming data stream and, when a congestion signal is detected, feedback congestion information to the sender (RP) of the data stream according to specific rules so that the sender can adjust its sending rate accordingly.

[0040] The operation of the NP algorithm is based on each independent data stream, and corresponding status information is maintained for each data stream. Its main implementation process revolves around the generation and sending logic of CNP, and this logic strictly follows the rate limit principle. The specific steps are as follows: (1) Congestion event detection: When the receiving end (NP) receives a packet marked with ECN congestion experience (ECN-Marked) in its IP header or transport protocol header, this event is regarded as a congestion indication signal.

[0041] (2) State variables and timers: For each data stream, at least one state variable and one timer are maintained internally in NP. The state variable is used to record whether a CNP has been sent for this data stream within the most recent time window, and the timer is used to implement the rate limit for CNP generation.

[0042] (3) CNP generation rule - first trigger: Judgment condition: When it is detected that a certain data stream (such as the target data stream) receives an ECN-marked packet, the NP algorithm first checks the status variable associated with this data stream to determine whether a CNP has been sent for this data stream within the time window of the most recent N microseconds.

[0043] Trigger action: If the judgment result is negative (i.e., no CNP has been sent within N microseconds), the NP algorithm immediately generates a CNP for this data stream and sends it back to the sender (RP) of this data stream through the network. At the same time, update the status variable of this stream (marked as "sent") and start (or reset) the N-microsecond timer.

[0044] (4) CNP generation rule - subsequent suppression and rate limit: Suppression phase: After successfully sending a CNP and starting the N-microsecond timer, enter a suppression period. Within this N microseconds, even if this data stream receives more ECN-marked packets, the NP algorithm will not immediately generate a new CNP.

[0045] Periodic Check and Generation: When the N-microsecond timer expires, the NP algorithm checks whether at least one ECN-marked packet of this data stream has been received within the just-passed N-microsecond time window. If the check result is yes, the NP algorithm generates at most one new CNP for this data stream and sends it to the sender, then resets the N-microsecond timer again and enters the next suppression and check cycle; if the check result is no, the NP algorithm does not generate a new CNP, and the state can return to the state of waiting for the first trigger. This mechanism ensures that for any data stream that continuously experiences congestion, the sending frequency of CNP is strictly limited to no more than one per N microseconds.

[0046] (5)Parameter Configuration and Optimization: The time window parameter N is configurable. For example, it can be set to 50 microseconds.

[0047] In summary, through monitoring ECN markings, state management based on each data stream, and a strict N-microsecond time window rate limit mechanism, the NP algorithm realizes effective detection and feedback of network congestion signals, and can ensure that congestion information can be transmitted to the sender in a timely manner (through the first trigger mechanism) and efficiently (through the rate limit mechanism).

[0048] In this application, the traffic types of the target data stream include cross-intelligent computing cluster traffic and intra-intelligent computing cluster traffic, and the congestion locations include the link connecting the first intelligent computing cluster and the second intelligent computing cluster and the internal network of the first intelligent computing cluster. Accordingly, step 203 may include: Step 2031: If the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the link connecting the first intelligent computing cluster and the second intelligent computing cluster, determine the first multiplicative slowdown factor according to the congestion degree and transmission progress, determine the first rate according to the first multiplicative slowdown factor, and control the transmission rate of the target data stream according to the first rate; Step 2032: If the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the internal network of the intelligent computing cluster, determine the second multiplicative slowdown factor according to the congestion degree and transmission progress, determine the second rate according to the second multiplicative slowdown factor, and control the transmission rate of the target data stream according to the second rate; Step 2033: If the type of the communication flow is intra-intelligent computing cluster traffic, determine the third multiplicative slowdown factor according to the congestion degree and transmission progress, determine the third rate according to the third multiplicative slowdown factor, and control the transmission rate of the target data stream according to the third rate.

[0049] Next, the control strategy (RP algorithm) of the sender (i.e., the first node) of this application will be introduced in detail.

[0050] 1. : In this application, time is divided into configurable time slots. Each time slot indicates whether a CNP arrives within that time slot. (i.e., ) The parameter is a continuously changing average value, which is the proportion of time slots in which CNP arrives (if more than one CNP arrives in the same time slot, it has the same effect as only one CNP arriving). At the end of each time slot, is updated by the following formula , where g is a constant parameter between 0 and 1, and CNP_arrived is a bit field used to indicate whether a CNP arrived in the previous time slot.

[0051] In this application, the function is a core component of the congestion control logic of the sender (RP). Its main goal is to dynamically calculate and maintain a key state parameter . This parameter is designed to quantify and smoothly reflect the degree of congestion experienced on the recent network path, based on the arrival frequency of congestion notification packets (CNPs) fed back by the receiver (NP). The calculated value will be used as the input basis for the RP algorithm to make subsequent rate adjustment decisions.

[0052] The implementation of the function is based on the following key mechanisms: (1) Time discretization and periodic execution: The algorithm divides continuous time into discrete, configurable-length "time slots". Figure 5 Shown in the RP state machine diagram in is used to define these time slots. The function is called and executed periodically, and its trigger event is expiration ( ), that is, it is executed once at the end of each time slot to calculate and update the value. Figure 5 is a schematic diagram of the state machine of a sender shown in an embodiment of this application.

[0053] (2) Congestion signal monitoring: In each time slot, the RP monitors whether at least one CNP arrives from the receiver. For the calculation of the value, only the presence of CNP in that time slot is concerned, rather than the specific number of CNPs that arrive.

[0054] (3) State variable : The RP maintains a boolean (or bit) state variable , at the end of each time slot, when Previously, this variable was set according to the monitoring results during the just-ended time interval: If a CNP arrived during the interval, it was set to 1; if no CNP arrived, it was set to 0.

[0055] (4)Exponential Moving Average (EMA) Update: The parameter itself represents an exponential moving average value, reflecting the proportion of the time interval containing CNP arrival events in the recent history. The function updates the value using the following formula: . is the new value calculated after the end of the current time interval. Among them, represents the old value calculated at the end of the previous time interval, g represents a pre-configured smoothing factor , which determines the weights of the historical value and the current interval state. The larger the g value, the greater the historical influence, and the smoother the change; conversely, it is more sensitive to recent changes. represents the indication value (0 or 1) determined in the above steps indicating whether a CNP arrived in the most recent time interval.

[0056] The EMA formula effectively combines the historical average congestion level ( ) with the latest congestion indication . If a CNP arrived in the most recent interval , then will be adjusted towards 1 (increase or slow down the decline); if no CNP arrived , then will be adjusted towards 0 (decrease or slow down the rise).

[0057] The RP algorithm converts discrete CNP arrival events into a continuous and smoothly changing congestion degree indication parameter through the function, using monitoring based on a fixed time interval and exponential moving average calculation. The parameter can reflect the recent congestion trend and is used by the RP algorithm for subsequent more refined rate control decisions, thereby achieving an adaptive response to the network congestion situation.

[0058] 2. Rate Reduction Strategy.

[0059] In this application, first, according to the congestion area is divided into two types, the traffic across intelligent computing clusters and the traffic within the intelligent computing cluster Then, based on the BTH identifier, the congestion of the inter-DC traffic on the links between data centers (intelligent computing clusters) is detected and the congestion of the inter-DC traffic on the links within the data centers (intelligent computing clusters) is detected Different traffic throttling strategies are executed respectively

[0060] CutRate represents traffic throttling. The CutRate throttling process is triggered when the sender (RP) receives the congestion notification packet (CNP) sent back by the receiver (NP). The core objective of this process is to respond to the detected network congestion by reducing the sending rate to alleviate the congestion. The key feature of LLM-CC is its implementation of a differentiated traffic throttling strategy, that is, different rate adjustment logics are applied according to the inferred type and location of the congested link

[0061] The execution process of CutRate follows a structured decision-making process as follows (1) Preliminary determination of traffic type (based on ): Distinguish whether the current congested communication flow (queue pair, QP) belongs to the traffic type across intelligent computing clusters / inter-data centers ( ) or the traffic type within the intelligent computing cluster / data center . The main basis for the determination is the baseline round-trip time of this data flow . By comparing with a preset threshold, long delays (usually corresponding to ) and short delays (usually corresponding to ) of the connection are distinguished. This determination corresponds to the first diamond decision node in the process Figure 6 , and its output classifies the traffic as or . Figure 6 Figure 36 is a schematic diagram of a traffic throttling control principle shown in an embodiment of the present application

[0062] (2) Refinement of congestion location (for , based on BTH): If the traffic type is determined to be in the previous step, the possible location where the congestion occurs needs to be further refined. This step uses specific information in the packet transmission protocol header, that is, the BTH identifier (corresponding to Figure 6 in the process ). By parsing the BTH information, the algorithm infers whether this congestion event occurs on the long-distance link connecting different data centers (intelligent computing clusters) (marked as ) or in the internal network of the remote data center (intelligent computing cluster) (marked as ). This determination corresponds to the second diamond decision node below in the process Figure 6 , only when Execute on the branch.

[0063] (3)Select and execute a specific rate reduction processing function: Based on the congestion scenario classification completed in the above steps, the algorithm will select and call one of the three independent rate reduction processing functions: : When it is determined that and the congestion type is call.

[0064] : When it is determined that and the congestion type is call.

[0065] : When the first step is determined to be and congestion is detected, call.

[0066] Each function encapsulates the multiplicative decrease (MD) logic customized for the corresponding congestion scenario. This process involves calculating a specific multiplicative decrease factor , which depends on parameters such as the congestion awareness (i.e., the degree of congestion), the transmission progress Ratio, etc., and applying this factor to reduce the current sending rate . The calculation formula or parameters used by different processing functions can be different to achieve different rate reduction effects.

[0067] In summary, 's rate reduction mechanism maps the received congestion signal to three specific congestion scenarios through a two-stage classification process (first distinguishing traffic flow, and then further dividing the congestion location for traffic). Subsequently, for each scenario, a specially designed processing function is called to perform differentiated multiplicative rate reduction operations. This refined processing aims to more precisely respond to different types of network congestion, minimizing the impact on non-congested paths or traffic with less contribution to congestion while alleviating congestion, thereby improving the overall network performance and resource utilization.

[0068] In this application, after step 203, the method further includes: When a speed increase event is first detected, control the transmission rate of the target data stream to enter the fast recovery mode. In the fast recovery mode, the transmission rate approaches the target rate logarithmically, and the target rate is the transmission rate of the target data before reducing the transmission rate of the target data stream; When the number of detected speed-up events meets the first preset condition, control the transmission rate of the target data stream to switch from the fast recovery mode to the aggressive increase mode. In the aggressive increase mode, the transmission rate approaches the target rate with a first preset constant value as the growth amplitude. When the number of detected speed-up events meets the second preset condition, control the transmission rate of the target data stream to switch from the aggressive increase mode to the super-aggressive increase mode. In the super-aggressive increase mode, the transmission rate approaches the target rate with a second preset constant value as the growth amplitude, and the second preset constant value is greater than the first preset constant value.

[0069] Among them, in any of the fast recovery mode, the aggressive increase mode, and the super-aggressive increase mode, if an explicit congestion notification packet returned by the second node is detected, re-execute the step of reducing the transmission rate of the target data stream.

[0070] In one implementation, a speed-up event is generated when the following rules are met: No explicit congestion notification packet is detected within a preset time interval and a preset amount of data bytes in the target data stream is successfully sent.

[0071] The speed-up strategy of the sender is introduced in detail below.

[0072] In this application, the speed-up logic is divided into three sequential stages: fast recovery, aggressive increase (keep probing), and super-aggressive increase (keep probing).

[0073] Moving from one stage to the next is defined by the number of speed-up event parameters counted in that stage. After the number of speed-up events in a stage exceeds a predefined threshold, the logic moves to the next stage. A speed-down event resets all counters related to speed-up and returns to the fast recovery stage. In addition, once speeded up, before speeded down, the current speed is saved in a parameter called Since the last speed-up, after a predefined time interval or a predefined number of sent bytes, if no speed-down event occurs, a speed-up event will occur. When in the fast recovery stage, for each speed-up event, the speed increases by half of the distance to (that is, logarithmically approaching, . This allows for a quick recovery to the speed at which congestion occurred at the beginning of the fast recovery stage, and then a more cautious increase in speed when the speed approaches the speed at which congestion occurred. In the latter two stages, once a speed-up event occurs, the speed increases by a constant value. This can obtain throughput when the bandwidth is released.

[0074] Specifically, in this application, the speed-up logic of the algorithm is at the sender Perform a speed - reduction operation once After that, and no congestion notification packets are received within a subsequent observation period It is activated. Its overall goal is to carefully increase the sending rate after confirming the alleviation or elimination of network congestion to detect and utilize potential available network bandwidth. The speed - up process is not continuous but is driven by discrete speed - up events .

[0075] In this application, a speed - up event is defined as: after the last rate adjustment (either speed - up or speed - reduction), satisfying any of the following conditions and no are received during this period: 1) A predefined time gap has passed (corresponding to Figure 7 in the process ); 2) A predefined number of bytes have been successfully sent (corresponding to Figure 7 in the process ). Figure 7 is a schematic diagram of a speed - up process shown in an embodiment of this application. Wherein, T is used to count the cumulative number of detected speed - up events, and BC is used to count the cumulative number of times that a predefined number of bytes have been successfully sent

[0076] The speed - up process of LLM - CC is designed into three sequentially - executed stages to balance the rate recovery speed and network stability: Stage 1: Quick recovery .

[0077] Entry condition: This is the default initial stage and is entered when a speed - up event is first triggered after each speed - reduction event

[0078] Core goal: Quickly restore the sending rate to a level close to that before the last congestion occurred. The algorithm uses a state variable , which stores the sending rate value before the last execution of operation

[0079] Rate update mechanism: In this stage, each time a speed - up event occurs, the current sending rate ( or RC) is updated according to the following formula: This way makes the rate approach logarithmically, that is, the initial increment is large, and as approaches , the increment gradually decreases, showing the characteristics of being fast at the beginning and cautious at the end

[0080] Stage 2: Aggressive increase ( / Keep probing).

[0081] Entry condition: When, during the fast recovery phase, it is determined that the first preset condition is met based on the number of speed-up events that have occurred cumulatively, the algorithm logic migrates to this phase, or as in Figure 7 wherein, when (i.e., the first preset condition) is met, the algorithm logic migrates to this phase.

[0082] Core objective: After the rate has recovered to a certain level, linearly detect the available bandwidth in a stable and relatively conservative manner.

[0083] Rate update mechanism: During this phase, each time a speed-up event occurs, the current transmission rate increases by a fixed small constant value . That is .

[0084] Phase 3: Ultra-aggressive increase ( / keep detecting).

[0085] Entry condition: When, during the aggressive increase phase, it is determined that the second preset condition is met based on the number of speed-up events that have occurred cumulatively, the algorithm logic migrates to this phase, or as in Figure 7 wherein, when (i.e., the second preset condition) is met, the algorithm logic migrates to this phase.

[0086] Core objective: When the network continues to perform stably, actively detect potential higher bandwidth at a faster rate.

[0087] Rate update mechanism: During this phase, each time a speed-up event occurs, the current transmission rate increases by a fixed large constant value . That is .

[0088] Phase transition and reset mechanism: Phase transition: The condition for transitioning from one phase to the next higher phase is determined based on the cumulative number of consecutive "speed-up events" that have occurred within that phase and the cumulative number of times a predefined number of bytes have been successfully transmitted. The first preset condition for transitioning from Phase 1 to Phase 2 is , and the second preset condition for transitioning from Phase 2 to Phase 3 is .

[0089] Global reset: At any time, once the sender (RP) receives a CNP (i.e., a "speed-down event" occurs), the entire three-phase speed-up logic will be immediately interrupted and reset. All states related to speed-up (including the current phase, the cumulative speed-up event counter T, BC, etc.) will be restored to the initial state. The next time the "speed-up event" condition is met, it will start executing again from the fast recovery of Phase 1.

[0090] In summary, the speed-up mechanism of LLM-CC achieves the rate increase through a structured three-stage process (quick recovery actively increasing super-actively increasing). This mechanism utilizes for quick initial recovery and then continuously probes the bandwidth through two additive increments with different magnitudes. The transition between stages is controlled by the cumulative number of non-interrupted speed-up events, and the occurrence of any congestion signal (CNP) triggers an immediate reset, ensuring the quick response and adaptability of the speed-up process to network congestion changes. This design aims to efficiently utilize network bandwidth while maintaining the stability of the probing process.

[0091] Next, according to Figure 8 , the meanings of the key mathematical formulas in the LLM-CC algorithm are explained in detail. Figure 8 is a schematic diagram of the parameters of the key mathematical formulas in the LLM-CC algorithm shown in an embodiment of this application.

[0092] (1) . Ratio is the transmission progress, representing the proportion of the data volume that has been sent in the current data transmission task to the total data volume that needs to be sent. For example, if a large file is 100MB in total and 30MB has been sent, then Ratio is 0.3 (or 30%). Ratio is used to adjust the intensity of speed reduction or speed increase.

[0093] (2) .

[0094] When applying this formula, calculate the multiplicative speed reduction factor and update the current rate . That is, the recent congestion degree perceived by the sender in the previous text.

[0095] (3) .

[0096] When applying this formula, calculate the multiplicative speed reduction factor and update the current rate .

[0097] (4) .

[0098] When applying this formula, .

[0099] (5) .

[0100] ​When applying this formula, . RC is the current rate, and RT is .

[0101] (6) .

[0102] When applying this formula, the additive growth rate adjustment factor , the target rate , the current actual rate . θ is a growth rate factor, used to fine-tune the increased step size. In the active increase phase, the target rate RT itself increases by an amount equal to a basic smaller growth rate step multiplied by the adjustment factor .

[0103] (7) .

[0104] When applying this formula, the additive growth rate adjustment factor , the target rate , the current actual rate . In the ultra-active increase phase, the amount by which the target rate RT increases is equal to a basic larger growth rate step multiplied by the adjustment factor .

[0105] Next, according to Figure 9 , the core control logic in the LLM-CC algorithm will be explained in detail. Figure 9 is a schematic diagram of the control logic of the LLM-CC algorithm shown in an embodiment of this application.

[0106] In Figure 9 , the pseudocode defines the core logic of the LLM-CC congestion control algorithm at the sender (RP). It integrates two main functions: congestion control (rate adjustment) and priority scheduling. The algorithm is implemented through three independent functions: for handling congestion control of cross-data center connections, for handling congestion control of intra-data center connections, and for implementing priority scheduling based on application (such as AI training) characteristics. In Figure 9 , respectively represent the dynamic adjustment factors calculated in multiplicative decrease and additive increase, reflects the smoothed average of the recent congestion level, is a factor used to adjust the effect of additive increase, is the current sending rate RC, It is the basic growth rate step length for the active increase and super active increase stages.

[0107] (function , lines 1-11): This function is responsible for handling cross-data center ( ) Queue pair Congestion control logic, that is, rate adjustment for long-distance, high-latency connections. The execution process is as follows: Line 3: Based on the ECN flag of the received data packet and basic transport header information , determine the current congestion type This corresponds to In the process The congestion location of the traffic is broken down into steps.

[0108] Line 4: Calculate the completion ratio of the current transfer task .

[0109] Lines 5-6 (processing ): If congestion is determined to occur on the link between data centers: Calculate the multiplicative speed reduction factor , the formula is . Apply multiplicative speed reduction: Current rate According to the calculated and congestion awareness To reduce.

[0110] Lines 7-8 (processing ): If congestion is determined to occur on the internal link of the remote data center: Calculate the multiplicative speed reduction factor , the formula is . Apply multiplicative speed reduction: .

[0111] Lines 9-11 (Processing ): If it is determined that there is no congestion, calculate the additive growth factor , the formula is , and indicate , indicating that Mainly determined by Ratio. Apply additive growth rate: The current rate Rcur increases by an amount that is determined by the basic speed increase step length. and regulatory factors This corresponds to the operation of the aggressive increasing or super aggressive increasing stage in the ramp-up logic.

[0112] (function , line 12 - 21): This function is responsible for handling the congestion control logic within the data center queue pair , that is, the rate adjustment of short - distance and low - latency connections. The execution process is as follows: Line 14: Obtain the congestion type (usually or ).

[0113] Line 15: Calculate the transmission progress .

[0114] Lines 16 - 18 (processing ): If it is determined that there is no congestion, calculate the additive increase factor , and apply additive increase: .

[0115] Lines 19 - 21 (processing ): If it is determined that there is congestion, calculate the multiplicative decrease factor , and the formula is . Apply multiplicative decrease: .

[0116] (Function , lines 22 - 27): This function is independent of congestion control and is responsible for prioritizing the packets to be sent according to the characteristics of the application (especially AI model training). The execution process is as follows: Line 24: Obtain the information of the current communication round (CommRound) from the application layer (Model_Framework).

[0117] Line 25: Calculate a priority value (Prior) based on the communication round CommRound. Here, the modulo operation is used , indicating that the priority may change periodically.

[0118] Line 26: Set the calculated priority value Prior to a specific field in the packet header (such as the DSCP field in the IP header).

[0119] Line 27: Put this packet with the priority mark into the corresponding hardware transmit queue on the network card (NIC) . The network card hardware will then determine the actual transmission order of the packets according to the queue priority, aiming to achieve fine - grained scheduling of specific data streams (such as different data blocks in AI training), give priority to sending critical data, and ensure the synchronization and efficiency of the application layer (such as model training).

[0120] In summary,Figure 9 In it, LLM-CC depicts an integrated sender algorithm framework. It, through and functions, realizes differential congestion awareness and rate adjustment (including multiplicative decrease and additive increase) based on the connection type and the congestion location . Meanwhile, through the function, an application-aware priority scheduling mechanism is introduced, and data packets are assigned to different sending queues according to their importance in a specific application (such as an AI training round). The combination of these two aspects aims to comprehensively optimize the end-to-end transmission performance of high-performance applications in a complex network environment (especially across multiple intelligent computing clusters in different domains).

[0121] When the traffic type of the target data stream in this application is cross-intelligent computing cluster traffic, the congestion location of the target data stream is further segmented and controlled accordingly, which can effectively cope with the long link delay and insufficient link bandwidth caused by cross-domain connections, can better adapt to the traffic pattern of the intelligent computing cluster network, give full play to the network performance, and then overcome the problem of insufficient network utilization of existing congestion control algorithms in the network environment of cross-domain multiple intelligent computing clusters, and can significantly improve the efficiency of tasks such as large model training based on intelligent computing clusters.

[0122] Next, the congestion control device for the data center network provided by this application will be described. The congestion control device described below can be correspondingly referred to the congestion control method described above.

[0123] Figure 10 is a structural block diagram of a congestion control device for a data center network shown in an embodiment of this application. Referring to Figure 10 , the congestion control device 1000 of this application is deployed on the first node in the first intelligent computing cluster in the multi-intelligent computing cluster network. The first node is communicatively connected to the second node, and the second node is located within the first intelligent computing cluster or in a second intelligent computing cluster different from the first intelligent computing cluster. The intelligent computing cluster is located in the data center. The first node is used to transmit target data to the second node. The network congestion control device 1000 includes: A first determination module 1001, configured to, when detecting an explicit congestion notification packet returned by the second node, determine the traffic type of the target data stream where congestion currently occurs according to the explicit congestion notification packet, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine the congestion location of the target data stream. The target data stream is any one of at least one data stream included in the target data; A second determination module 1002, configured to determine the congestion degree and transmission progress of the target data stream according to the information of the received historical explicit congestion notification packets; The speed reduction module 1003 is used to reduce the transmission rate of the target data stream according to the traffic type of the target data stream, the congestion location, the congestion degree, and the transmission progress.

[0124] According to the congestion control device 1000 of the present application, the traffic type of the target data stream includes cross-intelligent computing cluster traffic and intra-intelligent computing cluster traffic, and the congestion location includes the link connecting the first intelligent computing cluster and the second intelligent computing cluster and the internal network of the first intelligent computing cluster; the speed reduction module 1003 includes: The first speed reduction sub-module is used to determine a first multiplicative speed reduction factor according to the congestion degree and the transmission progress if the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the link connecting the first intelligent computing cluster and the second intelligent computing cluster, determine a first rate according to the first multiplicative speed reduction factor, and control the transmission rate of the target data stream according to the first rate; The second speed reduction sub-module is used to determine a second multiplicative speed reduction factor according to the congestion degree and the transmission progress if the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the internal network of the intelligent computing cluster, determine a second rate according to the second multiplicative speed reduction factor, and control the transmission rate of the target data stream according to the second rate; The third speed reduction sub-module is used to determine a third multiplicative speed reduction factor according to the congestion degree and the transmission progress if the type of the communication stream is intra-intelligent computing cluster traffic, determine a third rate according to the third multiplicative speed reduction factor, and control the transmission rate of the target data stream according to the third rate.

[0125] According to the congestion control device 1000 of the present application, it further includes: The first acceleration module is used to control the transmission rate of the target data stream to enter the fast recovery mode when a speed increase event is first detected. In the fast recovery mode, the transmission rate approaches the target rate in a logarithmic form, and the target rate is the transmission rate of the target data before reducing the transmission rate of the target data stream; The second acceleration module is used to control the transmission rate of the target data stream to switch from the fast recovery mode to the active increase mode when the number of detected speed increase events meets a first preset condition. In the active increase mode, the transmission rate approaches the target rate with a first preset constant value as the growth amplitude; A third acceleration module, configured to control the transmission rate of the target data stream to switch from the positive increase mode to the ultra-positive increase mode when the number of detected speed-up events meets a second preset condition. In the ultra-positive increase mode, the transmission rate approaches the target rate with a second preset constant value as the growth amplitude, and the second preset constant value is greater than the first preset constant value.

[0126] The congestion control device 1000 according to the present application further includes: A monitoring module, configured to re-perform the step of reducing the transmission rate of the target data stream if an explicit congestion notification packet returned by the second node is detected in any one of the fast recovery mode, the positive increase mode, and the ultra-positive increase mode.

[0127] According to the congestion control device 1000 of the present application, the speed-up event is generated when the following rules are met: No explicit congestion notification packet is detected within a preset time interval and a preset byte amount of the target data stream is successfully transmitted.

[0128] According to the congestion control device 1000 of the present application, the first node is communicatively connected to the second node through a switch, and the congestion control policy adopted in the switch is as follows: When receiving any data packet in the target data stream, determine the length of the queue at the exit end of the switch; If the length of the queue is less than a preset minimum threshold, do not explicitly mark the data packet and send the data packet to the second node; If the length of the queue is greater than or equal to the preset minimum threshold and less than or equal to a preset maximum threshold, explicitly mark the data packet according to a marking probability matching the length of the queue, and send the processed data packet to the second node; If the length of the queue at the exit end of the switch is greater than the preset minimum threshold, explicitly mark the data packet in the case where the marking probability is 1, and send the marked data packet to the second node.

[0129] According to the congestion control device 1000 of the present application, the congestion control policy adopted in the second node is as follows: After detecting a data packet carrying the explicit mark for the first time, generate an explicit congestion notification packet for the target data stream where the data packet is located, and return the explicit congestion notification packet to the first node; After an interval of a preset duration, determine whether at least one data packet carrying the explicit mark of the target data stream is received within the preset duration. If the result is yes, generate an explicit congestion notification packet for the target data stream and return the explicit congestion notification packet to the first node.

[0130] Figure 11 FIG. 4 is a schematic diagram of the physical structure of an electronic device according to an embodiment of the present application. As Figure 11 shown, the electronic device may include: a processor 1110, a communication interface 1120, a memory 1130, and a communication bus 1140. Among them, the processor 1110, the communication interface 1120, and the memory 1130 communicate with each other through the communication bus 1140. The processor 1110 may call logical instructions in the memory 1130 to execute a network congestion control method.

[0131] In addition, when the logical instructions in the above-mentioned memory 1130 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0132] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a congestion control method for a data center network provided by the above-mentioned various methods.

[0133] On the other hand, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute a congestion control method for a data center network provided by the above-mentioned various methods.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A congestion control method for a data center network, characterized in that Applied to the first node in the first intelligent computing cluster in a multi-intelligent computing cluster network, the first node is communicatively connected to a second node, the second node is located within the first intelligent computing cluster or in a second intelligent computing cluster different from the first intelligent computing cluster, the intelligent computing cluster is located in a data center, and the first node is used to transmit target data to the second node. The network congestion control method includes: When detecting an explicit congestion notification packet returned by the second node, according to the explicit congestion notification packet, determine the traffic type of the target data stream where congestion currently occurs, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine the congestion location of the target data stream. The target data stream is any one of at least one data stream included in the target data; According to the information of the received historical explicit congestion notification packets, determine the congestion degree and transmission progress of the target data stream; According to the traffic type, congestion location, congestion degree, and transmission progress of the target data stream, reduce the transmission rate of the target data stream.

2. The congestion control method according to claim 1, characterized in that The traffic type of the target data stream includes cross-intelligent computing cluster traffic and intra-intelligent computing cluster traffic, and the congestion location includes the link connecting the first intelligent computing cluster and the second intelligent computing cluster and the internal network of the first intelligent computing cluster. The reducing the transmission rate of the target data stream according to the traffic type, congestion location, congestion degree, and transmission progress of the target data stream includes: If the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the link connecting the first intelligent computing cluster and the second intelligent computing cluster, determine a first multiplicative slowdown factor according to the congestion degree and transmission progress, determine a first rate according to the first multiplicative slowdown factor, and control the transmission rate of the target data stream according to the first rate; If the traffic type of the target data stream is cross-intelligent computing cluster traffic and the congestion location is the internal network of the intelligent computing cluster, determine a second multiplicative slowdown factor according to the congestion degree and transmission progress, determine a second rate according to the second multiplicative slowdown factor, and control the transmission rate of the target data stream according to the second rate; If the type of the communication flow is intra-intelligent computing cluster traffic, determine a third multiplicative slowdown factor according to the congestion degree and transmission progress, determine a third rate according to the third multiplicative slowdown factor, and control the transmission rate of the target data stream according to the third rate.

3. The congestion control method according to claim 1, wherein After reducing the transmission rate of the target data stream, the method further includes: When detecting a speed-up event for the first time, control the transmission rate of the target data stream to enter a fast recovery mode. In the fast recovery mode, the transmission rate approaches the target rate in a logarithmic form, and the target rate is the transmission rate of the target data before reducing the transmission rate of the target data stream; When the number of detected speed-up events meets the first preset condition, control the transmission rate of the target data stream to switch from the fast recovery mode to the actively increasing mode. In the actively increasing mode, the transmission rate approaches the target rate with a first preset constant value as the growth amplitude. When the number of detected speed-up events meets the second preset condition, control the transmission rate of the target data stream to switch from the actively increasing mode to the super actively increasing mode. In the super actively increasing mode, the transmission rate approaches the target rate with a second preset constant value as the growth amplitude, and the second preset constant value is greater than the first preset constant value.

4. The congestion control method according to claim 3, wherein The method further includes: In any one of the fast recovery mode, the actively increasing mode, and the super actively increasing mode, if an explicit congestion notification packet returned by the second node is detected, re-execute the step of reducing the transmission rate of the target data stream.

5. The congestion control method according to claim 3, wherein The speed-up event is generated when the following rules are met: No explicit congestion notification packet is detected within a preset time interval and a preset amount of data bytes in the target data stream is successfully sent.

6. The congestion control method according to claim 1, wherein The first node is communicatively connected to the second node through a switch, and the congestion control policy adopted in the switch is as follows: When any data packet in the target data stream is received, determine the length of the queue at the outlet end of the switch; If the length of the queue is less than the preset minimum threshold, do not explicitly mark the data packet and send the data packet to the second node; If the length of the queue is greater than or equal to the preset minimum threshold and less than or equal to the preset maximum threshold, explicitly mark the data packet according to the marking probability matching the length of the queue, and send the processed data packet to the second node; If the length of the queue at the outlet end of the switch is greater than the preset minimum threshold, explicitly mark the data packet according to the case where the marking probability is 1, and send the marked data packet to the second node.

7. The congestion control method according to claim 6, wherein The congestion control policy adopted in the second node is as follows: After the data packet carrying the explicit mark is first detected, generate an explicit congestion notification packet for the target data stream where the data packet is located, and return the explicit congestion notification packet to the first node; After a preset time interval, determine whether at least one data packet carrying the explicit mark in the target data stream is received within the preset time interval. If the result is yes, generate an explicit congestion notification packet for the target data stream, and return the explicit congestion notification packet to the first node.

8. A congestion control device for a data center network, characterized in that The first node is configured in the first computing cluster of the multi-computing cluster network. The first node is communicatively connected to the second node. The second node is located within the first computing cluster or in a second computing cluster different from the first computing cluster. The computing cluster is located in the data center. The first node is used to transmit target data to the second node. The network congestion control device includes: A first determination module, configured to, when detecting an explicit congestion notification packet returned by the second node, determine a traffic type of a target data stream with current congestion according to the explicit congestion notification packet, and when the traffic type of the target data stream is cross-intelligent computing cluster traffic, determine a congestion location of the target data stream, where the target data stream is any one of at least one data stream included in the target data; A second determination module, configured to determine a congestion degree and a transmission progress of the target data stream according to information of received historical explicit congestion notification packets; A speed reduction module, configured to reduce a transmission rate of the target data stream according to the traffic type of the target data stream, the congestion location, the congestion degree, and the transmission progress; 9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, the processor implements a congestion control method for a data center network according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, the computer program implements a congestion control method for a data center network according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cross-data center RDMA network congestion control method and system

    CN117896324A

  • Congestion control method and device, communication system and related products

    CN118301075A

  • Congestion control method, device and equipment for intelligent calculation center network and storage medium

    CN119544627A

  • Network congestion control method, apparatus, device, and system, and storage medium

    US20230107366A1

  • Transmission rate control method and apparatus, sending device and receiving device

    WO2020042624A1