Congestion control method of artificial intelligence training cluster, chip and medium
By dynamically adjusting the packet transmission rate and the number of congestion windows, and optimizing the data transmission mechanism, the problem that traditional congestion control algorithms cannot meet the needs of artificial intelligence training clusters is solved, and efficient and reliable communication and good scalability are achieved.
Patent Information
- Application Number
- CN202510643157.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-08
AI Technical Summary
The congestion control algorithms in traditional data centers cannot meet the network needs in the artificial intelligence training cluster environment, resulting in low communication efficiency and poor scalability.
By analyzing the confirmation packets feedback from the data receiver, dynamically adjusting the packet transmission rate and the number of congestion windows, combining one-way queue delay and target queue time, optimizing the data transmission mechanism, which is suitable for decentralized deployment and avoiding modifications to switches and other network devices.
It realizes efficient and reliable communication of the artificial intelligence training cluster, has good scalability and adaptability, and can flexibly expand the network capacity and number of nodes when the cluster is expanded, solving the shortcomings of traditional congestion control algorithms in the artificial intelligence cluster environment.
Smart Images

Figure CN120281722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a congestion control method, chip and medium for an artificial intelligence training cluster. Background Art
[0002] The congestion control algorithms of traditional data centers show obvious limitations when facing artificial intelligence training workloads. Existing congestion control algorithms are mainly designed for the traffic patterns of general data centers and perform well in dealing with the diverse services of traditional data centers, but it is difficult to effectively meet the special network requirements in the artificial intelligence training environment. This is mainly because there are significant differences in traffic patterns and network architectures between traditional data centers and artificial intelligence clusters.
[0003] Traditional data centers provide a variety of network services, including web search, data storage and processing, cloud computing, and many other applications. Their network traffic shows a high degree of uncertainty, burstiness, and diversity. Therefore, traditional congestion control algorithms mainly focus on quickly responding to bursty traffic and fairly allocating bandwidth resources. However, the network communication mode of artificial intelligence training clusters is essentially different from that of traditional data centers. Existing congestion control mechanisms are mainly designed for the traffic patterns of general data centers rather than tailored for artificial intelligence workloads, resulting in the current congestion control algorithms being difficult to meet the congestion control requirements in the artificial intelligence cluster environment. Summary of the Invention
[0004] The present invention provides a congestion control method, chip and medium for an artificial intelligence training cluster to solve the problem that the congestion control algorithms of traditional data centers cannot meet the congestion control requirements in the artificial intelligence cluster environment.
[0005] According to one aspect of the present invention, there is provided a congestion control method for an artificial intelligence training cluster, which is executed by a data sender and includes:
[0006] Analyze the acknowledgment packet feedback by the data receiver;
[0007] When the reference rate carried in the acknowledgment packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjust the packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate;
[0008] Determine the number of congestion windows according to the target sending rate and the round-trip time, and determine the packet sending interval according to the number of congestion windows.
[0009] According to another aspect of the present invention, there is provided a congestion control method for an artificial intelligence training cluster, which is executed by a data receiver and includes:
[0010] Calculate the reference rate and one-way queue delay matching each data packet according to the packet headers of the data packets sent by at least one data sender, and add the reference rate and one-way queue delay matching each data packet to each acknowledgment packet;
[0011] Feedback the acknowledgment packet to the corresponding data sender so that the data sender executes the congestion control method of the corresponding artificial intelligence training cluster.
[0012] According to another aspect of the present invention, there is provided a congestion control device for an artificial intelligence training cluster, configured at a data sender, and the device includes:
[0013] An acknowledgment packet parsing module, configured to parse the acknowledgment packet fed back by the data receiver;
[0014] A first transmission rate determination module, configured to, when the reference rate carried in the acknowledgment packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjust the data packet transmission rate in stages according to the one-way queue delay and the target queuing time to obtain the target transmission rate;
[0015] A congestion control module, configured to determine the number of congestion windows according to the target transmission rate and the round-trip time, and determine the data packet transmission interval according to the number of congestion windows.
[0016] According to another aspect of the present invention, there is provided a congestion control device for an artificial intelligence training cluster, configured at a data receiver, and the device includes:
[0017] An acknowledgment packet generation module, which calculates the reference rate and one-way queue delay matching each data packet according to the packet headers of the data packets sent by at least one data sender, and adds the reference rate and one-way queue delay matching each data packet to each acknowledgment packet;
[0018] A data sending module, configured to feedback the acknowledgment packet to the corresponding data sender so that the data sender executes the congestion control method of the corresponding artificial intelligence training cluster.
[0019] According to another aspect of the present invention, there is provided an artificial intelligence chip, and the artificial intelligence chip is used to execute the congestion control method of the artificial intelligence training cluster executed by the data sender in any embodiment of the present invention, or execute the congestion control method of the artificial intelligence training cluster executed by the data receiver.
[0020] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for executing the congestion control method of the artificial intelligence training cluster performed by the data sending end according to any embodiment of the present invention, or the congestion control method of the artificial intelligence training cluster performed by the data receiving end according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, there is provided a computer program product including a computer program for implementing the congestion control method of the artificial intelligence training cluster performed by the data sending end according to any embodiment of the present invention, or the congestion control method of the artificial intelligence training cluster performed by the data receiving end according to any embodiment of the present invention.
[0022] In the technical solution of the embodiment of the present invention, by parsing the acknowledgement data packet fed back by the data receiving end at the data sending end, when the reference rate carried in the acknowledgement data packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, the packet sending rate is adjusted in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate, and then the congestion window number is determined according to the target sending rate and the round-trip time, and the packet sending interval is determined according to the congestion window number. In this solution, when the reference rate remains unchanged, the target sending rate of the packet is dynamically adjusted through the one-way queue delay and the target queuing time, and the congestion window number and the packet sending interval are adjusted in coordination, optimizing the data transmission mechanism, ensuring efficient and reliable communication between training nodes, being applicable to decentralized deployment, avoiding modification of network devices such as switches, adapting to heterogeneous environments, being able to flexibly expand the network capacity and the number of nodes when the cluster is expanded, solving the problem that the congestion control algorithms of traditional data centers cannot meet the congestion control requirements in the artificial intelligence cluster environment, being able to perform accurate congestion control on the artificial intelligence training cluster, ensuring efficient and reliable communication of training nodes, and having good scalability.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0025] Figure 1 Flowchart of a congestion control method for an artificial intelligence training cluster provided in Embodiment 1 of the present invention;
[0026] Figure 2 Flowchart of a congestion control method for an artificial intelligence training cluster provided in Embodiment 2 of the present invention;
[0027] Figure 3 Flowchart of a congestion control method for an artificial intelligence training cluster provided in Embodiment 3 of the present invention;
[0028] Figure 4 Schematic diagram of a data packet round-trip transmission path provided in Embodiment 3 of the present invention;
[0029] Figure 5 Schematic diagram of a data receiving end feeding back an ACK packet to a data sending end provided in Embodiment 3 of the present invention;
[0030] Figure 6 Schematic diagram of the structure of a congestion control device for an artificial intelligence training cluster provided in Embodiment 4 of the present invention;
[0031] Figure 7 Schematic diagram of the structure of another congestion control device for an artificial intelligence training cluster provided in Embodiment 4 of the present invention. Detailed implementation manners
[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] Embodiment 1
[0035] Figure 1 The following is a flowchart of a congestion control method for an artificial intelligence training cluster provided in Embodiment 1 of the present invention. This embodiment is applicable to the scenario of precise congestion control of an artificial intelligence training cluster, and this method can be executed by a data sender. As Figure 1 shown, the method includes:
[0036] Step 110: Analyze the acknowledgment packet feedback by the data receiver.
[0037] Among them, each data receiver in the artificial intelligence training cluster generally receives packets transmitted by multiple data senders. The data receiver receives, analyzes, and processes the data transmitted by the data sender.
[0038] In the embodiment of the present invention, the data sender in the artificial intelligence training cluster can receive the acknowledgment packet, that is, the ACK (Acknowledge Character) packet, feedback by the corresponding data receiver.
[0039] Step 120: When the reference rate carried in the acknowledgment packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjust the packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate.
[0040] Among them, the reference rate and the one-way queue delay are two types of information carried in the acknowledgment packet. The reference rate is used to describe the actual carrying capacity of the network link and the current load condition. The one-way queue delay can be used to describe the actual one-way queuing delay when the packet is transmitted forward, that is, the difference between the forward transmission path delay and the link static delay. The path from the data sender to the data receiver to send the packet is the forward transmission path. The one-way queue delay, as a congestion signal, can provide a more accurate feedback on the network congestion state and can effectively reduce the deployment complexity.
[0041] Among them, the current reference rate adjustment interval can be used to describe the time interval between the current moment and the last adjustment of the reference rate. The transmission rate adjustment interval threshold can be the time adjustment interval of the packet transmission rate when the reference rate remains unchanged. The target queuing time can be the allowed time for the packet to queue one-way. Optionally, the target queuing time can be preset in advance or dynamically adjusted based on the network topology characteristics and the number of competing flows. When a sudden congestion is detected, the target queuing time can be temporarily reduced to quickly relieve the network pressure. The target sending rate can be the packet sending rate determined by the data sender based on the currently received acknowledgment packet.
[0042] In an embodiment of the present invention, after the data receiving end parses and confirms the data packet, based on the reference rate carried in the current confirmed data packet and the reference rate in the previous confirmed data packet sent by the same data receiving end, it determines whether the reference rate has changed. If the reference rate carried in the confirmed data packet has not changed, it can further obtain the current reference rate adjustment interval and compare the current reference rate adjustment interval with the transmission rate adjustment interval threshold. If the current reference rate adjustment interval is equal to or greater than the transmission rate adjustment interval threshold, then further use the magnitude relationship between the one-way queue delay and the target queuing time as the basis for phased adjustment of the data packet sending rate, adjust the data packet sending rate (which is the reference rate in the confirmed data packet at this time), and use the adjusted data packet sending rate as the target sending rate, so that the data sending end sends data packets at the target sending rate.
[0043] Step 130: Determine the number of congestion windows according to the target sending rate and the round-trip time, and determine the data packet sending interval according to the number of congestion windows.
[0044] Among them, the round-trip time can be used to describe the time interval between the data sending end receiving the confirmed data packet and the data sending end sending the corresponding data packet. The number of congestion windows can be used to describe the size of the congestion window. The data packet sending interval can be the time interval between the data sending end sending two data packets.
[0045] In an embodiment of the present invention, the number of congestion windows can be determined according to the product value of the target sending rate and the round-trip time, and the congestion degree of the data packet can be judged based on the number of congestion windows. Furthermore, the data packet sending interval can be adjusted based on the congestion degree of the data packet to achieve fine-grained traffic control and achieve larger-scale congestion control.
[0046] Optionally, the product value of the target sending rate and the round-trip time can be directly used as the number of congestion windows, or the product value can be further multiplied by an adjustable coefficient and other forms to determine the number of congestion windows.
[0047] Exemplarily, the more serious the congestion degree of the data packet, the larger the data packet sending interval. Its specific rules can be set by itself. For example, the data packet sending interval can be set in segments based on the congestion degree of the data packet.
[0048] To improve resource utilization, this solution introduces the target queuing time and dynamically adjusts the target sending rate to make the one-way queue delay approach this target value, thus achieving a balance between network throughput and latency. The target sending rate adjustment mechanism aims to balance the responsiveness and stability of the algorithm. It has been verified that adjusting the congestion window number for each ACK packet is not the optimal choice because the effect of the new congestion window number takes at least one packet round-trip delay to be perceived by the data sender. Too frequent adjustments easily cause the system to oscillate between congestion and low utilization. Therefore, by setting the transmission rate adjustment interval threshold, the congestion window number is kept stable within the time corresponding to the transmission rate adjustment interval threshold, unless the reference rate changes, to ensure that the packets sent through the updated congestion window have sufficient time to be transmitted to the data receiver and obtain feedback.
[0049] In the technical solution of the embodiment of the present invention, by parsing the acknowledgment packet feedback by the data receiver at the data sender, when the reference rate carried in the acknowledgment packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, the packet sending rate is adjusted in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate. Then, the congestion window number is determined according to the target sending rate and the round-trip time, and the packet sending interval is determined according to the congestion window number. In this solution, when the reference rate remains unchanged, the target sending rate of the packet is dynamically adjusted through the one-way queue delay and the target queuing time, and the congestion window number and the packet sending interval are jointly adjusted, optimizing the data transmission mechanism, ensuring efficient and reliable communication between training nodes, being applicable to decentralized deployment, avoiding modification of network devices such as switches, adapting to heterogeneous environments, being able to flexibly expand the network capacity and the number of nodes during cluster expansion, solving the problem that the congestion control algorithm of the traditional data center cannot meet the congestion control requirements in the artificial intelligence cluster environment, being able to perform precise congestion control on the artificial intelligence training cluster, ensuring efficient and reliable communication of training nodes, and having good scalability.
[0050] Embodiment 2
[0051] Figure 2 It is a flowchart of a congestion control method for an artificial intelligence training cluster provided by Embodiment 2 of the present invention. This embodiment is specific based on the above embodiment and gives a specific and optional implementation manner for determining the packet sending interval according to the congestion window number. As Figure 2 shown, the method includes:
[0052] Step 210, parse the acknowledgment packet feedback by the data receiver.
[0053] In an alternative embodiment of the present invention, the congestion control method for an artificial intelligence training cluster may further include: when it is confirmed that the benchmark rate carried in the data packet changes, adjusting the data packet sending rate based on the current benchmark rate, and recording the adjustment time of the data packet sending rate.
[0054] Wherein, the current benchmark rate may be the benchmark rate currently carried in the confirmed data packet.
[0055] In an embodiment of the present invention, if the benchmark rate carried in the current confirmed data packet is different from the benchmark rate in the previous confirmed data packet sent by the same data receiver, it can be determined that the benchmark rate has changed. Then, the current benchmark rate is used as the data packet sending rate, and the update time of the data packet sending rate is recorded.
[0056] Step 220: When the benchmark rate carried in the confirmed data packet does not change and the current benchmark rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjust the data packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate.
[0057] In an alternative embodiment of the present invention, adjusting the data packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate may include: when the one-way queue delay is less than the target queuing time, adjusting the data packet sending rate to increase based on the acceleration parameter to obtain the target sending rate; when the one-way queue delay is greater than or equal to the target queuing time, adjusting the data packet sending rate to decrease based on the deceleration coefficient to obtain the target sending rate.
[0058] Wherein, the acceleration parameter may be a parameter for presetting to increase the data packet sending rate. The deceleration coefficient may be a parameter for decreasing the data packet sending rate. The deceleration coefficient is a value greater than 0 and less than 1.
[0059] Specifically, when the data sender compares that the one-way queue delay is less than the target queuing time, the acceleration parameter can be used as an addend to adjust the data packet sending rate to increase and obtain the target sending rate. When the data sender compares that the one-way queue delay is greater than or equal to the target queuing time, the deceleration coefficient is multiplied by the data packet sending rate to adjust the data packet sending rate to decrease and obtain the target sending rate.
[0060] In an alternative embodiment of the present invention, when the one-way queue delay is less than the target queuing time, based on the speed-up parameter, the packet sending rate is adjusted for speed-up to obtain the target sending rate, which may include: when the one-way queue delay is equal to zero, based on the first speed-up parameter in the speed-up parameter, the packet sending rate is adjusted for speed-up to obtain the target sending rate; when the one-way queue delay is greater than zero and less than the target queuing time, based on the second speed-up parameter in the speed-up parameter, the packet sending rate is adjusted for speed-up to obtain the target sending rate.
[0061] Wherein, both the first speed-up parameter and the second speed-up parameter belong to the speed-up parameter, and the first speed-up parameter is greater than the second speed-up parameter.
[0062] Correspondingly, when the one-way queue delay is equal to zero, the first speed-up parameter in the speed-up parameter can be used as an addend, and added to the packet sending rate to obtain the target sending rate. When the one-way queue delay is greater than zero and less than the target queuing time, the second speed-up parameter in the speed-up parameter is used as an addend and added to the packet sending rate to obtain the target sending rate.
[0063] In this solution, when the one-way queue delay is 0, it indicates that there is no queuing and congestion in the path of this packet, and the network resources are sufficient. The target sending rate enters the ultra-fast growth stage, only to improve the network utilization rate. When the one-way queuing delay is between 0 and the target queuing time, it means that there is queuing in the network but it is still relatively low. The target sending rate enters the relatively fast growth stage, and the remaining bandwidth needs to be carefully detected. When the one-way queue delay exceeds the target queuing time, it indicates that there is an obvious queuing phenomenon in the network. Corresponding deceleration measures are taken according to the degree of congestion, avoiding the problem of overshoot or undershoot that may be brought by a fixed deceleration factor. This segmented strategy not only provides more refined rate control, but also can quickly adapt to network state changes.
[0064] In an alternative embodiment of the present invention, before adjusting the packet sending rate for deceleration based on the deceleration coefficient to obtain the target sending rate, it may further include: determining an exponential reduction parameter and a multiplicative deceleration factor; calculating a first difference between the one-way queue delay and the target queuing time, and calculating a first sum value of the one-way queue delay and the static delay of the transmission link; multiplying the ratio of the first difference to the first sum value by the exponential reduction parameter to obtain a multiplicative deceleration factor to be compared; determining the deceleration coefficient according to the multiplicative deceleration factor and the multiplicative deceleration factor to be compared.
[0065] Wherein, the exponential reduction parameter may be a preset parameter for determining the deceleration coefficient. The multiplicative deceleration factor can be used to represent the maximum deceleration ratio. The first difference may be the difference between the one-way queue delay and the target queuing time. The first sum value may be the sum of the one-way queue delay and the static delay of the transmission link. The multiplicative deceleration factor to be compared may be the data compared with the multiplicative deceleration factor.
[0066] In an embodiment of the present invention, the data receiving end can obtain an exponentially decreasing parameter and a multiplicative slowdown factor set by a developer, and then subtract the static delay of the transmission link from the one-way queue delay to obtain a first difference, and perform a summation calculation on the one-way queue delay and the static delay of the transmission link to obtain a first sum value. Then, multiply the ratio of the first difference to the first sum value by the exponentially decreasing parameter to obtain a multiplicative slowdown factor to be compared. Thus, compare the multiplicative slowdown factor and the multiplicative slowdown factor to be compared to obtain the smaller value of the two, and use the complement of the smaller value as the slowdown coefficient.
[0067] Step 230: Determine the number of congestion windows according to the target sending rate and the round-trip time, and when the number of congestion windows is less than a preset window number, determine the packet sending interval based on the ratio of the round-trip time to the number of congestion windows.
[0068] Among them, the preset window number can be compared with the number of congestion windows to perform segmented adjustment on the packet sending interval. The preset window number can be set by itself according to needs. Exemplarily, the preset window number is allowed to be less than one packet, for example, it can be set to 0.0001 packets.
[0069] In an embodiment of the present invention, after determining the number of congestion windows according to the target sending rate and the round-trip time, the number of congestion windows can be further compared with the preset window number. When the number of congestion windows is less than the preset window number, the ratio of the round-trip time to the number of congestion windows can be used as the corresponding duration of the packet sending interval.
[0070] Step 240: When the number of congestion windows is greater than or equal to the preset window number, the packet sending interval is zero.
[0071] In an embodiment of the present invention, if the number of congestion windows is greater than or equal to the preset window number, it indicates that there is no blocking situation, and the packet sending interval can be set to zero to accelerate the sending of packets.
[0072] In an optional embodiment of the present invention, the congestion control method for an artificial intelligence training cluster may further include: after detecting an event of timeout for loss of acknowledgment packets, update the number of congestion windows according to the product of the multiplicative slowdown factor and the number of congestion windows, and configure the target sending rate to zero.
[0073] Among them, the event of timeout for loss of acknowledgment packets may be an event that the data sending end does not receive an acknowledgment packet within the timeout period.
[0074] In an embodiment of the present invention, when the data receiving end detects an acknowledgment packet loss timeout event, it triggers a timeout mechanism, calculates the product value of the multiplicative slowdown factor and the congestion window count, uses this product value as the congestion window count, and configures the target sending rate to zero.
[0075] To handle various network anomalies, the timeout mechanism enhances the system's robustness against physical-level failures and minimizes packet loss. This mechanism is mainly used to prevent the spread of network congestion caused by the loss of ACK packets carrying key information (benchmark rate and one-way queue delay). Although the reverse path delay of ACK packets does not affect the accuracy of congestion control, if an ACK packet is not received within an RTO (retransmission timeout), the system will determine it as a network congestion or packet loss event, that is, an acknowledgment packet loss timeout event.
[0076] After detecting an acknowledgment packet loss timeout event, the congestion window count is rapidly reduced by the multiplicative slowdown factor. This aggressive deceleration strategy can quickly alleviate potential network congestion, and reset the target sending rate of the corresponding flow to 0. This setting continues until an ACK packet carrying the benchmark rate and adjusted rate is received, and then the congestion window count can be adjusted to quickly adapt to a wide range of congestion. This reset mechanism takes into account the situation where the historical benchmark rate may not accurately reflect the current network state when a serious network failure occurs. By resetting the target sending rate, the system waits to receive a new ACK packet carrying the benchmark rate, re-establishes the understanding of the network state, and then readjusts the congestion window count to ensure accurate adaptation to the network congestion state and avoid using potentially outdated network state information.
[0077] To efficiently handle concurrent flows in large-scale artificial intelligence training scenarios, the timeout mechanism has an efficient timer management strategy. By using a hierarchical time wheel algorithm, it can efficiently manage timeout events of a large number of concurrent flows while maintaining a low central processing unit overhead. Even if the acknowledgments of multiple packets are merged into one ACK packet, this solution can still accurately track the transmission status and network feedback information of each packet. Optionally, a fault recovery mechanism can also be set. When a serious network anomaly (such as a hardware failure or severe congestion) is detected, especially when no ACK packet is received for a long time, the connection reconstruction process can be started, which is crucial for ensuring the stability of long-running artificial intelligence training tasks.
[0078] Applying the congestion control method of this solution, the packet loss rate caused by overflow in the normal operation state of the system is extremely low. Therefore, packet loss and timeout will not become the bottleneck of system performance, and it can maintain a stable and efficient state under various network conditions, providing a reliable transmission service for upper-layer applications.
[0079] The technical solution of the embodiment of the present invention analyzes the confirmation data packet fed back by the data receiving end. Then, when the reference rate carried in the confirmation data packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, the packet sending rate is adjusted in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate. Then, the congestion window number is determined according to the target sending rate and the round-trip time. When the congestion window number is less than the preset window number, the packet sending interval is determined based on the ratio of the round-trip time to the congestion window number. When the congestion window number is greater than or equal to the preset window number, the packet sending interval is zero. In this solution, when the reference rate remains unchanged, the target sending rate of the data packet is dynamically adjusted through the one-way queue delay and the target queuing time, and the congestion window number and the packet sending interval are jointly adjusted, optimizing the data transmission mechanism, ensuring efficient and reliable communication between training nodes, and being suitable for decentralized deployment. It avoids modifying network devices such as switches, adapts to heterogeneous environments, and can flexibly expand the network capacity and the number of nodes when the cluster is expanded. It solves the problem that the congestion control algorithm of the traditional data center cannot meet the congestion control requirements in the artificial intelligence cluster environment, can perform accurate congestion control on the artificial intelligence training cluster, ensure efficient and reliable communication of training nodes, and has good scalability.
[0080] Embodiment III
[0081] Figure 3 It is a flowchart of a congestion control method for an artificial intelligence training cluster provided by Embodiment III of the present invention. This method can be executed by the data receiving end. As Figure 3 shown, this method includes:
[0082] Step 310: Calculate the reference rate and the one-way queue delay matched by each data packet according to the packet headers of the data packets sent by at least one data sending end, and add the reference rate and the one-way queue delay matched by each data packet to each confirmation data packet.
[0083] In the embodiment of the present invention, the data receiving end can receive data packets sent by multiple data sending ends, parse the packet headers for the data packets sent by each data sending end, determine the category of the collective communication method of the data sending end set, and then obtain the reference rate matched by the data packet based on the mapping table between the collective communication method and the reference rate, and calculate the one-way queue delay. Thus, the reference rate and the one-way queue delay matched by each data packet are added to each confirmation data packet.
[0084] Step 320: Feed back the confirmation data packet to the corresponding data sending end.
[0085] In an embodiment of the present invention, after the data receiving end feeds back the acknowledgment packet to the corresponding data sending end, the data sending end may execute the congestion control method of the artificial intelligence training cluster in the foregoing embodiment.
[0086] In a specific example, the calculated reference rate and the one-way queue delay are combined as the congestion control mechanism. The calculation of the reference rate takes into account the actual carrying capacity of the network link and the current load condition, while the measurement of the one-way queue delay provides a direct feedback on the degree of network congestion. The combination of these two key indicators can accurately judge the network state, and through the dynamic adjustment mechanism, it can effectively cope with the sudden communication requirements during the training process. This flexible adjustment ability is particularly important for supporting collective communication operations under different parallel strategies. Different from the diverse and unpredictable traffic patterns in traditional data center networks, the artificial intelligence training workload mainly relies on specific distributed machine learning frameworks and standardized communication libraries, and this relatively fixed communication mode provides a theoretical basis for accurately calculating the reference rate.
[0087] By deeply integrating with the central controller of the cluster, each worker node can accurately grasp the specific situation of the incoming traffic, including the number of flows, types, and the collective communication methods used. Common collective communication methods include all-reduce, all-to-all, all-gather, and other. The reference rate corresponding to all-reduce can be linerate / n, the reference rate corresponding to all-to-all can be line rate / (n - 1), the reference rate corresponding to all-gather can be line rate / (n - 1), and the reference rate corresponding to other can be line rate / n. Here, line rate represents the line rate, and n represents the number of traffic flows, that is, the number of system nodes.
[0088] The data receiving end dynamically calculates the reference rate according to the current network state. This calculation process takes into account the physical bandwidth of the link and makes dynamic adjustments according to the number and type of flows. For example, when it is detected that the number of flows increases, the reference rate of each flow will be correspondingly reduced to ensure the fair allocation of network resources. The reference rate will be encapsulated in the ACK packet and returned to the data sending end.
[0089] Measuring the one-way queue delay relies on a high-precision time synchronization mechanism, which can achieve a time synchronization accuracy of 100 nanoseconds within the cluster without relying on dedicated hardware. This high-precision time synchronization supports the accurate measurement of the one-way queue delay.
[0090] Exemplarily, such as Figure 4As shown, the forward path delay on the forward transmission path includes the link static delay and the one-way queue delay. The link static delay represents the theoretical time required for each data packet to reach the destination without queuing, including static delays such as propagation delay and serialization delay on network cards and switches. By working in coordination with the cluster scheduler, the value of the link static delay can be updated in a timely manner to ensure the measurement accuracy is maintained during network reconstruction or routing changes.
[0091] When the data sending end sends out a data packet from the network card, it marks the initial timestamp t0 and adds it to the data packet. The data receiving end records the timestamp t1 when the data packet arrives. The end-to-end timestamp mechanism avoids the processing overhead of intermediate devices, ensuring that the measurement result accurately reflects the actual delay situation on the network path, and the forward path delay is obtained as t0 - t1. By subtracting the link static delay from this value, the one-way queue delay of the data packet in the network can be accurately obtained, that is, (t0 - t1) - link static delay. The one-way queue delay directly reflects the congestion state of the network and accurately reflects the actual queuing situation of switches on the network path.
[0092] Compared with the traditional measurement method of data packet round-trip delay, the one-way queue delay has unique advantages as a congestion signal. The measurement of the data packet round-trip delay is easily interfered by factors such as reverse path congestion and data receiving end processing delay, resulting in wrong judgments by congestion control algorithms. However, the one-way queue delay focuses on the forward path of data transmission, effectively avoiding these interference factors, providing more accurate network status information for the data sending end, and increasing the granularity of the control algorithm. In addition, this solution does not require special configuration or programming of switches, showing good adaptability and avoiding the complexity during deployment. As a multi-bit signal, the one-way queue delay can accurately quantify the degree of network congestion, rather than just reflecting whether there is congestion in the network.
[0093] The data receiving end encapsulates the measured one-way queue delay and the reference rate in the ACK packet and returns it to the data sending end. The one-way queue delay is used to quickly detect the network congestion state, and the reference rate provides a reference for the ideal transmission rate. The combination of these two key data can effectively prevent network congestion while maintaining high throughput in the system. Especially when sudden network congestion occurs, the transmission behavior can be adjusted in a timely manner relying on the fast response characteristics of the one-way queue delay.
[0094] At the data sending end, a reference rate record table can be maintained to track whether the reference rate of each active flow changes. When an ACK packet carrying the reference rate is received, the data sending end can compare it with the historical value in the record table. If a difference is detected, it means that the network topology or traffic pattern has changed, such as the addition of new traffic or the end of a traffic transmission. The data sending end will immediately adjust its congestion window size to the product of the new reference rate and the round-trip time to adapt to the new network conditions and update the reference rate record table. This fast response mechanism enables a rapid reach to a new equilibrium point when the network state changes.
[0095] During network stability, the reference rate embedded in the ACK packet serves as the reference rate for the data sender, based on which adjustments are made. When bursty traffic enters the network as Figure 5 shown, the data receiving end detects the fluctuations of this traffic, quickly modifies the reference rate and adds it to the ACK packet to be returned to the data sending end. After detecting the network change, the data sending end quickly updates the sending rate and the congestion window size.
[0096] When the network changes, the adaptive base rate calculation mechanism usually only needs one iteration to converge to a stable value to effectively manage congestion. In contrast, traditional congestion control algorithms may require multiple iterations to adapt to network changes, which may result in a large amount of link bandwidth being wasted in artificial intelligence training scenarios, or the packet delay becoming larger, causing serious performance degradation. And this solution demonstrates superior response capabilities. This fast convergence ability significantly reduces the flow completion time and tail latency, providing a stable and efficient network environment for artificial intelligence training clusters. Especially in large-scale distributed training scenarios, this feature is of great significance for maintaining the continuity of training and improving the overall training efficiency.
[0097] The congestion control method for the artificial intelligence training cluster provided by this solution fully considers the particularity of the artificial intelligence training workload and realizes the efficient utilization of network resources. Through precise flow control and dynamic adjustment mechanisms, excellent network performance is achieved without modifying the existing network infrastructure: 1) In terms of performance, it solves the network requirements during the artificial intelligence training process. By optimizing the data transmission mechanism, it ensures efficient communication between training nodes. Especially when dealing with complex parallel strategies such as tensor parallelism and mixture of experts parallelism, it provides low-latency and high-bandwidth transmission performance, improves the convergence speed and resource utilization efficiency of training, reduces the communication overhead during training, thereby minimizing the flow completion time and enhancing the overall training efficiency. 2) It provides good scalability. Adopting a decentralized deployment scheme, it only needs to be deployed at the network card end, avoiding the modification of network devices such as switches. This design not only simplifies the deployment process but also ensures adaptability in heterogeneous environments. It can flexibly expand the network capacity and the number of nodes when the cluster is expanded. At the same time, the lightweight design ensures stable performance when the cluster scale is expanded. 3) It provides high reliability, especially considering the continuous traffic characteristics of artificial intelligence training. In large-scale distributed training, network interruptions or performance fluctuations may lead to training failures or a decrease in the convergence speed. Therefore, a multi-level reliability guarantee mechanism is introduced, including precise flow control, dynamic congestion detection, and fast fault recovery, to ensure the stability of the training process.
[0098] Embodiment 4
[0099] Figure 6 It is a schematic structural diagram of a congestion control device for an artificial intelligence training cluster provided by Embodiment 4 of the present invention. The congestion control device of the artificial intelligence training cluster is configured at the data sending end. As Figure 6 shown, the device includes:
[0100] The acknowledgment packet parsing module 410 is used to parse the acknowledgment packet fed back by the data receiving end;
[0101] The first transmission rate determination module 420 is used to, when the reference rate carried in the acknowledgment packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjust the packet transmission rate in stages according to the one-way queue delay and the target queuing time to obtain the target transmission rate;
[0102] The congestion control module 430 is used to determine the number of congestion windows according to the target transmission rate and the round-trip time, and determine the packet sending interval according to the number of congestion windows.
[0103] In the technical solution of the embodiment of the present invention, by parsing the acknowledgment data packet fed back by the data receiving end, when the reference rate carried in the acknowledgment data packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, the packet sending rate is adjusted in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate. Then, the congestion window number is determined according to the target sending rate and the round-trip time, and the packet sending interval is determined according to the congestion window number. In this solution, when the reference rate remains unchanged, the target sending rate of the packet is dynamically adjusted through the one-way queue delay and the target queuing time, and the congestion window number and the packet sending interval are jointly adjusted, optimizing the data transmission mechanism, ensuring efficient and reliable communication between training nodes, being applicable to decentralized deployment, avoiding modification of network devices such as switches, adapting to heterogeneous environments, being able to flexibly expand the network capacity and the number of nodes during cluster expansion, solving the problem that the congestion control algorithm of the traditional data center cannot meet the congestion control requirements in the artificial intelligence cluster environment, being able to perform precise congestion control on the artificial intelligence training cluster, ensuring efficient and reliable communication of training nodes, and having good scalability.
[0104] Optionally, the congestion control device of the artificial intelligence training cluster includes a second sending rate determination module, configured to, when the reference rate carried in the acknowledgment data packet changes, adjust the packet sending rate based on the current reference rate and record the adjustment time of the packet sending rate.
[0105] Optionally, the first sending rate determination module 420 includes an acceleration determination unit and a deceleration determination unit. The acceleration determination unit is configured to, when the one-way queue delay is less than the target queuing time, adjust the packet sending rate for acceleration based on the acceleration parameter to obtain the target sending rate. The deceleration determination unit is configured to, when the one-way queue delay is greater than or equal to the target queuing time, adjust the packet sending rate for deceleration based on the deceleration coefficient to obtain the target sending rate.
[0106] Optionally, the acceleration determination unit is specifically configured to, when the one-way queue delay is zero, adjust the packet sending rate for acceleration based on the first acceleration parameter in the acceleration parameter to obtain the target sending rate; when the one-way queue delay is greater than zero and less than the target queuing time, adjust the packet sending rate for acceleration based on the second acceleration parameter in the acceleration parameter to obtain the target sending rate.
[0107] Optionally, the congestion control device of the artificial intelligence training cluster further includes a deceleration coefficient determination unit, configured to determine an exponential reduction parameter and a multiplicative deceleration factor; calculate a first difference between the one-way queue delay and the target queuing time, and calculate a first sum of the one-way queue delay and the transmission link static delay; multiply the ratio of the first difference to the first sum by the exponential reduction parameter to obtain a multiplicative deceleration factor to be compared; determine the deceleration coefficient according to the multiplicative deceleration factor and the multiplicative deceleration factor to be compared.
[0108] Optionally, the congestion control module 430 is specifically configured to, when the number of congestion windows is less than a preset number of windows, determine the data packet sending interval based on the ratio of the round-trip time to the number of congestion windows; when the number of congestion windows is greater than or equal to the preset number of windows, the data packet sending interval is zero.
[0109] Optionally, the congestion control device of the artificial intelligence training cluster further includes a timeout mechanism processing module, configured to, after detecting an acknowledgment packet loss timeout event, update the number of congestion windows according to the product of the multiplicative deceleration factor and the number of congestion windows, and configure the target sending rate to zero.
[0110] The congestion control device of the artificial intelligence training cluster configured at the data sending end provided by the embodiments of the present invention can execute the congestion control method of the artificial intelligence training cluster executed by the data sending end provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0111] Figure 7 FIG. 4 is a schematic structural diagram of another congestion control device of the artificial intelligence training cluster provided by Embodiment 4 of the present invention. The congestion control device of the artificial intelligence training cluster is configured at the data receiving end. As Figure 7 shown, the device includes:
[0112] An acknowledgment packet generation module 510, configured to calculate a reference rate and a one-way queue delay matching each data packet according to the packet headers of the data packets sent by at least one data sending end, and add the reference rate and the one-way queue delay matching each data packet to each acknowledgment packet.
[0113] A data sending module 520, configured to feed back the acknowledgment packet to the corresponding data sending end, so that the data sending end executes the corresponding congestion control method of the artificial intelligence training cluster.
[0114] The congestion control device of the artificial intelligence training cluster configured at the data receiving end provided by the embodiments of the present invention can execute the congestion control method of the artificial intelligence training cluster executed by the data receiving end provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0115] Embodiment 5
[0116] An embodiment of the present invention provides an artificial intelligence chip. The processor or integrated circuit module in the artificial intelligence chip can be used to execute the congestion control method of the artificial intelligence training cluster executed by the data sending end in any embodiment of the present invention, or execute the congestion control method of the artificial intelligence training cluster executed by the data receiving end.
[0117] In some embodiments, the congestion control method of the artificial intelligence training cluster executed by the data receiving end, or the congestion control method of the artificial intelligence training cluster executed by the data sending end, can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium. When the computer program is loaded into the artificial intelligence chip for execution, one or more steps of the congestion control method of the artificial intelligence training cluster executed by the data receiving end, or the congestion control method of the artificial intelligence training cluster executed by the data sending end described above can be executed.
[0118] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages.
[0119] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0120] This application embodiment also discloses a computer program product. The computer program product includes a computer program that, when executed, implements the congestion control method of the artificial intelligence training cluster executed by the data receiving end provided in any embodiment of this application, or the congestion control method of the artificial intelligence training cluster executed by the data sending end. This program product and the congestion control method of the artificial intelligence training cluster executed by the data receiving end or the data sending end disclosed in each embodiment of this application belong to the same inventive concept, and thus will not be elaborated herein.
[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0122] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A congestion control method for an artificial intelligence training cluster, characterized in that, Executed by the data sender, including: Parsing the acknowledgement data packet fed back by the data receiver; When the reference rate carried in the acknowledgement data packet remains unchanged and the current reference rate adjustment interval is not less than the transmission rate adjustment interval threshold, adjusting the data packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate; Determining the number of congestion windows according to the target sending rate and the round-trip time, and determining the data packet sending interval according to the number of congestion windows.
2. The congestion control method for an artificial intelligence training cluster according to claim 1, wherein Also including: When the reference rate carried in the acknowledgement data packet changes, adjusting the data packet sending rate based on the current reference rate and recording the adjustment time of the data packet sending rate.
3. The congestion control method for the artificial intelligence training cluster according to claim 1, characterized in that, Adjusting the data packet sending rate in stages according to the one-way queue delay and the target queuing time to obtain the target sending rate, including: When the one-way queue delay is less than the target queuing time, adjusting the data packet sending rate to increase based on the increasing speed parameter to obtain the target sending rate; When the one-way queue delay is greater than or equal to the target queuing time, adjusting the data packet sending rate to decrease based on the decreasing speed coefficient to obtain the target sending rate.
4. The congestion control method for an artificial intelligence training cluster according to claim 3, wherein When the one-way queue delay is less than the target queuing time, adjusting the data packet sending rate to increase based on the increasing speed parameter to obtain the target sending rate, including: When the one-way queue delay is zero, adjusting the data packet sending rate to increase based on the first increasing speed parameter in the increasing speed parameter to obtain the target sending rate; When the one-way queue delay is greater than zero and less than the target queuing time, adjusting the data packet sending rate to increase based on the second increasing speed parameter in the increasing speed parameter to obtain the target sending rate.
5. The congestion control method of the artificial intelligence training cluster according to claim 3, wherein Before adjusting the data packet sending rate to decrease based on the decreasing speed coefficient to obtain the target sending rate, further including: Determining the exponential decrease parameter and the multiplicative decrease factor; Calculating the first difference between the one-way queue delay and the target queuing time, and calculating the first sum value of the one-way queue delay and the static delay of the transmission link; Multiplying the ratio of the first difference to the first sum value by the exponential decrease parameter to obtain the multiplicative decrease factor to be compared; Determining the decreasing speed coefficient according to the multiplicative decrease factor and the multiplicative decrease factor to be compared.
6. The congestion control method for the artificial intelligence training cluster according to claim 1, characterized in that Determining the data packet sending interval according to the number of congestion windows, including: When the number of congestion windows is less than the preset window number, determining the data packet sending interval based on the ratio of the round-trip time to the number of congestion windows; When the number of congestion windows is greater than or equal to the preset window number, the data packet sending interval is zero.
7. The congestion control method for an artificial intelligence training cluster according to any one of claims 1-6, characterized in that Also including: After detecting the acknowledgement data packet loss timeout event, updating the number of congestion windows according to the product of the multiplicative decrease factor and the number of congestion windows, and configuring the target sending rate to zero.
8. A congestion control method for an artificial intelligence training cluster, characterized in that, Executed by the data receiver, including: Calculate the reference rate and one-way queue delay matching each data packet according to the packet headers of the data packets sent by at least one data sender, and add the reference rate and the one-way queue delay matching each data packet to each acknowledgment packet; Feed back the acknowledgment packets to the corresponding data senders, so that the data senders execute the method according to any one of claims 1-7.
9. An artificial intelligence chip, characterized in that, The artificial intelligence chip is used to execute the congestion control method of the artificial intelligence training cluster performed by the data sender according to any one of claims 1-7, or execute the congestion control method of the artificial intelligence training cluster performed by the data receiver according to claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to execute and implement the congestion control method of the artificial intelligence training cluster performed by the data sender according to any one of claims 1-7, or implement the congestion control method of the artificial intelligence training cluster performed by the data receiver according to claim 8.
11. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is used to execute and implement the congestion control method of the artificial intelligence training cluster performed by the data sender according to any one of claims 1-7, or implement the congestion control method of the artificial intelligence training cluster performed by the data receiver according to claim 8.