Congestion control method for intra-frame network aggregation
By introducing a congestion control method for rack intranet aggregation in distributed computing, using ECN tags and congestion notification packet mechanisms, the problem of traditional congestion control methods being too slow to respond is solved, and more efficient network resource utilization and data aggregation efficiency is achieved.
Patent Information
- Application Number
- CN202510448662.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-13
AI Technical Summary
In large-scale distributed computing, a large number of ensemble communication operations in parallel data training may lead to network congestion, resulting in increased latency, decreased throughput and unstable communication performance. The traditional congestion control method responds too slowly and cannot follow the sudden large traffic changes in AI training, and the abnormal working nodes lead to inefficient aggregation.
A congestion control method for rack intranet aggregation is proposed. By starting each round of data aggregation, each working node conducts additive exploration at a set rate until a congestion notification message is received. The switching unit forms a dynamic queue, labels the data packets with ECN tags according to the queue length, and broadcasts congestion notification messages. The worker node adjusts the rate according to the notification to alleviate congestion.
Effectively reduce latency, improve bandwidth utilization, improve on-network aggregation efficiency, can quickly respond to changes in training traffic, and ensure synchronous training of each working node.
Smart Images

Figure CN120151285A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence control technology, and particularly to a congestion control method for in-rack in-network aggregation. Background Art
[0002] With the continuous development of deep learning, the scale of neural networks is becoming increasingly large. When facing the training of more and more ultra-large-scale neural networks, distributed training methods are usually required, which requires multiple machines to work together. Currently, data parallelism is a common distributed training method. In data parallelism, collective communication operations play a crucial role in parallel computing. These collective communication operations are usually used to coordinate and integrate data from multiple computing units to obtain global results. A common collective communication operation is the "AllReduce" upward communication, which plays a key role in distributed computing. The purpose of this communication operation is to aggregate the local calculation results in each computing unit into a global result. Upward communication refers to sending the local calculation results from each computing unit to the central node of the collective operation for global data aggregation. In deep learning, this usually involves transmitting local gradients from each computing unit to the collective node for global gradient calculation and model parameter update. Another communication operation is the "AllReduce" downward communication, which is also a common operation in distributed computing. It is usually used to propagate the global result obtained from the collective operation back to each computing unit. This is the second stage of the "AllReduce" operation, aiming to distribute the globally aggregated result to each computing unit so that they can update their local states or model parameters. This downward communication operation ensures the synchronous update of the global result, enabling all computing units to obtain the latest model parameters. In summary, as the scale of neural networks increases, distributed training becomes a necessary choice. Data parallelism is a common distributed training method, in which collective communication operations such as "AllReduce" upward communication and downward communication play important roles in distributed computing, used to coordinate and integrate data to obtain global results and keep the model parameters of each computing unit synchronized.
[0003] SwitchML is a commonly used system for distributed deep learning training, aiming to optimize data transmission and communication to improve training performance and efficiency. It is typically used to accelerate large-scale data transmission in machine learning workloads. For example, during the training process of deep learning models, it aims to maximize the bandwidth and efficiency of data transmission by effectively utilizing high-speed network switches (such as RDMA-enabled Ethernet switches), thereby accelerating data transmission. By helping to reduce the time required for data transmission, it accelerates the model training process. The architecture diagram of the SwitchML framework is usually relatively simple because its core idea is to accelerate distributed deep learning training through high-performance network switches and reduce the communication overhead between computing units. In this architecture, there is usually no dedicated switch, but the switch itself plays an important role and performs data aggregation operations. SwitchML is used in the scenario of data parallel distributed training, where multiple computing units are connected to a programmable switch. The switch aggregates the sending data of multiple computing units and communicates with the computing units.
[0004] Congestion control is an important issue in large-scale distributed computing. Due to a large number of collective communication operations involved in data parallel training, when data is transmitted in the network, network congestion may occur. Network congestion leads to increased latency, decreased throughput, and instability of communication performance. The root cause of the congestion control problem is the finiteness of network resources and the imbalance of data transmission. When multiple computing units simultaneously send a large amount of data to the collective node, the network link may be overloaded, resulting in congestion. This leads to packet loss, retransmission, and increased latency, thus affecting the overall training performance.
[0005] To solve this problem, the SwitchML system usually adopts a series of congestion control mechanisms. These include techniques such as flow control, congestion detection, and congestion avoidance. Flow control can limit the sending rate to prevent too much data from flowing into the network. Congestion detection can timely detect and respond to congestion situations by monitoring congestion flags in the network. Once congestion is detected, the congestion avoidance mechanism will take corresponding measures, such as reducing the transmission rate or adjusting the data distribution strategy, to relieve congestion and restore network performance. Traditional congestion control algorithms have the following problems: (1) During AI training, the traffic has strong time-varying characteristics and is prone to sudden extremely large traffic. Traditional congestion control methods respond too slowly and cannot well follow the dynamic changes of training traffic characteristics; (2) The training speeds between worker nodes are not synchronized, and the transmission progress of each node is also different. Nodes that complete transmission in advance need to wait for slower nodes to perform the next round of aggregation, resulting in low aggregation efficiency. Summary of the Invention Based on the technical problems existing in the background art, the present invention proposes a congestion control method for in-rack intra-network aggregation, which can effectively reduce latency and improve bandwidth utilization, thereby enhancing the efficiency of intra-network aggregation.
[0006] A congestion control method for in-rack intra-network aggregation proposed by the present invention includes the following steps: Step 1: At the beginning of each round of data aggregation, each worker node performs additive exploration at a set rate, and the rate increment value for each additive exploration is a fixed value , until a congestion notification packet forwarded by the switching unit is received; Step 2: The switching unit obtains the data packets uploaded by each worker node and forms a dynamic queue in order. When the length of the dynamic queue is within a set range, with a probability add an ECN label to the data packet, and send the data packets uploaded by each worker node to the computing unit; Step 3: When the computing unit receives a data packet with an ECN label and does not receive other data packets with an ECN label within a time interval , update the maintained bitmap table, carry the updated bitmap table into the payload of the congestion notification packet, and broadcast it to all worker nodes through the switching unit; Step 4: The worker node adjusts the rate according to the updated bitmap table carried in the congestion notification packet to relieve line congestion.
[0007] Further, in Step 1, the set rate is calculated as follows: ; where is the set number of worker nodes, is the total bandwidth received by all worker nodes from the computing unit.
[0008] Further, in Step 2, the switching unit sets a lower limit and an upper limit for the queue length. The switching unit obtains the data packets uploaded by each worker node and forms a dynamic queue in order. When the length of the dynamic queue is greater than the lower limit , and less than or equal to the upper limit , add an ECN label to the data packet with a probability .
[0009] Further, adding an ECN label to the data packet with a probability specifically means: modifying the ECN field in the IP data packet header from 0x00 to 0x11, and the probability increases as the length of the dynamic queue increases Further, in step two, the uploaded data packets include the data packets with ECN tags attached.
[0010] Further, in step three, the calculation unit maintains an original bitmap table for recording the data to be aggregated in each round. By default, all entries in the original bitmap table are 0. After receiving the data packets forwarded by the switching unit, the entries in the original bitmap table corresponding to the uploaded data packets are modified to 1.
[0011] Further, in step four, specifically: the worker nodes adjust the rate according to the updated bitmap table carried in the congestion notification message. Each worker node uses the number of unfinished data packets of the current node among the unfinished data packets of all worker nodes as the weight to allocate the total bandwidth to each worker node after weighting to relieve line congestion.
[0012] Further, in step four, each worker node adjusts the rate according to the following formula: ; where, represents the rate adjustment value of the th worker node, is the total number of data packets aggregated in each round, represents a constant, represents the th data packet of the th worker node in the updated bitmap table, represents the th worker node, represents the th data packet of the th worker node in the updated bitmap table, represents the number of data packets that all worker nodes have completed, represents the number of data packets that all worker nodes need to process, represents the product, represents the number of unfinished data packets of all worker nodes.
[0013] A computer-readable storage medium stores a number of classification programs thereon. The number of classification programs is used to be called by a processor and execute the congestion control method as described above.
[0014] The advantages of a congestion control method for in-rack in-network aggregation provided by the present invention are as follows: In the structure of the present invention, a congestion control method for in-rack in-network aggregation is provided. Based on the ECN mechanism, the switching unit marks network data packets, so that the computing unit can quickly perceive the current network state, make a quick and timely response, and notify other working nodes to perform congestion control. The training speeds of the working nodes are not synchronized. Therefore, the working node with a fast transmission progress slows down its transmission speed to wait for the working node with a slow training speed, and distributes this extra bandwidth resource to the working node with a slow transmission speed to enable it to increase its transmission speed and catch up with the working node with a fast progress; By reasonably allocating bandwidth resources, the aggregation speed of data can be improved; Therefore, this congestion control method can effectively reduce latency and improve the utilization rate of bandwidth, thereby improving the in-network aggregation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic structural diagram of the present invention; Figure 2 is a schematic diagram of the congestion control method on the FPGA state machine; Figure 3 is a schematic diagram of data exchange among working nodes, switching units, and computing units in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0017] As Figures 1 to 3 shown, a congestion control method for in-rack in-network aggregation proposed by the present invention includes the following steps: Step 1: Set the number of working nodes as , the total bandwidth received by all working nodes from the computing unit is . At the beginning of each round of data aggregation, each working node performs additive probing at a rate of , and the rate increment value for each additive probing is a fixed value until a congestion notification message forwarded by the switching unit is received; Step 2: The switching unit sets a lower limit and an upper limit for the queue length. The switching unit obtains the data packets uploaded by each working node and forms a dynamic queue in order. When the length of the dynamic queue is greater than the lower limit , with a probability of Tag the data packet with an ECN label, that is, modify the ECN field in the IP data packet header from 0x00 to 0x11, and send the data packets uploaded by each worker node to the computing unit. The uploaded data packets include the data packets tagged with the ECN label; According to the probability In tagging the data packet with an ECN label, the data packets forming the dynamic queue exceed the lower limit , that is, it causes a certain degree of data congestion. In the subsequent processing, the rate of the worker node needs to be adjusted. The probability will increase as the queue length increases. When the length of the dynamic queue is greater than the upper limit , all are directly tagged with the ECN label, that is, the probability p = 100%.
[0018] Step 3: The computing unit maintains an original bitmap table for recording the data to be aggregated in each round. By default, all entries in the original bitmap table are 0. After receiving the data packets forwarded by the switching unit, modify the entry in the original bitmap table corresponding to the uploaded data packet to 1. When receiving a data packet with an ECN label and not receiving other data packets with an ECN label within the time interval , obtain the updated bitmap table. The computing unit generates a congestion notification message, and carries the updated bitmap table in the payload of the congestion notification message, and broadcasts it to all worker nodes through the switching unit; The computing unit obtains the th data packet of the th worker node from the switching unit. In the original bitmap table, is modified to 1. That is, the computing unit marks the obtained data packet. This marking method is to modify the entry corresponding to the data packet in the original bitmap table from 0 to 1, so as to obtain the updated bitmap table. The entries marked as 1 in the updated bitmap table indicate the data packets that have been processed and are being processed, and the entries marked as 0 in the updated bitmap table indicate the data packets that have not been processed yet.
[0019] Step 4: The worker node adjusts its rate according to the updated bitmap table carried in the congestion notification message. Each worker node uses the number of uncompleted data packets of the current node as the weight among the uncompleted data packets of all worker nodes, and distributes the total bandwidth of after weighting to each worker node to relieve line congestion.
[0020] Therefore, all worker nodes adjust their rates together. Each worker node adjusts its rate according to the following formula: ; Among them, Indicates the rate adjustment value of the th working node, is the total number of data packets aggregated in each round, represents a constant, represents the th data packet of the th working node in the updated bitmap table, Indicates the number of uncompleted data packets of the th working node, represents the th data packet of the th working node in the updated bitmap table, represents the number of data packets that all working nodes have completed, represents the number of data packets that all working nodes need to process, represents the product,
[0021] It should be noted that represents a constant. Since congestion has occurred, a constant value needs to be subtracted to ensure that the data packets in the egress queue buffer are emptied.
[0022] Next, each node loops through the above steps to send the remaining data packets.
[0023] As the training tasks of ultra-large-scale neural networks continue to increase, adopting a distributed training method has become a common choice. The training process needs to ensure the stability and reliability of the network to improve the performance and efficiency of distributed deep learning training. Reasonable congestion control can prevent congestion from occurring in the network, that is, when the network link or node is overloaded, resulting in packet loss, retransmission, and increased latency. Therefore, this embodiment proposes a congestion control algorithm with centralized management, enabling the system to avoid performance degradation and training instability caused by congestion and ensuring the timely transmission and processing of data.
[0024] Through Steps 1 to 4, based on the ECN mechanism, the switching unit marks the network data packets, enabling the computing unit to quickly perceive the current network state, respond quickly and in a timely manner, and notify other working nodes to perform congestion control. The training speeds of the working nodes are not synchronized. Therefore, the working node with a fast transmission progress slows down its transmission speed to wait for the working node with a slow training speed, and distributes this extra bandwidth resource to the working node with a slow transmission speed to enable it to increase its transmission speed and catch up with the working node with a fast progress. By reasonably allocating bandwidth resources, the aggregation speed of data can be improved.
[0025] As an embodiment This solution is a centralized notification congestion control architecture based on an FPGA aggregator, with a computing unit (FPGA) as the central management node. It is connected to multiple worker nodes simultaneously through a switching unit. During the training process of a deep learning model, to support the dynamic changes in the training progress and synchronization on different nodes in the training system, it manages and coordinates the gradient aggregation rate among various workers, adjusts the transmission rate on each node in real time, and achieves the goal of a shorter total training time. To test the feasibility of the congestion control method in this embodiment, this solution is deployed in a real environment, such as Figure 2 As shown, four hosts and one programmable switch are used. One host simulates the computing unit for aggregating data and maintaining the bitmap table, and three hosts are used to simulate the worker nodes for sending training data (corresponding to Worker1, Worker2, and Worker3 in the attached figure). The programmable switch simulates the switching unit, which is used to transmit training data and implement the congestion control algorithm of this embodiment. In this example, through simulating distributed training of three worker nodes, the programmable switch is used for data transmission and congestion control, and data aggregation is performed in the simulated computing unit. By simulating data transmission during distributed training and measuring the time required to complete the transmission multiple times, it is compared with the traditional congestion control algorithm to verify the feasibility of this algorithm.
[0026] In this example, it is assumed that the bandwidth of the computing unit is 300 MBps. The servers of the three simulated worker nodes are respectively connected to three ports of the switching unit (programmable switch), and the number of data packets sent by each node is 1000. The server of one simulated computing unit is connected to one port of the switching unit. In the initial stage, each worker node sends data packets to the port of the switching unit at a speed of 100 MBps by running a script. At the same time, each worker node also needs to run a listening script to listen for the congestion notification packet (CNP) signal from the computing unit. The switching unit is responsible for forwarding the data packets of the worker nodes to the computing unit, and at the same time needs to detect the congestion situation in the switching unit and mark the data packets with the ECN label according to the congestion situation with a probability to mark the data packets with the ECN label.
[0027] For example, when the number of data packets in the dynamic queue of the switching unit is greater than 50% (lower limit ) of the queue length and less than 60% (upper limit ) of the queue length, the switch will mark the data packets with the ECN label with a probability of 1 / 2, that is, mark half of the data packets in the dynamic queue with the ECN label. The computing unit maintains a two-dimensional array bitmap[3]
[1000] , and the default entries in the original bitmap table are all 0. Whenever it receives the th data packet from the When there are packets, modify the
[0028] in the original bitmap table to 1 to obtain an updated bitmap table. The computing unit will monitor the ports of the connection switching unit, detect the ECN label of each packet. If it finds that the ECN label is 0x11 (corresponding to the ECN label being added to the packet), it will send a congestion notification message (congestion control signal CNP) to all worker nodes through the switching unit. This congestion notification message will carry the updated bitmap table. At the same time, the computing unit will also aggregate the received packets, that is, perform calculations on the received data. When a worker node receives a congestion notification message from the computing unit, it will adjust the rate according to the information in the updated bitmap table to relieve the congestion in the line. Each worker node will adjust the rate to: ; Continue to loop through the previous steps to send the remaining packets.
[0029] As described above, only the preferred specific implementation manner of the present invention is provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A congestion control method for intra-rack network aggregation, characterized in that: The steps include: Step 1: At the beginning of each round of data aggregation, each working node performs additive exploration at a set rate, and the rate increase of each additive exploration is a fixed value. , until receiving a congestion notification message forwarded by the switching unit; Step 2: The switching unit obtains the data packets uploaded by each working node and forms a dynamic queue in sequence. When the length of the dynamic queue is within the set range, the probability Add ECN labels to the data packets and send the data packets uploaded by each working node to the computing unit; Step 3: The computing unit receives a data packet with an ECN label and When no other data packets with ECN labels are received within the time limit, the maintained bitmap table is updated, the updated bitmap table is carried into the payload of the congestion notification message, and is broadcasted to all working nodes through the switching unit; Step 4: The working node adjusts the rate according to the updated bitmap table carried in the congestion notification message to alleviate line congestion.
2. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step 1, set the rate The calculation is as follows: ; in, To set the number of working nodes, It is the total bandwidth of computing units received by all working nodes.
3. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step 2, the switching unit sets a lower limit on the queue length and upper limit The switching unit obtains the data packets uploaded by each working node and forms a dynamic queue in sequence. When the length of the dynamic queue is greater than the lower limit , and less than or equal to the upper limit When, according to the probability Add an ECN label to the data packet.
4. The congestion control method for intra-rack network aggregation according to claim 3, characterized in that: According to probability The specific steps of labeling the data packet with ECN tag are as follows: modify the ECN field in the IP data packet header from 0x00 to 0x11, the probability Increases as the length of the dynamic queue increases.
5. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step 2, the uploaded data packets include data packets marked with ECN labels.
6. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step three, the computing unit maintains an original bitmap table for recording the data that needs to be aggregated in each round.
7. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: The original bitmap table defaults to all 0s. After receiving the data packet forwarded by the switching unit, the entry corresponding to the uploaded data packet in the original bitmap table is modified to 1.
8. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step 4, specifically: the working node adjusts the rate according to the updated bitmap table carried in the congestion notification message. Each working node uses the current number of unfinished data packets in the number of unfinished data packets of all working nodes as the weight, and adjusts the total bandwidth to Each working node is assigned after weighting to alleviate line congestion.
9. The congestion control method for intra-rack network aggregation according to claim 1, characterized in that: In step 4, each working node adjusts the rate according to the following formula: in, Indicates The rate adjustment value of the worker nodes, is the total number of packets aggregated in each round, represents a constant, Indicates updating the bitmap table. The first worker node Data packets, Indicates The number of unfinished packets per worker node, Indicates updating the bitmap table. The first worker node Data packets, Indicates the number of data packets that have been completed by all working nodes. Indicates the number of data packets that all working nodes need to process, represents the product, Indicates the number of outstanding packets for all worker nodes.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of classification programs, and the plurality of classification programs are used to be called by a processor and execute the congestion control method as claimed in any one of claims 1 to 9.