Efficient data center network congestion control method and system for realizing O (1) convergence
By using batch processors to calculate network status evaluation indicators and dynamically adjusting the transmission rate and window size, the bottleneck of the existing data center network congestion control technology in a high-speed environment is solved, and efficient, fair and stable network transmission is achieved.
Patent Information
- Application Number
- CN202510423320.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-13
AI Technical Summary
The existing data center network congestion control technology has bottlenecks in convergence speed and computing efficiency in high-speed environments, resulting in insufficient link utilization, unfairness of RTT, and Bufferbloat phenomenon, which is difficult to meet the performance needs of the new generation of data center networks.
The network status is obtained by batch processor, and the evaluation indicators of the network status are calculated in batches, such as average delay, delay gradient, average data volume and transmission rate, and dynamically adjust the transmission rate and window size to achieve the O(1) step convergence control mechanism.
It improves network resource utilization and transmission efficiency, reduces latency and congestion risks, achieves fairness and stability, and is suitable for the efficient congestion control needs of modern data centers.
Smart Images

Figure CN120151283A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of high-speed data center network transmission control, and particularly relates to congestion control technology in high-speed networks (>100Gbps). The present invention breaks through the bottlenecks of the convergence speed and computational efficiency of traditional congestion control schemes and is applicable to the construction of new-generation infrastructures such as distributed machine learning. Background Art
[0002] In the process of data center network evolution, it is an inevitable trend for data center networks to develop towards ultra-large scale. Data center congestion control algorithms play a crucial role in providing high-throughput and low-latency data transmission for critical services. However, in the current context of the rapid evolution of ultra-large scale data center networks, traditional congestion control mechanisms face many challenges. At present, many traditional congestion control technology systems have multi-dimensional systematic bottlenecks, such as distributed AI model training, high I / O speed distributed storage, resource decomposition, etc. To meet these goals, data center congestion control algorithms usually aim at full bandwidth utilization and minimum queuing delay and strive to converge to these goals as quickly as possible. With the rapid development of network bandwidth and popular switch service models, the convergence speed of data center congestion control algorithms has become increasingly critical.
[0003] The convergence speed of data center congestion control algorithms involves many aspects of bottlenecks. There is a significant gap between the rapid growth of network interface bandwidth and the improvement speed of protocol stack processing capabilities. When the network interface is upgraded from a lower bandwidth to a higher bandwidth, the transmission delay is greatly compressed, which makes classical algorithms based on packet loss detection show obvious performance limitations in high-speed environments, and many problems are exposed in the core mechanism in high-bandwidth and low-latency network environments.
[0004] Firstly, the link utilization rate is insufficient. Packet loss events lead to a sharp drop in bandwidth, which in turn causes a significant decrease in link utilization rate. Actual test data shows that in high-bandwidth links, the average throughput of traditional algorithms is much lower than the theoretical value, and the utilization rate cannot reach the expected level. In addition, the RTT unfairness of traditional algorithms has always been a bottleneck problem, and the throughput of long flows may be suppressed by short flows. These bottlenecks will ultimately lead to the deterioration of the Bufferbloat phenomenon, a significant increase in the switch cache queue delay, and further exacerbate network congestion and performance degradation.
[0005] Secondly, the current state-of-the-art latency-sensitive solutions also have technical bottlenecks, which also illustrates the importance of the convergence speed of data center network congestion control algorithms. The current solutions aim to address the deficiencies of traditional packet loss detection mechanisms and attempt to reflect the congestion state through latency signals. However, these solutions still face many technical bottlenecks in practical applications, such as poor hardware offloading adaptability. At the hardware offloading level, existing solutions are limited by storage and computing resources. When dealing with a large number of concurrent flows, the state storage requirements may exceed the actual capacity of the hardware, thus requiring an expansion of the storage space, which will significantly increase the memory access latency. In addition, the serial processing mode of traditional algorithms does not match the parallel architecture of smart network cards, and the parallel capabilities of the hardware cannot be fully utilized, resulting in the actual performance being far lower than the theoretical value.
[0006] Furthermore, the emerging in-band network telemetry technology introduces new constraints. As an emerging technology, it inserts metadata into the packet header to achieve real-time monitoring of network status, but it also brings many problems, such as too much data carried in the packet header, resulting in a loss of effective bandwidth. There is also a relatively high closed-loop latency in transmission control, making it difficult to meet the strict real-time requirements of certain scenarios.
[0007] In summary, the existing congestion control technologies have significant bottlenecks in multiple dimensions, severely restricting the expansion of the performance boundaries of the new generation of data center networks. Some solutions that improve efficiency through precise rate calculation can optimize performance to a certain extent, but their control models may cause delay jitter under dynamic loads; other innovative solutions based on feedback mechanisms require higher-precision clock synchronization and more network state information capture, which will significantly increase the deployment cost. These systematic defects severely restrict the technological evolution of key scenarios (such as distributed machine learning, high-frequency trading, etc.). The trade-off between latency and throughput makes it difficult for network performance to break through the existing ceiling, there is a fundamental contradiction between fairness requirements and convergence speed, and the hybrid scheduling strategy exacerbates the performance degradation of certain traffic types. There is a lack of a unified framework for technical solutions among different business scenarios, and it is impossible to effectively meet the differentiated requirements of scenarios such as AI training, databases, and video streams. These multi-dimensional technical bottlenecks severely restrict the expansion of the performance boundaries of the new generation of data center networks, and there is an urgent need to build a breakthrough fast-converging intelligent congestion control system. Summary of the Invention
[0008] The technical problem to be solved by the present invention is: how to implement an easily deployable congestion control (CC) mechanism that can achieve O(1) step convergence without relying on advanced hardware functions or the latest network devices.
[0009] The present invention mainly addresses the following main problems faced by existing fast-converging congestion control algorithms in actual deployment:
[0010] Data center congestion control algorithm based on precise INT: High-performance hardware needs to be adopted in the network to process complex INT header information, which not only affects the forwarding rate of the switch but also may cause the packet length to exceed the MTU when a link fails, thereby reducing the effective payload ratio of the network;
[0011] Data center congestion control algorithm based on the receiver: It relies on symmetric network topologies and packet spraying techniques. However, frequent link failures make it difficult to maintain symmetric topologies. At the same time, the problem of out-of-order packets requires the support of the latest model network cards, increasing the cost of hardware upgrades;
[0012] Data center congestion control algorithm based on the switch: The switch needs to have computing capabilities, but most commercial switches do not support this function.
[0013] To solve the hardware dependence problem of the above O(1)-step convergence CC in actual deployment, the present invention provides an efficient data center network congestion control method and system that achieve O(1) convergence. Without relying on advanced hardware functions or the latest network devices, a congestion control (CC) mechanism that is easy to deploy and can achieve O(1)-step convergence is realized.
[0014] To achieve the above object, the present invention adopts the following technical solutions:
[0015] In a first aspect, the present invention provides an efficient data center network congestion control method that achieves O(1) convergence, including:
[0016] Using a batch processor to obtain the network state and batch-calculate evaluation metrics of the network state, including average latency, latency gradient, average in-transit data volume, and transmission rate;
[0017] According to the obtained network state, call the network state estimation function to analyze the current latency, bandwidth, and queue situation of the network, and trigger window loop control and rate loop control of the network when the network latency indicates potential underutilization of bandwidth.
[0018] Optionally, the batch calculation of the evaluation metrics of the network state is triggered in the following manner:
[0019] The sender embeds the sending timestamp in the header of the packet;
[0020] The receiver echoes the sending timestamp to the ACK packet;
[0021] The sender calculates the round-trip delay through the difference between the timestamp of the received ACK packet and the sending timestamp;
[0022] Record cumulative values, including the cumulative value of the sending timestamp, the cumulative value of the round-trip delay, the cumulative value of the square of the sending timestamp, the cumulative value of the product of the sending timestamp and the round-trip delay, and the packet count;
[0023] Each time the sender receives an ACK packet, calculate the round-trip delay and update the cumulative values; when the cumulative number of ACK packets reaches the set threshold and the current time exceeds the set window from the batch start time, trigger batch calculation.
[0024] Optionally, the evaluation metrics of the network state are batch-calculated based on the recorded cumulative values. After the calculation is completed, clear the cumulative values and update the batch start time.
[0025] Optionally, the average delay is calculated by averaging the cumulative value of the round-trip delay, and the delay gradient is calculated by the following formula:
[0026]
[0027] In the formula, g represents the delay gradient, x i represents the sending timestamp, y i represents the round-trip delay, and n represents the number of packets used for batch calculation.
[0028] Optionally, the average in-transit data volume is calculated by dividing the cumulative in-transit data volume by the packet count, and the sending rate is calculated by multiplying the packet count by the MTU and dividing by the sending timestamp interval.
[0029] Optionally, the window loop control adjusts the transmission window size according to the current in-transit data volume and round-trip delay.
[0030] Optionally, the rate loop control dynamically adjusts the sending rate according to the current sending rate and delay gradient.
[0031] In a second aspect, the present invention provides an efficient data center network congestion control system that achieves O(1) convergence, including:
[0032] A batch processor module for using a batch processor to obtain the network state and batch-calculate the evaluation metrics of the network state, including average delay, delay gradient, average in-transit data volume, and sending rate;
[0033] A control loop module for calling a network state estimation function according to the obtained network state, analyzing the current delay, bandwidth, and queue situation of the network, and triggering window loop control and rate loop control of the network when the network delay indicates potential underutilization of bandwidth.
[0034] The beneficial effects of the present invention are:
[0035] 1. Improve network resource utilization and transmission efficiency: Through real-time network state estimation and dual-loop control, dynamically adjust the sending rate and window size, make full use of high-bandwidth network resources, significantly improve throughput, and reduce bandwidth waste.
[0036] 2. Reduce latency and congestion risk: Based on the dual-loop control mechanism, give priority to accelerating transmission when the latency is low, and actively decelerate when the latency is high, avoiding high latency caused by queue backlog, and ensuring low-latency and efficient transmission.
[0037] 3. Achieve fairness and stability: Through the fairness adjustment module and target latency control, avoid the problem of resource monopolization in multi-flow competition, and at the same time limit excessive adjustment to ensure the stability and fairness of the transmission process. Description of the Drawings
[0038] Figure 1 is the pseudocode of the working process of the batch processor module.
[0039] Figure 2 is the pseudocode of the working process of the control loop module. Detailed Implementation Manner
[0040] Now, the present invention will be further described in detail with reference to the accompanying drawings.
[0041] The present invention proposes an efficient data center network congestion control system that achieves O(1) convergence, which is used to execute an efficient data center network congestion control method that achieves O(1) convergence, mainly including the design of the following modules.
[0042] I. Batch Processor Module
[0043] The present invention uses a batch processor to obtain the network state. The algorithm pseudocode of this module is as Figure 1 shown. The batch processor module of the present invention performs network state evaluation through Batched Least Squares (BLS), mainly used to calculate congestion signals (such as latency and latency gradient) and reference states. This method is designed for modern data center networks with high bandwidth and low latency, and has the characteristics of high precision and low overhead. The design goal of the module is to estimate the latency gradient with O(1) complexity to meet the real-time control requirements of high-performance networks.
[0044] As Figure 1As shown in the pseudocode, the batch processor uses RTT (Round-Trip Time) as the core metric for delay measurement. The sender embeds the send timestamp (sendTs) in the header of the data packet, which occupies 4 bytes. The receiver echoes sendTs to the ACK packet. When calculating RTT, the sender calculates RTT by the difference between the timestamp of the received ACK (recvTs) and sendTs. This method avoids storing additional states on the network card and meets the high-precision requirements at the same time.
[0045] When estimating the delay gradient, in a high-bandwidth and low-latency network, the traditional delay difference method is vulnerable to noise. Especially when the sending interval is extremely short, it may cause the calculation error to be amplified. The present invention estimates the delay gradient by the batch least squares method, which has higher precision and robustness. The present invention regards time as the independent variable and delay as the dependent variable, and calculates the relationship between time and delay by the least squares method. To reduce the storage overhead, the module only needs to record the following 5 cumulative values: the cumulative value of the send timestamp, the cumulative value of the delay, the cumulative value of the square of the send timestamp, the cumulative value of the product of the send timestamp and the delay, and the packet count. By means of batch calculation, the per-packet calculation is avoided, reducing the computational complexity. The calculation method is shown in the following formula.
[0046]
[0047] where x i represents the send timestamp of the data packet, y i represents the delay of the data packet, and n represents the number of data packets used for batch estimation. It can be found from the formula that only five cumulative values need to be stored for calculating the delay gradient. Therefore, the module adopts the batch calculation method, processes the cumulative values every fixed time window τ, calculates RTT and updates the above cumulative values each time an ACK is received. When the number of accumulated ACKs reaches more than 3 and the current time exceeds the set window from the batch start time, batch calculation is triggered. The batch calculation mainly calculates the average delay, delay gradient, average in-flight data volume, and sending rate, as the core metrics for network state evaluation. After the calculation is completed, the cumulative values are cleared and the batch start time is updated. Since the timestamp is stored in 4 bytes and the unit is nanosecond, the timestamp will overflow after about 4.3 seconds. To solve this problem, the module adds 1-bit storage overhead to identify whether the timestamp overflows.
[0048] The batch processor module of the present invention realizes efficient and accurate network state evaluation through the batch least squares method, and is especially suitable for modern data center networks with high bandwidth and low latency. Its design takes into account both hardware resource limitations and high-precision control requirements, providing reliable technical support for congestion control.
[0049] In multiple-increase / multiple-decrease (MIMD) updates, the correct selection of the reference state is crucial for avoiding network fluctuations and overreactions. Traditional methods (such as HPCC) perform updates based on a reference window before one RTT. This lagging design is prone to overreactions. The present invention overcomes this defect of traditional methods by introducing a reference state synchronized with the congestion signal. When an ACK is received, the module uses a batch processor to calculate the average in-flight data volume and the sending rate, and updates the window based on this, thereby avoiding overreactions.
[0050] The idea of how the batch processor module solves the problem is as Figure 1 shown in the pseudocode. When sending data packets, the current in-flight data volume (pktInfl) is embedded in the header, and the receiving end echoes this information in the ACK. The batch processor calculates the average in-flight data volume and the sending rate based on the returned ACK as the reference state update. The batch processor calculates the average in-flight data volume and the sending rate by accumulating pktInfl values and timestamps, rather than directly using the window and rate specified by the CC algorithm. Several key metrics are involved in the calculation. The average in-flight data volume is calculated by accumulating pktInfl values and dividing by the packet count. The actual sending rate is calculated by multiplying the packet count by the MTU and dividing by the sending timestamp interval.
[0051] II. Control Loop Module
[0052] This module dynamically adjusts the data transmission window and the sending rate through an innovative dual-loop control mechanism to optimize throughput under different network conditions, achieve low-latency transmission, and avoid high latency caused by queue backlogs.
[0053] The present invention uses two independent control loops to establish its control law. To coordinate the control loops, first, it is determined how to coordinate the window and the rate to ensure that they can converge to the same target. During the coordination process of the window and the rate, the present invention determines which control loop to adopt according to the relationship between the current network state and the target network state. The pseudocode of the control law is given in Figure 2 the following. When a new ACK is received, the control loop module feeds the information of the packet to the batch estimator. The congestion control action is executed after each estimation.
[0054] The working process of the module is as shown in the pseudocode Figure 2 below. First, based on the network state estimation obtained from the batch processor module, through timestamps and the average in-flight data volume, the network state estimation function is called to analyze the current delay, bandwidth, and queue situation of the network. If the network state cannot be effectively estimated, it directly returns to avoid incorrect adjustment. If the network delay indicates potential underutilization of bandwidth (i.e., no queue backlog), the module triggers dual-loop control to quickly increase the sending rate to fully utilize the network bandwidth.
[0055] Dual-loop control refers to the window loop and the rate loop. Window loop: Adjusts the transmission window size based on the current in-transit data volume and latency to ensure that data is not over-sent during network congestion. Rate loop: Dynamically adjusts the sending rate according to the current sending rate and the trend of latency change (latency gradient) to adapt to changes in network conditions. The two loops use a coordination mechanism to prioritize acceleration at low latency and deceleration at high latency to achieve balance. On the basis of dual-loop adjustment, the module will introduce an additional fairness adjustment item to ensure reasonable resource allocation in multi-flow competition and avoid a single flow occupying too much bandwidth.
[0056] Finally, the module calculates the new window size and sending rate based on the adjusted utilization rate, while ensuring that these values do not exceed the basic capabilities of the network.
[0057] Next, a specific example is used to illustrate the execution of an efficient data center network congestion control method (hereinafter referred to as "algorithm") for achieving O(1) convergence proposed by the present invention. This algorithm can be implemented on different platforms, such as being implemented based on DPDK or based on RNIC. The following describes the implementation on DPDK and calculates its implementation overhead in step 3:
[0058] Step 1: Implement the congestion control method of the present invention in the Linux 5.4.0 kernel, compile it using DPDK 25.03 and GCC8.4, and set the optimization level to the highest.
[0059] Step 2: Since the network card does not support hardware timestamps, the present invention uses the CPU's Time Stamp Counter to record the timestamps required for RTT calculation.
[0060] Step 3: To evaluate the computational overhead of the algorithm, run it one million times and take the average. The algorithm mainly involves addition, multiplication, and comparison operations, which are triggered by each ACK; the batch processor and rate update are triggered every fixed time τ and include 6 division operations. The test results show that: without batch estimation (w / o BE): the overhead is 34 CPU cycles; with batch estimation (w / BE): the overhead is 52 CPU cycles; in contrast, the computational overhead of congestion control algorithms based on exact INT (such as HPCC and PowerTCP) is significantly higher (assuming a typical data center hop count of 5, they are 92 and 159 cycles respectively).
[0061] In addition, this algorithm can also be implemented on an RDMA network card (RNIC). The main challenges in implementing congestion control on RNIC lie in the use of memory space and timers. The present invention avoids the need for additional timers through optimized design and uses timestamp comparison to trigger the batch estimator batch processor.
[0062] The algorithm of the present invention is implemented on the RNIC and only needs to additionally store 7 variables: two of the variables are stored using 8 bytes; the remaining variables are stored using 4 bytes, totaling 36 bytes. Compared with the existing RNIC implementation of DCQCN (about 60 bytes), the storage overhead of the present invention is lower, having a significant resource advantage.
[0063] Generally speaking, the present invention significantly reduces the computational and storage overhead through algorithm complexity optimization, while avoiding the need for additional timers. The experimental results show that the algorithm is superior to the prior art in terms of performance and resource utilization, and is suitable for the efficient congestion control requirements of modern data centers.
[0064] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. An efficient data center network congestion control method for achieving 0(1) convergence, characterized in that: include: Use batch processors to obtain network status and batch calculate network status evaluation indicators, including average delay, delay gradient, average in-transit data volume, and sending rate; According to the obtained network status, the network status estimation function is called to analyze the current delay, bandwidth and queue status of the network, and when the network delay indicates that there is potential underutilization of bandwidth, the window loop control and rate loop control of the network are triggered.
2. The method for efficient data center network congestion control for achieving 0(1) convergence according to claim 1, characterized in that: The evaluation index of the batch computing network status is triggered in the following way: The sender embeds a sending timestamp in the header of the data packet; The receiving end will echo the sending timestamp to the ACK message; The sender calculates the round-trip delay by the difference between the timestamp of the received ACK message and the timestamp of the sent message; Record cumulative values, including the cumulative value of the sending timestamp, the cumulative value of the round-trip delay, the cumulative value of the sending timestamp squared, the cumulative value of the sending timestamp multiplied by the round-trip delay, and the packet count; Each time the sender receives an ACK message, it calculates the round-trip delay and updates the cumulative value; when the accumulated number of ACK messages reaches the set threshold and the current time and the batch start time exceed the set window, batch calculation is triggered.
3. The method for efficient data center network congestion control achieving O(1) convergence according to claim 2, characterized in that: The evaluation index of the network status is calculated in batches according to the recorded cumulative values. After the calculation is completed, the cumulative values are cleared and the batch start time is updated.
4. The method for efficient data center network congestion control for achieving 0(1) convergence as claimed in claim 2, characterized in that: The average delay is calculated by averaging the accumulated values of the round-trip delays, and the delay gradient is calculated by the following formula: In the formula, g represents the delayed gradient, x i Indicates the sending timestamp, y i represents the round-trip delay, and n represents the number of data packets used for batch calculation.
5. The method for efficient data center network congestion control achieving O(1) convergence as claimed in claim 3, characterized in that: The average in-transit data volume is calculated by dividing the accumulated in-transit data volume by the packet count, and the sending rate is calculated by multiplying the packet count by the MTU and dividing the result by the sending timestamp interval.
6. The method for efficient data center network congestion control for achieving 0(1) convergence according to claim 1, characterized in that: The window loop control is to adjust the transmission window size according to the current amount of in-transit data and the round-trip delay.
7. The method for efficient data center network congestion control for achieving 0(1) convergence according to claim 1, characterized in that: The rate loop control dynamically adjusts the sending rate according to the current sending rate and delay gradient.
8. An efficient data center network congestion control system that achieves 0(1) convergence, characterized in that: include: The batch processor module is used to obtain the network status by using the batch processor and batch calculate the evaluation indicators of the network status, including the average delay, delay gradient, average in-transit data volume and sending rate; The control loop module is used to call the network status estimation function according to the acquired network status, analyze the current delay, bandwidth and queue status of the network, and trigger the window loop control and rate loop control of the network when the network delay indicates that there is potential underutilization of bandwidth.