Data center network transmission enhancement method based on bandwidth resource management

By implementing the virtual queue and delay range allocation mechanism in the data center network, the problem of difficult multi-priority scheduling in the prior art is solved, and the bandwidth utilization rate and QoS guarantee of the network are improved.

CN120110993APending Publication Date: 2025-06-06NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249364.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When existing data center networks support multiple traffic types, they are limited by protocols and hardware resources, making it difficult to achieve flexible multi-priority scheduling, resulting in limited refinement of QoS guarantees.

Method used

By implementing virtual queues and delay range allocation mechanisms in a single physical queue, prioritization and delay management are performed, and network transmission rules are dynamically adjusted to achieve multi-priority scheduling.

Benefits of technology

It effectively guarantees the performance of high-priority traffic, and at the same time improves the bandwidth utilization of low-priority traffic, improves the overall bandwidth utilization of the network, and makes it easy to deploy without hardware changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110993A_ABST
    Figure CN120110993A_ABST
Patent Text Reader

Abstract

The invention discloses a data center network transmission enhancement method based on bandwidth resource management, and belongs to the technical field of network transmission, a data center is provided with physical priority queues, and each physical priority queue comprises a plurality of network transmission flows. Each network transmission flow has a corresponding original congestion control algorithm and ideal round-trip time, and the transmission enhancement algorithm comprises the following steps: for a physical priority queue, performing priority ranking on all the network transmission flows contained in the physical priority queue to form a virtual queue, and allocating independent delay ranges to all the network transmission flows of the virtual queue; measuring the current time delay of each network transmission flow; and judging the size relationship among the current time delay, the time delay range and the ideal round-trip time, and obtaining a transmission rule for real-time network transmission. According to the method, the performance of high-priority traffic is guaranteed, the bandwidth utilization rate of low-priority traffic is improved, meanwhile, hardware does not need to be changed, and the method is easy to deploy and efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network transmission, and in particular relates to a data center network transmission enhancement method based on bandwidth resource management. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, data centers have become the core of modern information technology infrastructure. As a key component, data center networks carry the transmission and computing tasks of massive data. Their performance directly determines the overall quality of service (QoS) of data centers. In data center networks, traffic scheduling and congestion control technologies are the core technical fields for ensuring network performance. Their main goals are to optimize resource utilization, reduce network latency, improve throughput, and meet the QoS requirements of different services. In addition, as the scale and complexity of data center networks have increased significantly, the types of services carried have become more diverse, including but not limited to the following categories: Delay-sensitive real-time services (such as voice communication, video conferencing, etc.) are extremely sensitive to network latency and require strict low latency guarantees. Computing tasks that rely on coflows (such as distributed computing, batch processing tasks, etc.) need to coordinate the transmission of multiple flows to ensure the completion time of the overall task. Throughput-sensitive backend tasks (such as big data analysis, storage backup, etc.) mainly focus on the efficient use of network bandwidth.

[0003] The above services have strict and diverse requirements for quality of service (QoS), such as latency, bandwidth, throughput, etc., and the QoS requirements of different services may vary significantly. Therefore, in order to meet the QoS requirements of different traffic types, data center networks usually adopt a priority queue mechanism to achieve traffic isolation through weighted bandwidth sharing to avoid mutual interference between different traffic. Priority queues can not only provide isolation between different types of traffic, but also support priority-based scheduling algorithms to improve application performance. For example, in the same type of traffic, priority queues can achieve strict priority scheduling of independent flows, collaborative flows (coflow), remote procedure calls (RPCs), and distributed machine learning model training tasks.

[0004] However, although priority-based isolation and scheduling can significantly optimize the performance of data center networks in theory, its practical application faces the following limitations:

[0005] (1) Limitations in protocol support. The protocols commonly used in current data center networks have significant limitations on the number of priorities they support. For example, the DSCP field in the IP header: This field is used to specify the priority, but according to existing standards, only 12 priorities are supported. The same is true for the PFC protocol, which is a key protocol for providing lossless priority for remote direct memory access (RDMA), but it only supports 8 priorities. This limitation at the protocol level directly leads to a shortage of physical priority queues, making it impossible for data centers to flexibly assign sufficient priorities to multiple traffic types, thus affecting the level of refinement of QoS guarantees.

[0006] (2) Hardware resource limitations. As data center network bandwidth continues to grow, the expansion rate of switch buffers lags far behind the growth rate of bandwidth. This contradiction between bandwidth growth and insufficient buffer resources further exacerbates the difficulty of supporting more priorities. On the Trident 2 switch, Microsoft only used two lossless priorities to provide coarse-grained isolation for real-time traffic and bulk transmission traffic, respectively, to save buffers allocated for PFC headroom. However, the buffer-to-bandwidth ratio of the latest generation of switches has been reduced by half compared to Trident 2, making it more difficult to support more lossless priorities.

[0007] (3) The scarcity of priority queues leads to traffic mixing problems and scheduling algorithm constraints. Due to the limited number of physical priority queues, data centers face many problems in traffic scheduling. The traffic mixing problem is caused by the fact that due to the insufficient number of priority queues, traffic using different congestion control (CC) algorithms has to be mixed into the same queue. This way of mixing traffic increases the complexity of deploying new congestion control algorithms and reduces the effectiveness of traffic isolation. The scheduling algorithm constraint problem is that the insufficient number of priority queues hinders the application of priority-based traffic scheduling algorithms in large-scale data centers, limiting their potential to improve network performance and flexibility. Summary of the invention

[0008] In view of the deficiencies in the prior art, the present invention provides a data center network transmission enhancement method based on bandwidth resource management, which can implement strict multi-priority scheduling in a single physical queue, thereby ensuring the performance of high-priority traffic and improving the bandwidth utilization of low-priority traffic. At the same time, no hardware modification is required, and the method is easy to deploy and highly efficient.

[0009] The present invention provides the following technical solutions:

[0010] A data center network transmission enhancement method based on bandwidth resource management, wherein the data center has at least one independent physical priority queue, each of which includes a number of network transmission flows, and each of which has a corresponding original congestion control algorithm and a corresponding ideal round trip time. Data center network transmission enhancement methods include:

[0011] For the physical priority queue, all network transmission traffic contained in it is prioritized to form a virtual queue, and all network transmission traffic in the virtual queue is assigned an independent delay range;

[0012] Based on the sampling frequency, the network test tool is used to measure the current delay corresponding to each network transmission flow i to be transmitted. i ;

[0013] For each network transmission flow i, determine its corresponding current delay i Delay range, ideal round-trip time The size relationship between them is obtained, the transmission rules are obtained, and network transmission is performed in real time according to the obtained transmission rules.

[0014] Optionally, the delay range of network transmission flow i includes the target delay and limit delay The ideal round trip time The relationship with the delay range is:

[0015] Optionally, for each network transmission flow i, the corresponding current delay delay is determined i Delay range, ideal round-trip time The size relationship between them is used to obtain the transmission rules, which are as follows:

[0016] like The original congestion control algorithm of the current network transmission traffic is used for network transmission; otherwise, network transmission is performed according to the priority order of the virtual queue and the congestion of the transmission bandwidth.

[0017] Optionally, the network transmission is performed according to the priority order of the virtual queues and the congestion of the transmission bandwidth, specifically:

[0018] like Then the transmission of the current network transmission traffic is suspended;

[0019] like Then use the linear startup strategy to speed up the transmission of the current network transmission traffic;

[0020] like Then use the adaptive delay increase strategy to increase the current delay i Increase to target delay

[0021] Optionally, the use of a linear start strategy to accelerate the transmission of the current network transmission traffic is specifically: the congestion control window size within the current RTT period is increased by a set congestion control window step size before transmission.

[0022] Optionally, the adaptive delay increase strategy is used to increase the current delay delay i Increase to target delay Specifically: increase the congestion control window size in the current RTT period by the adaptive congestion control window step size N i Then transmit;

[0023] The adaptive congestion control window step size N i :

[0024]

[0025] Among them, cwnd is the size of the congestion control window in the current RTT period.

[0026] Optionally, after pausing the transmission of the current network transmission traffic, a detection packet of set bytes is sent to obtain the delay at the next time point in a manner of low bandwidth occupancy.

[0027] Optionally, the next time point is: at the current delay i On the basis of the corresponding time point, the time point after the random waiting time is added.

[0028] Optionally, the network transmission flow i limits the delay and target delay The difference between them is the sum of the fluctuation value of congestion control and the correction value of delay noise; the fluctuation value of congestion control and the correction value of delay noise are both measured statistical values; the target delay of network transmission flow i The limit delay with network transmission flow i-1 The difference is half of the fluctuation value of congestion control, where the priority of network transmission flow i is higher than that of network transmission flow i-1.

[0029] Optionally, the network transmission flow i limits the delay and target delay The difference between them is 4 microseconds, of which the fluctuation value of congestion control is 3.2 microseconds and the correction value of delay noise is 0.8 microseconds; the target delay of the network transmission flow i The limit delay with network transmission flow i-1 The difference is 1.6 microseconds.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] (1) The present invention implements strict multi-priority scheduling through a delay range allocation mechanism, effectively guarantees the performance of high-priority traffic, and at the same time, fully utilizes the remaining bandwidth of low-priority traffic without affecting high-priority traffic, thereby improving the overall bandwidth utilization of the network. In addition, the present invention can be seamlessly integrated with the existing delay-based congestion control algorithm, does not require additional hardware support, is easy to deploy, and is suitable for modern data center networks.

[0032] (2) The adaptive delay increase strategy proposed in the present invention avoids overreaction and delay fluctuation caused by the lag in network condition changes, and further optimizes the smoothness of traffic scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a structural block diagram of the physical priority queue and virtual queue of the present invention.

[0034] Figure 2 It is a flow chart of a data center network transmission enhancement method based on bandwidth resource management of the present invention.

[0035] Figure 3 It is a bar chart of simulation results in a multi-stream scenario of the present invention. DETAILED DESCRIPTION

[0036] The present invention will be further described below in conjunction with the accompanying drawings. The following examples are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention. It should be noted that the term "comprising" and any variation thereof in the specification and claims of the present invention are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0037] A data center network transmission enhancement method based on bandwidth resource management. The data center of the present application has at least one independent physical priority queue. Each physical priority queue includes a number of network transmission flows. The physical queue refers to the queue that actually exists in the switch hardware and is used to store and process network flows. Each queue can be allocated a certain bandwidth and buffer to support flows of different priorities. For a data center, there are usually a fixed number of physical priority queues (usually 8 or 12). These queues correspond to specific hardware resources. For details, please refer to the prior art. Each network transmission flow of the present application has a corresponding original congestion control algorithm and a corresponding ideal round-trip time. The original congestion control algorithms such as Swift and LEDBAT are not the same for all network transmission traffic in a physical priority queue; the ideal round-trip time corresponding to the network transmission traffic It refers to the most ideal network transmission RTT (round trip time) when the network is unobstructed.

[0038] like Figure 1 and Figure 2 As shown, a data center network transmission enhancement method based on bandwidth resource management includes the following steps:

[0039] S1: For the physical priority queue, all network transmission traffic contained in it is prioritized to form a virtual queue, and an independent delay range is allocated to all network transmission traffic in the virtual queue.

[0040] A virtual queue is a logical queue that simulates multiple levels of priority through software or algorithms without relying on the physical queues that actually exist in the hardware. Virtual queues can achieve more priorities through dynamic allocation and scheduling based on limited physical queues. Virtual queues are priority mechanisms extended through software logic and are suitable for scenarios that require more fine-grained priority scheduling.

[0041] Specifically, the delay range of network transmission flow i includes the target delay and limit delay Ideal round trip time The relationship with the delay range is:

[0042] Target delay is a key parameter used to define priority traffic. We use the same parameters as the original congestion control algorithm for easy integration. It also represents the ideal delay level that the traffic hopes to achieve.

[0043] The values ​​of target delay and limit delay are usually determined by factors such as the nature of the traffic and QoS. For high-priority channels, we hope to allocate a higher delay range to ensure the performance of high-priority traffic. For low-priority channels, when the delay exceeds the limit, low-priority traffic will suspend transmission to release bandwidth for high-priority traffic.

[0044] For two network traffic flows of different levels, the target latency of the higher-level network traffic is Higher than the limited latency of low-level network transmission traffic

[0045] More specifically, since the delay range must accommodate normal fluctuations of congestion control (CC) and delay noise (the noise may come from jitter in software and hardware, protocol offload technology (such as TSO), etc.), the delay fluctuation caused by CC oscillates around the target delay, and the noise is superimposed on the actual delay. Therefore, the network transmission traffic i limit the delay. and target delay The difference between them is the sum of the fluctuation value of congestion control and the correction value of delay noise; the fluctuation value of congestion control can refer to the fluctuation value of the original congestion control algorithm of the network transmission flow, and the fluctuation value of congestion control and the correction value of delay noise are both measured statistical values; in addition, for the two flows with a higher priority than network transmission flow i-1, the target delay of network transmission flow i is The limit delay with network transmission flow i-1 The difference is half of the fluctuation value of congestion control.

[0046] As an option, the network transmission flow i limits the delay and target delay The difference between them is 4 microseconds, of which the fluctuation value of congestion control is 3.2 microseconds and the correction value of delay noise is 0.8 microseconds; the target delay of network transmission flow i The limit delay with network transmission flow i-1 The difference is 1.6 microseconds.

[0047] S2: Based on the sampling frequency, use the network test tool to measure the current delay corresponding to each network transmission flow i to be transmitted i .

[0048] The network testing tool may refer to existing technologies, such as ping or traceroute, etc. Of course, the current delay measurement value of the original congestion control algorithm of the network transmission flow i may also be used.

[0049] S3: For each network transmission flow i, determine its corresponding current delayi Delay range, ideal round-trip time The size relationship between them is obtained, the transmission rules are obtained, and network transmission is performed in real time according to the obtained transmission rules.

[0050] Specifically, the present invention is integrated with most delay-based congestion control algorithms (such as Swift and LEDBAT), assigns target delays to traffic, adjusts windows or rates, and fine-tunes the step size of additive increases (AI). Delay metrics based on RTT (round-trip time) or OWD (one-way delay) are supported. This application needs to find the best balance between high-priority guarantees and low-priority utilization, such as maintaining the frequency of congestion signals after bandwidth is released, or quickly adjusting the delay to a set delay range.

[0051] That is, when a network flow transmits data, it has been assigned a certain delay range for this flow, and the operating system can measure its delay. i At this time, the present invention uses different delays i Different treatments are given. Specifically: The original congestion control algorithm of the current network transmission traffic is used for network transmission. Otherwise, network transmission is performed according to the priority order of the virtual queue and the congestion of the transmission bandwidth.

[0052] More specifically, it can be divided into four processing methods:

[0053] like Figure 2 As shown, (1) if The original congestion control algorithm of the current network transmission traffic is used for network transmission. That is, in this case, the transmission of the current network transmission traffic i has met the delay range set for it by the virtual queue, so network transmission can be carried out according to the transmission rules of the original congestion control algorithm.

[0054] (2) If At this time, it is inferred that there is network transmission traffic with a higher priority than the current network transmission traffic being transmitted in the data center. Therefore, the current network transmission traffic is equivalent to low-priority traffic, and the transmission of the current network transmission traffic is suspended to release bandwidth.

[0055] In this embodiment, after pausing the transmission of the current network transmission traffic, a detection packet of set bytes is sent for detection to obtain the delay at the next time point in a low bandwidth occupation manner. The detection packet of set bytes should be as small as possible, for example, 64 bytes. The next time point is: at the current delay delay iOn the basis of the corresponding time point, the time point after the random waiting time is added. For details, reference can be made to the priority-assisted CSMA / CA technology in the existing WLAN, and the probability of detection collision can be reduced by random waiting time.

[0056] (3) If If the network transmission of the data center is in a smooth state (the path bandwidth may be idle or just fully utilized), the linear start strategy is used to accelerate the transmission of the current network transmission traffic.

[0057] Increase the congestion control window size within the current RTT period and set the congestion control window step to start, avoid excessive buffer occupancy, and quickly improve bandwidth utilization.

[0058] (4) If If there is a network transmission traffic with a lower priority than the current network transmission traffic in the data center, the adaptive delay increase strategy is used to increase the current delay. i Increase to target delay

[0059] Specifically: increase the congestion control window size in the current RTT period by the adaptive congestion control window step size N i to make a transmission;

[0060] The adaptive congestion control window step size N i :

[0061]

[0062] Among them, cwnd is the size of the congestion control window in the current RTT period.

[0063] The purpose of using the adaptive delay increase strategy is to allow this flow to "grab" the bandwidth of traffic with lower priority than this flow more quickly, while avoiding network fluctuations or instability. It dynamically adjusts the window step size based on the current delay and window size to avoid overreaction. The specific approach is to first calculate how much the window size needs to be increased based on the current delay and target delay values. The increase in window step size is This allows the rate to reach D as quickly as possible. target According to this calculation method, if the current delay is far from the target delay, increase it a little more; if the current delay is close to the target value, increase it a little less.

[0064] At the transmitting end, the present invention implements the congestion control algorithm designed by the present invention based on the DPDK framework, and allocates target delays and limit delays to traffic of different priorities through a channel allocation mechanism. The target delay and limit delay of high-priority traffic are set higher to ensure its performance; the target delay and limit delay of low-priority traffic are set lower to ensure that high-priority traffic uses bandwidth first. During network transmission, the operating system monitors the actual delay of the traffic in real time; when high-priority traffic appears, low-priority traffic quickly gives up bandwidth based on delay feedback to ensure the performance of high-priority traffic. When high-priority traffic ends, low-priority traffic immediately resumes bandwidth use to make full use of the remaining bandwidth. If the actual delay exceeds the limit delay, the queue length is controlled by reducing the traffic sending rate; if the delay is close to the target delay, the sending rate is gradually increased to ensure efficient use of bandwidth. In addition, the present invention implements the functions of the present invention through hardware (such as RDMA NIC) at the receiving end; uses the memory and timer resources of RNIC to store the nine variables (a total of 13 bytes) required for the congestion control algorithm described above and an additional timer for the traffic conflict avoidance mechanism; the RNIC hardware cooperates with the sending end through delay feedback to further optimize the congestion control performance.

[0065] After combining with the existing congestion control algorithm, the present invention significantly improves the performance and resource utilization of the data center network. The performance of the present invention is evaluated using the NS3 simulator. The present invention performs fine-grained tests to prove that the present invention can support multiple priorities under heavy traffic loads and keep latency fluctuations within a specified range during severe incast. The simulation results are shown in Figure 2. Figure 3 As shown. The vertical axis is the acceleration ratio, and the horizontal axis is the three groups of experiments from left to right. The three groups of experimental results are: the acceleration ratio of high priority traffic, the acceleration ratio of low priority traffic, and the total average acceleration ratio. The six configurations of each group of experiments are from left to right: the present invention + Swift; Physical + Swift; Swift; D2TCP; the present invention + LEDBAT; LEDBAT. Among them, Swift, D2TCP, etc. are traditional congestion control algorithms, and the present invention can be deployed on their basis to enhance their performance. As can be seen from the figure, after configuring the present invention, the total acceleration ratio is significantly improved.

[0066] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention are essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.

[0067] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.

Claims

1. A data center network transmission enhancement method based on bandwidth resource management, wherein the data center has at least one independent physical priority queue, each of which includes a number of network transmission flows, and each of which has a corresponding original congestion control algorithm and a corresponding ideal round trip time It is characterized in that include: For the physical priority queue, all network transmission traffic contained in it is prioritized to form a virtual queue, and all network transmission traffic in the virtual queue is assigned an independent delay range; Based on the sampling frequency, the network test tool is used to measure the current delay corresponding to each network transmission flow i to be transmitted. i ; For each network transmission flow i, determine its corresponding current delay i Delay range, ideal round-trip time The size relationship between them is obtained, the transmission rules are obtained, and network transmission is performed in real time according to the obtained transmission rules.

2. The data center network transmission enhancement method based on bandwidth resource management according to claim 1 is characterized in that: The delay range of network transmission flow i includes the target delay and limit delay The ideal round trip time The relationship with the delay range is:

3. The data center network transmission enhancement method based on bandwidth resource management according to claim 2 is characterized in that: For each network transmission flow i, determine its corresponding current delay delay i Delay range, ideal round-trip time The size relationship between them is used to obtain the transmission rules, which are as follows: like The original congestion control algorithm of the current network transmission traffic is used for network transmission; otherwise, network transmission is performed according to the priority order of the virtual queue and the congestion of the transmission bandwidth.

4. The data center network transmission enhancement method based on bandwidth resource management according to claim 3 is characterized in that: The network transmission is performed according to the priority order of the virtual queues and the congestion of the transmission bandwidth, specifically: like Then the transmission of the current network transmission traffic is suspended; like Then use the linear startup strategy to speed up the transmission of the current network transmission traffic; like Then use the adaptive delay increase strategy to increase the current delay i Increase to target delay 5. The data center network transmission enhancement method based on bandwidth resource management according to claim 4 is characterized in that: The use of a linear start strategy to accelerate the transmission of the current network transmission traffic is specifically: the congestion control window size within the current RTT period is increased by a set congestion control window step size before transmission.

6. The data center network transmission enhancement method based on bandwidth resource management according to claim 4 is characterized in that: The adaptive delay increase strategy is used to increase the current delay i Increase to target delay Specifically: increase the congestion control window size in the current RTT period by the adaptive congestion control window step size N i Then transmit; The adaptive congestion control window step size N i : Among them, cwnd is the size of the congestion control window in the current RTT period.

7. The data center network transmission enhancement method based on bandwidth resource management according to claim 4 is characterized in that: After pausing the transmission of the current network transmission flow, a detection packet of set bytes is sent to obtain the delay at the next time point in a manner of low bandwidth occupancy.

8. The data center network transmission enhancement method based on bandwidth resource management according to claim 7, characterized in that: The next time point is: at the current delay i On the basis of the corresponding time point, the time point after the random waiting time is added.

9. The data center network transmission enhancement method based on bandwidth resource management according to claim 2, characterized in that: Network transmission flow i limit delay and target delay The difference between them is the sum of the fluctuation value of congestion control and the correction value of delay noise; the fluctuation value of congestion control and the correction value of delay noise are both measured statistical values; Target delay of network transmission flow i The limit delay with network transmission flow i-1 The difference is half of the fluctuation value of congestion control, where the priority of network transmission flow i is higher than that of network transmission flow i-1.

10. The data center network transmission enhancement method based on bandwidth resource management according to claim 9, characterized in that: Network transmission flow i limit delay and target delay The difference between them is 4 microseconds, of which the fluctuation value of congestion control is 3.2 microseconds and the correction value of delay noise is 0.8 microseconds; the target delay of the network transmission flow i The limit delay with network transmission flow i-1 The difference is 1.6 microseconds.