A data center transmission control system and method capable of filling network bandwidth

By combining the data center network transmission protocol of centralized transmission control and active transmission control, bandwidth scheduling is used to use the global view of the centralized controller, the problem of global optimal decisions cannot be made in the existing active transmission control solution, and the effect of high bandwidth utilization and low computing overhead is achieved.

CN115733755BActive Publication Date: 2025-06-27TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211421327.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-06-27
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

In the data center network, the existing active transmission control solution can only make decisions based on local information, so it is impossible to make global optimal decisions, resulting in waste of bandwidth and inflexible streaming.

Method used

A data center network transmission protocol combining centralized transmission control and active transmission control is proposed. The centralized controller collects global traffic information, detects whether there is bandwidth wasted or bandwidth not being used in the network, and enables the host to establish a data transmission channel through a centralized scheduling packet to ensure the high bandwidth utilization of the network.

Benefits of technology

It realizes the rapid and gentle use of the remaining bandwidth of the entire network, avoids congestion and packet loss caused by multiple rounds of matching delays and Overcommitment mechanisms, and ensures low computing overhead and high bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115733755B_ABST
    Figure CN115733755B_ABST
Patent Text Reader

Abstract

The present invention discloses a data center transmission control system and method capable of filling network bandwidth. The system includes a centralized controller, a host-side transmission controller, and a data center network. The centralized controller maintains the real-time state of the network using flow information packets, connects to the corresponding host-side using flows, divides the entire network into multiple sub-bipartite graphs, calculates the sub-bipartite graphs when the flow information is updated, performs distributed control on simple and deterministic traffic by using the active transmission control strategy of the host-side, searches for available bandwidth, sends control messages to the sender and receiver of the newly established transmission, and completes the filling of the entire network bandwidth. Compared with the prior art, the present invention can utilize the remaining bandwidth of the entire network quickly and gently, avoiding congestion and packet loss in the network; 2) it can ensure high bandwidth utilization of the entire network while ensuring very low computational overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of flow control or transmission control, and particularly to a data center network transmission protocol that combines centralized transmission control and active transmission control. Background Art

[0002] With the growth of Internet services such as real-time audio and video, e-commerce, online games, stock trading, and virtual reality, the requirement for the latency performance of data center networks has become increasingly high, changing from the past second level to the current microsecond level. Latency directly affects user satisfaction and thus affects the revenue of enterprises. Therefore, designing a low-latency data center transmission control scheme has become a hot issue in data center networks.

[0003] Regarding the transmission control problem of data center networks, a large number of solutions have emerged in recent years. Traditional reactive congestion control (RCC) schemes need to probe the link status, blindly send data into the network, and then adjust the rate according to the transmission signal. For example, DCTCP adjusts the sending window according to the proportion of packets marked with ECN; TIMELY adjusts the sending rate according to the accurately measured round-trip time (RTT). However, this is a post-congestion adjustment method, that is, when adjusting the speed, situations such as switch queue congestion and even packet loss have already occurred, which seriously affects the latency of the network.

[0004] With the continuous development of data center network bandwidth, from 100Gbps to 400Gbps, active transmission control has increasingly demonstrated excellent effects. Active transmission control is essentially a receiver-driven transmission control scheme that uses the receiver to collect sender flow information and determines when to send which flow based on the flow size and the bandwidth of the receiver-side TOR switch to ensure high link utilization, low latency, and zero packet loss. However, in the actual deployment of data center networks, the receiver collects local flow information for all destinations of the receiver, and this flow information is only a part of the entire network topology. All receivers can only make locally optimal decisions based on local information, and the locally optimal decisions of different receivers will conflict. For example, when receiver A and receiver B both match GRANT to sender C at the same time, sender C can only choose to send data to the smaller flow of A or B, which will lead to the waste of a matched GRANT at one of the receivers and further result in bandwidth waste. To solve the above problems, some solutions have been proposed in current advanced active transmission control schemes, but they all have defects. For example, the Overcommitment mechanism proposed by Homa allows the receiver to overuse the TOR downlink bandwidth in a restricted manner and send redundant GRANTs to ensure high utilization of the receiver link. However, this aggressive strategy will cause packet accumulation at the receiver-side TOR, leading to packet loss and increased tail latency. To ensure a definite high bandwidth utilization of the entire network, Dcpim uses a multi-round matching method to match senders and receivers to ensure high bandwidth utilization of the entire network. However, to ensure the matching effect, each round of matching in Dcpim requires 5RTT, which results in a long flow waiting time for matching and inflexible flow transmission. For example, flow A is a large flow. After A starts, it needs to wait for 5RTT of matching delay before transmission. At the 6th RTT, flow B with the same receiver as A starts. Even if the flow size of B is much smaller than A, it still needs to wait until the end of the second round of matching (the 10th RTT) to start transmission. Therefore, the current state-of-the-art active transmission control cannot well solve this problem.

[0005] The essence of the above problems in active transmission control is that the receiving end only has a part of the entire network traffic information and cannot make a globally optimal decision. To solve this problem, a centralized scheduling scheme with a global view is an ideal solution, but the existing centralized scheduling schemes also have their unavoidable defects - high overhead. For example, in the centralized transmission control algorithm represented by Fastpass, all flows need to send flow information to the centralized scheduler before sending. The centralized scheduler divides time into time slices and uses time slices as the smallest decision unit to determine the flows that can be sent within each time slice and the sending rate of the sending flows. Although Fastpass can achieve globally optimal scheduling, its decision space is huge. It needs to decide the sending situation of each flow in each time slice, and as the number of network nodes in the data center and the number of flows increase continuously, its decision space will become difficult to calculate. Summary of the Invention

[0006] To solve the above problems, the present invention aims to propose a data center transmission control scheme that can fill network bandwidth. By running an active transmission control scheme at the host side to perform distributed control when the link is idle and the flow will not cause bandwidth waste, the decision space of the centralized controller is reduced, thereby reducing the scheduling and computing overhead of the centralized controller; at the same time, the centralized controller collects traffic information, constructs a global view and detects whether there is bandwidth waste and unused bandwidth in the network. If available bandwidth is found, the host side is made to establish a data transmission channel through a centralized scheduling packet to ensure high bandwidth utilization of the entire network.

[0007] The present invention is implemented by the following technical solutions:

[0008] A data center transmission control system that can fill network bandwidth, the system includes a centralized controller, a host-side transmission controller and a data center network; wherein:

[0009] The centralized controller is used to collect flow information, monitor network status and quickly utilize idle bandwidth; the centralized controller specifically includes a globally connected information collection module and a conflict detection and collection module; the globally connected information collection module collects sender flow information packets and receiver flow information packets to construct global flow information data; the conflict detection and collection module is used to detect whether there is bandwidth waste and unused bandwidth in the network and generate centralized scheduling packets;

[0010] The host - side transmission controller is used to execute the active transmission control strategy. The host - side transmission controller specifically includes a receiving end and a sending end. The sending end is used to send flow information packets and receive and execute control messages. The sending end further includes a flow generation module, a first flow information collection module, and a sending control module that are connected in sequence. The flow generation module is connected to the sending control module. Among them, the flow generation module is used to generate flow information packets, and the generated flow information packets are respectively transmitted to the first flow information collection module and the sending control module. The flow information collection module collects flow information packets and is used to transmit the flow information packets of the sending end. The sending control module is used to transmit flow information packets. The receiving end is used to collect flow information packets, control the sending of flow permission packets, and receive centralized scheduling. The receiving end specifically includes a second flow information collection module, a transmission control module, and a flow information normal sending detection module that are connected in sequence. Among them, the second flow information collection module is used to receive the flow information packets, centralized scheduling packets, and flow permission packets of the sending end. The flow information normal sending detection module is used to output the flow information packets of the receiving end to the data center network.

[0011] The data center network is used to send the flow information packets and flow data packets of the sending end from the sending end to the centralized controller, send the flow permission packets and centralized scheduling packets from the centralized controller to the sending end of the host - side transmission controller, send the flow information packets, centralized scheduling packets, and flow data packets of the sending end from the centralized controller to the receiving end of the host - side transmission controller, and send the flow permission packets from the receiving end of the host - side transmission controller to the centralized controller.

[0012] A data center transmission control method for filling network bandwidth includes the following steps:

[0013] The centralized controller uses the flow information packets to maintain the real - time network state, connects to the host - side corresponding to the flow connection, divides the entire network into multiple sub - bipartite graphs. When the flow information is updated, it calculates the sub - bipartite graphs, and through the active transmission control strategy on the end - host side, it performs distributed control on simple and deterministic traffic, searches for available bandwidth, sends control messages to the sending end and receiving end of the newly established transmission, and completes the rapid and accurate filling of the entire network bandwidth.

[0014] Compared with the prior art, the present invention can achieve the following positive technical effects:

[0015] 1) It can quickly and gently utilize the remaining bandwidth of the entire network, without waiting for the multi - round matching delay of dcpim, and avoid making the network suffer from congestion and packet loss brought by the Overcommitment mechanism of Homa.

[0016] 2) It can quickly and gently fill the available bandwidth of the entire network while ensuring very low computational overhead, thereby ensuring high bandwidth utilization of the entire network. Description of the Drawings

[0017] Figure 1 is an architecture diagram of a data center transmission control system capable of filling network bandwidth according to the present invention;

[0018] Figure 2 is a flowchart of the centralized controller;

[0019] Figure 3 is a flowchart of the sending end;

[0020] Figure 4 is a flowchart of the receiving end. Detailed Embodiments

[0021] The present invention will be further described in detail below in conjunction with the drawings and embodiments.

[0022] As Figure 1 shown, it is an architecture diagram of a data center transmission control system capable of filling network bandwidth according to the present invention. The system mainly includes a centralized controller 100, a host-side transmission controller 200, and a data center network 300.

[0023] The centralized controller 100 is used to collect flow information, monitor network status, and quickly utilize idle bandwidth. The centralized controller 100 specifically includes a global information collection module 110 and a conflict detection and collection module 120 that are connected to each other. The data center network 300 has two output ends connected to the global information collection module 110, and is used to transmit sender flow information packets and receiver flow information packets. The global information collection module 110 collects sender flow information packets and receiver flow information packets from the data center network 300 to construct a global view. One output of the conflict detection and collection module 120 is connected to the data center network 300, and is used to detect whether there is bandwidth waste and unused bandwidth in the network, and send centralized scheduling packets to the data center network for flow data packet control.

[0024] The host-side transmission controller 200: The host-side transmission controller 200 specifically includes a receiving end 210 and a sending end 220. The sending end 210 is used to send flow information packets, and receive and execute control messages. The sending end 210 further includes a flow generation module 211, a first flow information collection module 212, and a sending control module 213;

[0025] The flow generation module 211, the first flow information collection module 212, and the sending control module 213 are connected in sequence, and the flow generation module 211 is directly connected to the sending control module 213 through another path; wherein, the flow generation module is used to generate flow information packets, and the generated flow information packets are respectively transmitted to the first flow information collection module 212 and the sending control module 213. The flow information collection module 212 collects flow information packets. This module includes two outputs, one connected to the input end of the sending control module, and the other connected to the data center network for transmitting the sender flow information packets. One output of the sending control module 213 is to the data center network 300 for transmitting flow information packets. The data center network 300 has two outputs connected to the sending control module 213 for transmitting flow permission packets and centralized scheduling packets to the sending control module 213.

[0026] The receiving end 220 is used to collect flow information packets, control the sending of GRANT (permission), and receive centralized scheduling. The receiving end 220 further includes a second flow information collection module 221, a transmission control module 222, and a flow information normal sending detection module 223. The second flow information collection module 221, the transmission control module 222, and the flow information normal sending detection module 223 are connected in sequence; wherein, the second flow information collection module 221 receives three outputs from the data center network 300, namely, including sender flow information packets, centralized scheduling packets, and flow permission packets. The flow information normal sending detection module 223 is used to output the receiving end information packets to the data center network 300;

[0027] The data center network 300 is used to send the sender flow information packets and flow data packets from the sender 210 to the centralized controller 100, send the flow permission packets and centralized scheduling packets from the centralized controller 100 to the sender 210 of the host-side transmission controller 200, send the sender flow information packets, centralized scheduling packets, and flow data packets from the centralized controller 100 to the receiving end 220 of the host-side transmission controller 200, and send the flow permission packets from the receiving end 220 of the host-side transmission controller 200 to the centralized controller 100.

[0028] A data center transmission control method capable of filling network bandwidth according to the present invention uses an active transmission control strategy at the host side (the host side is located on the service side) to perform distributed control on simple and deterministic traffic, reducing the number of end hosts that the centralized controller needs to control and the number of flows that need to be calculated, greatly reducing the decision-making scope and calculation amount of the centralized controller. At the same time, the centralized controller uses flow information packets to maintain the real-time state of the network, uses flow connections to correspond to end hosts, divides the entire network into multiple sub-bipartite graphs, and when the flow information is updated, only needs to calculate the sub-bipartite graphs to find out whether there is free bandwidth. Greatly reducing the calculation amount and control delay of the centralized controller.

[0029] As shown Figure 2 in the figure, it is the flowchart of the centralized controller of the present invention.

[0030] Step S21: The centralized controller receives a flow information packet of a flow, and reads the source / destination IP in the flow information packet;

[0031] Step S2: Store the flow information, that is, the global flow information data, and update the global traffic information data to the corresponding

[0032] host-side information to maintain the consistency of the global flow view;

[0033] Step S23: Use the recorded global traffic information to detect whether the host side meets the autonomous control (the basis for detection is to check that the new flow is the minimum flow at both the receiving end and the sending end, indicating that the flow will establish a connection autonomously on the host side). If it meets, then execute Step S27 to update the network information including the flow information of the sending end and the receiving end;

[0034] If it does not meet, execute Step S24 to store the new flow information;

[0035] Step S25: Establish a bipartite graph of the sending end and the receiving end for the entire current network, record the traffic to the bipartite graph where the receiving end and the sending end are located, and perform a weighted maximum matching on the bipartite graph; among them, in the bipartite graph, whether there is a connection between two points depends on whether there is a flow between the sending end and the receiving end, and the length value of the connection is determined by the minimum flow size between the receiving end and the sending end. The larger the minimum flow, the smaller the length value of the connection, and vice versa, the larger the length value of the connection. Use the KM algorithm to calculate the maximum matching of the entire weighted bipartite graph; the process of constructing the bipartite graph is as follows: One side of the bipartite graph consists of all the sending ends, and the other side is all the receiving ends. Traverse all the sending ends, traverse each flow for each sending end, and establish a connection between the sending end and the receiving end of the flow. The connection value is determined by the remaining flow size of the flow. The larger the flow size, the smaller the value (all greater than 0). For each sending end, if there is no flow to other receiving end nodes, then also establish a connection, and the connection value is 0; after establishing the bipartite graph, run the KM algorithm;

[0036] Step S26: Check the difference between the new matching result and the current network state, compare the new matching with the actual traffic in the current network, send a centralized scheduling packet to the nodes with the newly established matching, establish a connection on this link, utilize the idle link, improve the bandwidth utilization rate of the entire network, and thus improve the network throughput;

[0037] Step S27, update network information; check if there is available idle bandwidth. If so, send control messages to the sender and receiver of the newly established transmission to quickly and accurately fill the entire network bandwidth.

[0038] As Figure 3 shown, it is the sender flowchart.

[0039] Step S31, when a new flow arrives at the sender, first use the flow information collection module to collect the flow information packet, which includes the flow sender IP, flow receiver IP, flow size, and the ranking of the flow size at the sender, etc. Send this flow information packet to the centralized controller and the receiver, and at the same time wait for the control message.

[0040] Step S32, determine whether to receive the centralized scheduling packet, that is, the control message; specifically, when the sender receives the control message, the control message includes the flow grant (GRANT) packet sent by the receiver and the centralized scheduling packet sent by the centralized controller.

[0041] Step S33, if the sender receives the flow grant packet from the receiver, record it as the active control state; record the current node sending state as the active transmission control controlled by the receiver, and record the current sending traffic information, including the flow five-tuple, flow size, etc.

[0042] Step S34, if the sender receives the centralized control message sent by the centralized controller, record it as the centralized control state; record the current node sending state as the centralized scheduling controlled by the centralized controller, and record the current sending traffic information, including the flow five-tuple, flow size, etc.

[0043] Step S35, send the flow information packet, that is, send the corresponding flow information packet at the line speed (the line speed value refers to the theoretical maximum speed of the physical link, which is determined by the performance of the sender network card and the network link bandwidth connected to the sender network card. Take the smaller value of the network card performance and the link bandwidth. For example, for a 10G network card and a 100G link bandwidth, the line speed is 10G, which is the theoretical value that can be achieved) according to the indication of the control message.

[0044] In the above process, if the flow ranking order on the sender side changes due to the addition of a new flow or the end of a flow, then the new smallest flow will send an updated flow information packet to the centralized controller and the receiver to synchronize the change in traffic information.

[0045] As Figure 4 shown, it is the receiver flowchart.

[0046] Step S41, judge the type of the data packet received by the receiver; if it is a certain data packet type, directly execute Step S45.

[0047] Step S42: If it is a flow information packet, store the flow information in the flow information table, and then check whether the flow information size is the smallest flow at the sending end (specifically, check the size ranking of this flow in the sending end in this flow information packet, and whether the flow size ranking is the smallest flow at the sending end); specifically, if the flow size is less than one BDP, then the host side marks this flow as a small flow and directly sends this flow at line speed. If the flow size is greater than one BDP, then the flow information collection module on the host side collects new flow information;

[0048] If this flow is the smallest flow at the sending end side, execute Step S43: Judge the current receiving end status; if the current receiving end is in the idle state, execute Step S44, send a flow permission packet, and count the flow information; if there is a flow being received in the current receiving end status, then execute Step S47, and judge whether the current flow is greater than the flow being received at the receiving end;

[0049] If the current flow is greater than the flow being received at the receiving end, then execute Step S46, store the flow information; if the current flow is small, indicating that this flow is also the smallest flow at the receiving end side, then execute Step S44 to send a flow permission packet, and the receiving end updates the flow being received to this flow;

[0050] If this flow is not the smallest flow at the sending end side, execute Step S46, store the flow information, store the flow information in the flow information table, record the sending and receiving record information of this data flow, and return the ACK of this data packet;

[0051] Step S45: Receive flow data packets until the flow ends.

[0052] In summary, in the above process, when the receiving end receives a flow information packet of a new flow, store the flow information packet of this flow, detect the data packet type; and query whether the flow size is the smallest flow at both the receiving end and the sending end at the same time. If so, then send GRANT to the sending end of this flow. If not, then continue to send the current traffic or wait idle for scheduling. When the receiving end receives the centralized scheduling message from the centralized controller, receive the corresponding traffic at line speed according to the indication of the control message.

[0053] Large-scale simulation experiments were conducted on YAPS in this invention. Real workloads such as WebServer, CacheFollower, WebSearch, and DataMining were adopted. Using YAPS for simulation experiments, the network topology is a leaf-spine topology, including 144 servers with a 100G bandwidth; Dcpim was implemented on YAPS as a comparative experiment to compare the effects of this method and Dcpim under different network loads. GoodPut was used as the main measurement index. This index divides the flow size after all flows end by the flow completion time. This index can be understood as the effective network throughput for applications and is a more effective index than comparing throughput because when retransmissions are included in throughput, the network bandwidth is not effectively utilized. Under different load conditions from 0.1 to 0.8, the Goodput of the entire network increased by 20 - 50%, which in turn led to a reduction in the overall flow FCT. And because this invention is more flexible than dcpim, the completion time of smaller flows (less than or equal to 10 * BDP) decreased by 2 - 3 times, which is a great performance improvement for delay-sensitive small flows.

[0054] All in all, this invention can quickly detect the usage status of network links with relatively low computational overhead, accurately and gently fill the network idle bandwidth, and ensure high throughput, low latency, and low packet loss of the entire network.

[0055] The above description of the technical solution does not limit the technical content of this invention. For those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A data center transmission control system capable of filling network bandwidth, characterized in that, The system includes a centralized controller, a host-side transmission controller, and a data center network; where: The centralized controller is used to collect flow information, monitor network status, and quickly utilize idle bandwidth; the centralized controller specifically includes a globally information collection module and a conflict detection and collection module connected to each other; the globally information collection module collects sender flow information packets and receiver flow information packets to construct global flow information data; the conflict detection and collection module is used to detect whether there is bandwidth waste and unused bandwidth in the network, generate a centralized scheduling packet, and send the centralized scheduling packet to the data center network for flow data packet control; The host-side transmission controller is used to execute an active transmission control strategy. The host-side transmission controller specifically includes a receiver and a sender; the sender is used to send flow information packets, and receive and execute control messages; the sender further includes a flow generation module, a first flow information collection module, and a transmission control module connected in sequence, and the flow generation module is connected to the transmission control module; where, the flow generation module is used to generate flow information packets, and the generated flow information packets are respectively transmitted to the first flow information collection module and the transmission control module. The flow information collection module collects flow information packets and is used to transmit sender flow information packets; the transmission control module is used to transmit flow information packets; the receiver is used to collect flow information packets, control the sending of flow permission packets, and receive centralized scheduling; the receiver specifically includes a second flow information collection module, a transmission control module, and a flow information normal sending detection module connected in sequence; where, the second flow information collection module is used to receive sender flow information packets, centralized scheduling packets, and flow permission packets; the flow information normal sending detection module is used to output receiver flow information packets to the data center network; The data center network is used to send sender flow information packets and flow data packets from the sender to the centralized controller, send flow permission packets and centralized scheduling packets from the centralized controller to the sender of the host-side transmission controller, send sender flow information packets, centralized scheduling packets, and flow data packets from the centralized controller to the receiver of the host-side transmission controller, and send flow permission packets from the receiver of the host-side transmission controller to the centralized controller.

2. The data center transmission control system capable of filling network bandwidth according to claim 1, wherein The process on the centralized controller side specifically includes the following steps: Receive a flow information packet of a flow, store it as global flow information data, and update the global traffic information data to the corresponding host-side information; establish a bipartite graph of senders and receivers for the current entire network, record the traffic to the bipartite graph where the receiver and sender are located, and perform a weighted maximum matching on the bipartite graph; compare the obtained weighted maximum matching result with the actual traffic in the current network as a new matching, send a centralized scheduling packet to the nodes of the newly established matching, establish a connection on the link to realize the utilization of idle links; update network information; in the case of existing idle bandwidth, send a control message to the new matching, the sender and receiver of the newly established transmission, to complete the rapid and accurate filling of the entire network bandwidth. The process on the sender side specifically includes the following steps:

3. The data center transmission control system capable of filling network bandwidth according to claim 1, characterized in that The sender process specifically includes the following steps: When a new flow arrives at the sender, the flow information packet is sent to the centralized controller and the receiver, waiting for the flow permission packet sent by the receiver and the centralized scheduling packet sent by the centralized controller; if the sender receives a flow permission packet, it is recorded as the active control state, and if the sender receives a centralized scheduling packet, it is recorded as the centralized control state; the corresponding flow information packet is sent at line speed according to the indication of the control message.

4. The data center transmission control system capable of filling network bandwidth according to claim 3, characterized in that, At the sender, if the flow ranking order on the sender side changes due to the addition of a new flow or the end of a flow, then the new minimum flow will send an updated flow information packet to the centralized controller and the receiver to synchronize the change in traffic information.

5. The data center transmission control system capable of filling network bandwidth according to claim 1, characterized in that, The receiver process includes the following steps: Perform policy analysis on the flow information packet received by the receiver: for the flow information of the minimum flow of the sender, when the current receiver state is the idle state, send a flow permission packet and count the flow information; when the current receiver state is that there is a flow being received, for the current flow that is the minimum flow on the receiver side, send a flow permission packet; for the current flow that is not the minimum flow on the receiver side, store the flow information, store the flow information packet in the flow information table, record the sending and receiving record information of the flow information packet, and return the confirmation of the flow information packet until the flow ends.