An Active Transmission Method for Long-Distance Data Centers to Accelerate the Transmission in the Pre-Credit Phase

By using random packet spraying and explicit path marking in the data center network, combining the connection status table and the host granularity expected bandwidth table, the problem of insufficient latency and bandwidth utilization in long-distance data center networks is solved, and efficient and low-latency lossless transmission is achieved.

CN120050242BActive Publication Date: 2025-07-22ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510499784.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing active transmission schemes have basic latency problems in long-distance data center networks, resulting in insufficient bandwidth utilization and poor latency characteristics, which cannot meet the needs of high-performance computing and artificial intelligence tasks. At the same time, the passive transmission schemes have problems with queue accumulation and high protocol complexity.

Method used

The method of random packet spraying and explicit path marking is adopted to maintain the connection status table and the host granularity expected bandwidth table in the switch, and accelerate the transmission of pre-credit phase through idle credit, ensuring that data packets are transmitted along symmetric paths and avoiding discarding and retransmission of unscheduled packets.

Benefits of technology

It realizes lossless transmission with zero packet loss in long-distance data center networks, reduces basic latency, reduces network deployment costs, improves bandwidth utilization efficiency, and avoids queue accumulation and protocol complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050242B_ABST
    Figure CN120050242B_ABST
Patent Text Reader

Abstract

The present invention discloses an active transmission method for long-distance data centers to accelerate the transmission in the pre-credit stage, which relates to the technical field of network transmission control. The method includes: when forwarding credit packets, the sending switch uses a random packet spraying scheme and explicit path marking, and randomly forwards the credit packets to the next hop in the routing table; maintaining a connection status table of the local data center and a host granularity expected bandwidth table of the remote data center at the cross-data center egress switch; using the connection status table to judge the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, determine, according to the host granularity expected bandwidth table and the connection status table, the sender of the most recently established request under the same ToR switch as the idle credit belongs to as the destination host, and forward the idle credit to the destination host; this active transmission method for long-distance data centers accelerates the credit transmission in the pre-credit stage and minimizes the network deployment cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network transmission control, and particularly to an active transmission method for long-distance data centers that accelerates the transmission in the pre-credit phase. Background Art

[0002] In the past decade, with the continuous development of business, data centers have become key infrastructure to support the digital economy. Computing, storage, and network are the three elements of a data center, and they promote each other and develop together: within the past decade, computing power has increased by hundreds of times, and the access performance of solid-state drives has increased by a hundred times compared to mechanical hard drives. The development of computing and storage has put forward new requirements for the network, demanding greater bandwidth and lower latency. More and more different types of traditional services, big data processing, high-performance computing, and artificial intelligence tasks are running on data centers, requiring more efficient and reliable data centers. For the transmission of data center networks, the TCP / IP system of traditional wide area networks is not used in modern data centers because it is designed for general complex and highly dynamic environments and requires high CPU overhead. A transmission method more suitable for the business needs and hardware facilities of data centers is needed.

[0003] At the same time, due to considerations of energy, scalability, and disaster recovery backup in data center networking, a single data center cannot be infinitely expanded, and data centers with geographical separation need to be considered. Traditional geographically separated cloud data centers are interconnected via wide area networks and carry traditional services such as email services and web services, which have low requirements for network performance. Using wide area network links from network service providers can meet the needs, but there are problems such as protocol conversion and heterogeneous network environments.

[0004] However, current artificial intelligence, such as large language models, requires ultra-large-scale cluster networking training with higher network requirements and reliable latency bandwidth conditions. More and more data center operators are starting to build dedicated cross-data center optical fibers to connect data centers. This shift from wide area network links to dedicated links gives greater freedom to the design of cross-data center transmission, allowing for transmission schemes optimized specifically for long-distance interconnection of data centers.

[0005] At present, the data center transmission solutions mainly adopted in the industry are passive, relying on passive congestion control to control traffic, avoid high network latency, and also require Priority-based Flow Control (PFC) technology to provide a lossless network with zero packet loss. The basic design intention of the passive transmission solution is to achieve low latency. In order to enable the small data flows in the data center to complete transmission as soon as possible, the method of directly sending at the maximum line speed at the time of sending is adopted. According to the network state caused by the traffic injected into the network, congestion signals are returned to the sender through different congestion signals, and the sender then passively adjusts the sending rate (or window) according to the received congestion signals. Typical congestion control solutions such as DCQCN use Explicit Congestion Notification (ECN), Swift (congestion control mechanism for data center-level networks) uses latency, and HPCC (congestion control protocol) uses more advanced in-band network telemetry signals. These passive methods all require the feedback delay of the end-to-end round-trip time and also take longer to converge, posing great challenges to the switch queues. Therefore, the PFC hop-by-hop flow control technology is usually adopted to design a lossless data center network, but PFC will bring problems such as deadlocks and head-of-line blocking. Among them, DCQCN (Data Center Quantized Congestion Notification) is a congestion control algorithm for lossless data transmission applied to the RoCEv2 network. RoCEv2 (RDMA over Converged Ethernet version 2) is a high-performance network protocol based on Ethernet, which realizes low-latency and high-throughput remote direct memory access functions through lossless Ethernet.

[0006] Compared with the traditional passive transmission solution, active transmission works in a different working paradigm of "request-allocation". Compared with the "try-backoff" method of the passive solution, active transmission has an inherent advantage in avoiding congestion due to its characteristic of pre-allocating bandwidth resources. By reasonably scheduling and allocating bandwidth resources in advance, the shallow queues in the network can be controlled, and at the same time, the requirements of low latency and high bandwidth can be met. Active transmission is divided into two categories: centralized and distributed. The centralized type uses a central arbiter to schedule all requests and data packets to optimize the behavior of data packets across the network. The advantage is that packet-level scheduling is performed, and the disadvantage is that a central node is required, lacking flexibility and scalability. The distributed type performs resource allocation across the network through an end-to-end request-allocation mechanism, avoiding the disadvantages of the centralized type. The distributed solution usually uses a frame of the minimum Ethernet size when allocating bandwidth resources, which is usually called a credit or token. There is a one-to-one correspondence between credit packets and data packets.

[0007] Although the active transmission scheme has advantages such as strong congestion avoidance and the ability to handle the incast phenomenon well, its biggest drawback is that it requires an end-to-end request mechanism in advance and a process of notifying the receiver to start resource allocation. This non-data-sending notification process consumes a basic end-to-end transmission time (Round-Trip Time, RTT). During this period, data is not sent because no credit packet is received. It is considered that bandwidth may be wasted during this RTT waiting time, and from the perspective of latency characteristics, the final latency characteristics cannot be improved compared to the passive scheme, and no advantage can be brought to the data center business from the perspective of the final task effect.

[0008] Some work, such as Aeolus, adopts the method of allowing the transmission of an unscheduled data packet with the quantity of Bandwidth Delay Production (BDP) before the credit arrives, to utilize the RTT idle time of active transmission by occupying the idle bandwidth. However, the existence of this unscheduled packet will affect the effectiveness of the original scheduled packet, reduce the congestion avoidance characteristics of active transmission, and cause performance degradation. To avoid these problems, Aeolus adopts the method of discarding the unscheduled data packet to ensure the scheduled data packet that conforms to the active transmission working mode. This inevitably introduces the detection and retransmission of discarded data packets, increasing the complexity of the protocol. As the bandwidth of the data center network is evolving towards 400Gbps and even 800Gbps, the BDP that grows linearly with bandwidth and latency makes the direct transmission of unscheduled packets face increasingly severe problems, and the effectiveness of this Aeolus scheme is increasingly challenged and more difficult to cope with the increasingly tense switch buffer resources. Among them, Aeolus is a research paper on data center network congestion control published by Alibaba Cloud at the SIGCOMM conference in 2020.

[0009] Under long-distance conditions, the effective bandwidth utilization of long-distance links is a key indicator, and at the same time, the impact caused by the basic long-distance latency introduced by active transmission is greater. Although active transmission can effectively utilize shallow buffer switches to avoid the BDP buffer requirements at the long-distance link level introduced by the PFC mechanism, an effective method to reduce the latency of active transmission is still needed. Summary of the Invention

[0010] Based on the technical problems existing in the background technology, the present invention proposes a long-distance data center active transmission method for accelerating the transmission in the pre-credit stage, which accelerates the credit transmission in the pre-credit stage and minimizes the network deployment cost.

[0011] A long-distance data center active transmission method for accelerating the transmission in the pre-credit stage proposed by the present invention is executed on a cross-data center network, and the method includes:

[0012] The switch set between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding credit packets, and randomly forwards the credit packets to the next hop in the routing table;

[0013] Maintain the connection status table of the local data center at the cross-data center egress switch, and maintain the host granularity expected bandwidth table of the remote data center. The host granularity expected bandwidth table records the number of credit packets from each remote host, and updates at fixed intervals Time update;

[0014] Use the connection status table to judge the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, determine the sender of the time-closest established request under the same ToR switch as the idle credit according to the host granularity expected bandwidth table and the connection status table as the destination host, and forward the idle credit to the destination host to realize the secure acceleration of the pre-credit stage of the transmission using the idle credit.

[0015] Furthermore, there are three stages in the active transmission communication between the sender and the receiver: the pre-credit stage, the credit stage, and the post-credit stage;

[0016] In the pre-credit stage, the sending host sends a credit request to the receiving host until the first credit from the receiving host is received. In the pre-credit stage, the sending host defaults to a waiting state and does not transmit any data packets;

[0017] In the credit stage, the sending host receives the first credit from the receiving host, and upon reaching the credit, triggers the transmission of a packet of one MTU (maximum transmission unit) until all the data corresponding to this connection of the sending host is completely transmitted;

[0018] In the post-credit stage, the sending host sends all the data and sends a credit stop control message, stating that there is no traffic to send on this connection. The receiving host stops sending credits at the latest after receiving the credit stop control message.

[0019] Furthermore, when the switch uses the random packet spraying scheme to forward credit packets, it adheres to a symmetric path based on credit active transmission, and achieves the result of a shallow queue through packet-level symmetric path constraints.

[0020] Furthermore, a symmetric path is implemented based on the explicit path marking method. The explicit path marking method is specifically:

[0021] Explicitly record the forwarding path of the switch for the credit in the credit packet header;

[0022] In the explicit path marking strategy, when the switch receives a credit packet, it determines the ingress port number where the credit packet is received, adds the ingress port number to the explicit path marking field in the credit packet header, and increments the hop counter by one. The counter records the number of forwarding times the credit packet has experienced.

[0023] After receiving the credit packet, the sender copies the explicit path marking field into the data packet, and the data packet reaches the receiver hop by hop along the symmetric path according to the ingress port number recorded in the explicit path marking field.

[0024] Furthermore, the host granularity expected bandwidth table constraint The number of credits for each host within a certain time is less than , where is the uplink bandwidth of the host and MTU is the maximum transmission unit.

[0025] Furthermore, in the process of using the connection state table to judge the state of the credit packet received by the egress switch from the remote data center, two types of states of the connection are identified from the three stages of active transmission communication: the sending state and the ending state;

[0026] The sending state represents the sendable characteristic of the connection sender, and it can send when it receives a credit;

[0027] The ending state means that the sending of this connection is completed, declaring the idle characteristic of subsequent credits.

[0028] Furthermore, the connection state table includes:

[0029] Basic tuple information of the connection: source IP address, destination IP address, source port, destination port;

[0030] Connection state flag , and the sending state and the ending state have opposite bit positions;

[0031] Traffic size , which is used for the volume of this traffic, in bytes;

[0032] Establishment request time , which records the timestamp at the time of the establishment request, indicating that the latest connection needs to be accelerated first.

[0033] Furthermore, according to the host granularity expected bandwidth table and the connection state table, determine the sender of the connection with the largest acceleration benefit coefficient under the same ToR switch as the idle credit, that is, in the transmission of the idle credit security acceleration pre-credit stage, use the acceleration benefit coefficient to measure the effectiveness:

[0034] ;

[0035] Among them, is the current moment during the pre-credit stage transmission process, is the establishment request time, is the end-to-end transmission time, is the host bandwidth, is the traffic size.

[0036] Furthermore, during the process of accelerating the pre-credit stage transmission by using idle credit security, the maximum data queue of the downlink ports under the same ToR switch is:

[0037] ;

[0038] Among them, is the long-distance link bandwidth, is the number of hosts under the same ToR switch.

[0039] Furthermore, when there is no queue risk;

[0040] When , that is, when the long-distance link bandwidth is greater than the host bandwidth, there is a risk of queue establishment, and it is constrained by the expected bandwidth table of the host granularity:

[0041] ;

[0042] Among them, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the expected bandwidth table of the host granularity, is the source IP address of the host.

[0043] The advantage of the active transmission method for long-distance data centers that accelerates the pre-credit stage transmission provided by the present invention is as follows: It realizes active transmission that completely relies on credit scheduling, so there is no need to introduce a loss detection and retransmission mechanism caused by active discarding. Through the explicit path marking function supported by ordinary switches and the maintenance of the expected bandwidth table of host granularity and some connection state tables supported by the egress switch, the credit transmission in the pre-credit stage is accelerated. This implementation utilizes the characteristics supported by ordinary switches, and the table entry maintenance function only requires the support of a minimum number of egress switches, minimizing the network deployment cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic flowchart of the present invention;

[0045] Figure 2 is a schematic diagram of the three stages of active transmission;

[0046] Figure 3 Schematic diagram of the working process of credit random spraying

[0047] Figure 4 Schematic diagram of the principle of explicit path marking for long - distance data centers

[0048] Figure 5 Schematic diagram of the message structures of credit packets and data packets; where (a) is the schematic diagram of the message structure of credit packets and (b) is the schematic diagram of the message structure of data packets

[0049] Figure 6 Schematic diagram of the architecture of the egress switch for accelerating the transmission in the pre - credit stage

[0050] Figure 7 Topology diagram and connection schematic diagram of a long - distance active transmission Detailed implementation manners

[0051] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0052] As Figures 1 to 7 shown, an active transmission method for a long - distance data center that accelerates the pre - credit stage transmission proposed by the present invention is executed on a cross - data - center network, and the method includes:

[0053] Step 1: The switch set between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding credit packets, and randomly forwards the credit packets to the next hop in the routing table.

[0054] Step 2: The egress switch for the cross - data - center maintains a connection status table of the local data center and a host - granularity expected bandwidth table of the remote data center. The host - granularity expected bandwidth table records the number of credit packets from each remote host and is updated at fixed - interval time.

[0055] Step 3: Use the connection status table to determine the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, determine the sender of the time - closest established request under the same ToR switch as the idle credit according to the host - granularity expected bandwidth table and the connection status table as the destination host, and forward the idle credit to the destination host to achieve secure acceleration of the pre - credit stage transmission using the idle credit.

[0056] Active transmission can avoid congestion, enabling long-distance interconnection to achieve data center interconnection with only general shallow buffer switches and enabling a lossless network with zero packet loss. Compared with passive solutions, it can greatly reduce the cost of deep buffer switching devices. However, due to the basic latency characteristics of active transmission, the final latency performance is poor, resulting in degraded cross-data center application performance. Against this background, the purpose of this embodiment is: based on the existing credit-based active transmission mechanism between data centers, to propose a method for accelerating the transmission in the pre-credit stage, effectively reducing the basic latency of cross-data center active transmission without affecting the original active allocation mechanism, and thus without the need to introduce a loss detection and retransmission mechanism. This method is easy to deploy, has minimal requirements for network-side switches, only adds some packet headers, and does not involve complex processing; only the minimum number of long-distance egress switches are required to participate in the specific function implementation, and only the minimum number of device replacements are needed. Through the connection status table recorded by the egress switch, the transmission of the recently created connection is safely accelerated using idle credits. Since the transmission is carried out using existing idle credits, there are no unscheduled packets and no problem of increased interference bandwidth deteriorating the advantages of active transmission due to the increase in the switch queue.

[0057] In the scenario of long-distance data center interconnection in this embodiment, in order to make full use of the strong congestion avoidance and shallow queue characteristics of active transmission, this embodiment is based on the data center active transmission scheme. The connection status table of the local data center and the host granularity expected bandwidth table of the remote data center are maintained in the cross-data center egress switch, which plays a key role in reducing the basic latency of the active transmission mechanism: when the egress switch detects the idle credit indicated by the connection status table, it detects the host connection with idle bandwidth indicated by the expected bandwidth table, and makes the necessary fields of the packet header of the idle credit packet modified so that it is finally forwarded to the available host to trigger packet transmission. Through the operations of the egress switch based on these states, the pre-credit acceleration effect of active transmission on the long-distance link is finally obtained.

[0058] (1) In one of the embodiments, this embodiment is based on credit-based active transmission, specifically:

[0059] The working paradigm of active transmission determines that it has the disadvantage of basic end-to-end delay, but in exchange for advantages that are difficult to achieve simultaneously by other passive solutions: when the sender initiates a connection, it first sends a connection establishment request to the receiver, carrying necessary connection information, and the frame size of the request is the minimum Ethernet frame size; after receiving the request, the receiver starts to send credit packets in the reverse direction, also set to the minimum Ethernet frame size. Each port of the switch requires two queue overheads, a credit queue and a data queue, and the whole network limits the speed of the credit queue, with an interval of the maximum Ethernet size (Maximum Transmission Unit, MTU) transmission time. From the perspective of the sender, there are three stages in active transmission communication: the pre-credit stage, the credit stage, and the post-credit stage:

[0060] In the pre-credit stage, the sending host sends a credit request to the receiving host. Until the first credit from the receiving host is received, in the pre-credit stage, the sending host is in a waiting state by default and does not transmit any data packets;

[0061] In the credit stage, the sending host receives the first credit from the receiving host. When the credit arrives, it triggers the transmission of a data packet of one MTU (Maximum Transmission Unit). This stage ends when all the data corresponding to this connection of the sending host has been transmitted;

[0062] In the post-credit stage, the sending host has sent all the data and sends a credit stop control message, stating that there is no traffic to send on this connection. The receiving host stops sending credits at the latest after receiving the credit stop control message.

[0063] The most significant impact of long-distance data center interconnection on transmission is the huge increase in delay (5 us / km) brought by the long-distance link. This first brings a millisecond-level basic delay to active transmission, masking the advantages of active transmission under long-distance conditions. At the same time, the long delay also brings problems such as a large number of idle credits and difficult control in the post-credit stage. However, the existence of idle credits brings a turning point for solving the basic high delay.

[0064] This embodiment is described based on the per-hop credit design that follows the credit mechanism for active congestion control. The three stages of active transmission are as Figure 2 shown. The sender first sends a request or continuous simulation packet to the receiver, assuming the size information of the unknown flow, which is more practical and makes this embodiment have better scalability.

[0065] Then, after the receiver receives the request, it returns credits to the sender or returns credits for each simulated packet received. Both the request packet and the credit packet are the smallest Ethernet frames of MinTU bytes, and the data packet is the largest Ethernet frame of MTU bytes. The host network card and the switch limit the speed of the credit packet, so that the forwarding interval of the credit packet is the transmission time of 1MTU, so as to constrain the triggered data packet queue. Each switch port contains a credit queue and a data queue, and the credit queue has a shallow threshold , when the threshold is exceeded, the credit packet is randomly discarded. The credit packets sharing the bottleneck can fairly obtain the bandwidth share. When the sender receives the credit packet, it immediately sends a data packet to the receiver. In the ideal case (ultra-large flow, each credit triggers a data packet of 1MTU size), the additional overhead of active control is:

[0066] .

[0067] (2) In one embodiment, the credit packet forwarding mechanism with explicit path marking;

[0068] To accelerate the transmission in the pre-credit stage of long-distance active transmission, first, a forwarding strategy with explicit path marking is adopted. Active transmission based on credits requires the data packets triggered by reverse credits to follow the same forwarding path in order to achieve advantages such as strong congestion avoidance. In the network, symmetric forwarding paths are easy to implement and have no impact on performance. To address the potential core load imbalance caused by hash load balancing and to fully reduce the basic latency of active transmission, a forwarding strategy called explicit path marking is adopted.

[0069] When the switch forwards the credit packet, it uses the random packet spraying scheme. The credit is randomly sprayed as Figure 3 ,

[0070] The switch randomly forwards the credit packet to the next hop in the routing table. The advantage of random packet spraying is that it has almost the best core load balancing effect. And because the equal-cost multi-path mechanism is actually redundant with explicit path marking, it is also necessary to introduce the load balancing of multi-path. The out-of-order rearrangement problem involved in random packet spraying itself is minimized by the queuing and discarding characteristics of credits, and only the smallest data buffer rearrangement overhead of the receiver is required, which is much smaller than the out-of-order overhead in the passive scheme.

[0071] The basis for using random packet spraying is that the credit active transmission still follows the requirement of symmetric paths. By constraining the symmetric paths at the packet level, the effect of a shallow queue is also achieved. To achieve symmetric paths and subsequent designs, the implementation scheme of explicit path marking is: as Figure 4As shown, all switches in the long-distance interconnection topology of the data center need to perform random packet spraying of credits and explicit path marking. The core of this operation is to explicitly record the forwarding path of the switch for the credit in the credit packet header for the forwarded credit packet, making up for the missing path information due to random credit packet spraying by marking the path of each credit packet. To obtain the reverse data symmetric path, in the explicit path marking strategy, when a switch in the network receives a credit packet, it needs to determine the incoming port number of the received credit packet (directly obtained from the data plane, which can be supported by almost all modern commercial switches), add this incoming port number to the explicit path marking field in the credit packet header, and increment the hop counter by one. This counter records the number of forwarding times the credit packet has experienced. For example, in the Figure 4 credit packet sent by the receiving direction sender in the attached figure, it enters the same TOR switch through link ① with sequence number, and the forwarding path of the TOR switch for the credit is explicitly recorded in the credit packet header, corresponding to the credit packet with sequence number ②. Then the credit packet enters the egress switch 2 through link ③, and the forwarding path of the egress switch 2 for the credit is explicitly recorded in the credit packet header, corresponding to the credit packet with sequence number ④. Subsequently, the credit packet enters the egress switch 1 through the long-distance link ⑤, and the forwarding path of the egress switch 1 for the credit is explicitly recorded in the credit packet header, corresponding to the credit packet with sequence number ⑥. The credit packet enters the same TOR switch through link ⑦, and the forwarding path of the same TOR switch for the credit is explicitly recorded in the credit packet header, corresponding to the credit packet with sequence number ⑧. Finally, the sender receives the credit packet with sequence number ⑧.

[0072] The message structures of the credit packet and the data packet are as Figure 5 . After receiving the credit packet, the sender copies the explicit path marking field into the data packet in the form of a replicated field with sequence number ⑨, obtaining the data packet with the explicit path marking field carried in the header corresponding to sequence number ⑩. When the switch forwards the data packet, it reads the recorded port number according to the hop counter and forwards it from the specified port, and decrements the hop counter by one. Thus, the data packet reaches the receiver hop by hop along the symmetric path according to the incoming port number recorded in the field. This function involves write and read operations on the message headers of all switches in the cross-data center environment and is easy to implement. The explicit path mechanism only utilizes the advantages of multi-path. To reduce the basic latency of the active transmission flow, the following key mechanisms are also required: the connection status table and the host granularity expected bandwidth table.

[0073] (3) In one embodiment, the maintenance of the connection status table and the host granularity expected bandwidth table;

[0074] To effectively utilize the idle credit and reduce the base latency of active transmission, the egress switch maintains a Host Throughput Table (HTT) for the hosts in the remote data center at the host granularity. Since modern data centers are oversubscribed to varying degrees, the core bandwidth is usually less than the uplink bandwidth of all server nodes, so that the sending hosts under the active transmission scheduling feature will not all be fully occupied simultaneously. The passive scheme causes the oversubscribed core to have traffic exceeding its carrying capacity due to the uncontrolled line-speed startup, resulting in queue accumulation and buffering of excessive traffic, which is also the main congestion problem that the passive transmission attempts to solve; while the active transmission, due to the characteristics of consuming time for bandwidth allocation and the advantage of a limited shallow data queue, ideally makes the network core of the active transmission be fully utilized. However, due to the long latency caused by long distances, the link may be in a low-utilization state, which is due to the discardable nature of the credit and the post-credit stage.

[0075] Therefore, there may be idle bandwidth in long-distance active transmission, making it possible to accelerate the transmission in the pre-credit stage of active transmission. Compared with the heuristic method of Aeolus for occupying idle bandwidth, the key point of this embodiment is to accurately discover the idle host bandwidth through the host granularity expected bandwidth table to avoid conflicts between occupying bandwidth in the core and the bandwidth allocation of the credit, and use the idle credit causing the idle bandwidth to accelerate the transmission in the pre-credit stage. To effectively judge the idle bandwidth, the egress switch maintains the expected bandwidth at the host granularity of the remote data center and maintains the host granularity expected bandwidth table of the remote hosts according to the IP address of the credit sent by the egress switch to the local. The host granularity expected bandwidth table of the egress switch records the number of credit packets belonging to each remote host and updates at a fixed interval time. Because of the update at a fixed time interval, there is no need to perform division operations, only simple counting operations are required, avoiding the overhead of the switch doing division. The host granularity expected bandwidth table restricts the number of credits of each host within a time to be less than , where is the uplink bandwidth of the host and MTU is the maximum transmission unit.

[0076] The egress switch manages the cross-data center connections from the local data center and maintains a connection status table as shown in Table 1:

[0077] Table 1

[0078]

[0079] The maintenance information of the status table can be divided into the following three segments:

[0080] Basic tuple information of the connection: source IP address ( ) and destination IP address ( ), source port ( ), destination port ( );

[0081] Connection status flag , where the send state and the end state have opposite bit positions;

[0082] Traffic volume , which is used for the volume of this traffic, in bytes;

[0083] Connection request time , recording the timestamp when the connection request is established, indicating that the latest connection needs to be accelerated first.

[0084] Therefore, compared with a general switch, the egress switch in this embodiment also needs to perform two additional operations. One is to maintain the throughput at the host granularity of the remote data center (i.e., maintain the expected bandwidth table at the host granularity). The egress switch counts the bandwidth of the host corresponding to the IP address in the header of the credit packet sent to the local data center (i.e., maintain the connection status table), which requires a fixed amount of resource overhead of the egress switch and can be performed in the control plane. The egress switch counts the number of credit packets of each host (IP address) at fixed time intervals. The expected bandwidth table at the host granularity is actually a counter because of the time update characteristic, without the need for specific division operations when calculating the bandwidth, reducing the calculation overhead of the switch.

[0085] (4) The egress switch accelerates the transmission in the pre-credit stage;

[0086] According to the aforementioned maintenance of the expected throughput table at the host granularity, when the egress switch has observed a host with the ability to receive packets with idle bandwidth, it can receive data packets. And because the network core is occupied by credit packets, starting from the idle credit, the network core bandwidth can be scheduled by the explicit path constraint-based scheduling, and the transmission of the flow in the pre-credit stage can be safely accelerated. Due to the bandwidth scheduling of the network core for the idle credit, it can ensure that the core bandwidth will not be overused. At the same time, due to the oversubscription characteristic of the topology of modern data centers, the uplink of the host can be effectively accelerated. Considering such a typical scenario, a certain number of hosts are connected under a Top of Rack (ToR) switch. When these hosts are receivers, the credits sent under the same ToR switch are forwarded through the core layer of the long-distance data center to the egress switch of the local data center. There are a relatively large number of hosts connected under the same ToR switch. If the idle credit is changed to the connection of other hosts under the same ToR switch, since the implementation of the path label does not depend on the common five-tuple information-based hash load balancing in the data center, it can ensure that the paths in the core layer remain consistent after changing the IP and port.

[0087] For hosts under the same ToR switch with inconsistent first hops, the host granularity expected bandwidth table is used for processing. If there is idle bandwidth, credit can be safely converted. Since the existing path information is retained, the paths from different hosts to the same ToR switch can also safely perform credit conversion because the egress switch detects idle bandwidth.

[0088] To enable the egress switch to have knowledge of the connection level in the local data center, a connection status table needs to be maintained in the egress switch to record the connection status initiated by the local switch, including basic connection tuple information (source IP address, destination IP address, source port number, destination port number, etc.), request establishment time, and status flag bits. There are two types of states that identify connections in the three stages of active transmission: the sending state and the ending state. The sending state indicates the sendable characteristics of the connection sender, and it can send upon receiving credit; the ending state indicates that the connection has been sent and declares the idle characteristics of subsequent credit.

[0089] For the credit received by the egress switch from the remote end, after checking the connection status table of the egress switch and finding that it is in the ending state, the credit is forwarded to the sender that initiated the request most recently under the same ToR switch to which the host belongs. Under the best conditions, the basic latency of active transmission can be reduced from millisecond-level long-distance end-to-end latency to microsecond-level latency within two or three hops in the data center, greatly reducing the basic latency problem.

[0090] Specifically: Figure 6 This is the architecture diagram of the egress switch. According to the connection status table of the egress switch, the process of safely accelerating the pre-credit stage of transmission using idle credit is as follows:

[0091] For the credit packet received by the egress switch from the remote data center, query the connection status table. If the current connection is already in the ending state, it indicates that this credit is idle credit and idle bandwidth has been allocated. If it is not utilized, it will damage the core bandwidth across data centers. Then, based on the host granularity expected bandwidth table and the connection recorded in the connection status table with the destination host under the same ToR switch to which the idle credit belongs, for the destination host with idle bandwidth, select the connection with the most recent establishment time sent to the idle host, and obtain the necessary IP and port information of this connection. The egress switch rewrites the IP and port information of the credit, converts it from idle credit to credit for a connection in the sending state, and retains the existing explicit path marking header. Subsequent forwarding continues with explicit path marking without affecting the implementation of the symmetric path.

[0092] In step three, according to the host granularity expected bandwidth table and the connection status table, determine the acceleration benefit coefficient under the same ToR switch to which the idle credit belongs The largest connected sender is used as the destination host. That is, in the transmission of the pre-credit stage using idle credit for secure acceleration, to clarify the effectiveness of accelerating the transmission in the pre-credit stage, an acceleration benefit coefficient is used to measure as follows:

[0093] ; (1)

[0094] Among them, is the current moment during the transmission in the pre-credit stage, is the request establishment time, is the end-to-end transmission time, is the host bandwidth, is the traffic size.

[0095] This embodiment adds constraints: , such that the calculation range of is from 0 to 50%. Among them, the larger the value, the more effective the acceleration of the transmission in the pre-credit stage. The core objective of formula (1) is to quantify the effect of early transmission in the pre-credit stage on reducing latency. The on the right side of the equal sign represents the total time required to complete credit allocation and transmission in the traditional protocol (two RTTs + transmission time); the on the right side of the equal sign represents the estimated transmission time for accelerating this connection (including the time after request establishment, RTT, and theoretical transmission time). If is close to 50%, it means that the secure acceleration pre-credit stage in this embodiment almost completely avoids the waiting time in the traditional protocol. Therefore, to maximize , traverse the connection status flags in the connection status table that are in the sending state, and select the connection that makes the largest. The parameters in formula (1) jointly characterize the transmission efficiency in the pre-credit stage. By maximizing the acceleration benefit coefficient , the latency of short flows can be significantly reduced, and the overall performance of the data center network can be improved.

[0096] Formula (1) approximately calculates the reduction ratio of latency brought by the early transmission of traffic in the pre-credit stage. To accelerate the transmission in the pre-credit stage, it is necessary to select the corresponding request that maximizes the acceleration benefit coefficient . At the same time, since the idle credit under the same ToR switch may be used to accelerate the traffic of a certain host, which may lead to possible bandwidth overage and queue establishment, the maximum data queue of the downstream link port under the same ToR switch is:

[0097] ; (2)

[0098] Among them, is the long - distance link bandwidth, is the number of hosts under the same ToR switch.

[0099] When there is no queue risk;

[0100] But in a more general scenario in the scenario where the long - distance link bandwidth is greater than the host bandwidth, there is a large queue establishment, and it is necessary to use the aforementioned host - granularity expected bandwidth table for constraint:

[0101] ; (3)

[0102] Among them, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the host - granularity expected bandwidth table, is the source IP address of the host.

[0103] In formula (3), the core goal is to ensure that the transmission of the acceleration flow does not cause additional queue accumulation. The optimization goal of formula (3) can be divided into two levels. (a1) Maximize the acceleration benefit coefficient (that is, ) : That is, maximize the reduction ratio of the transmission in the pre - credit phase of the traffic to the delay, which will reduce the waiting time in the credit negotiation phase and maximize the reduction of the overall delay. (a2) Meet the bandwidth constraint of the destination host : Control the traffic bandwidth received by the host within time not to exceed the capacity of the link, avoid additional queue establishment, so that the transmission in the acceleration pre - credit phase will not bring additional queue accumulation risk and increase the queuing delay.

[0104] By imposing the condition that the host bandwidth does not exceed the limit and then maximizing the delay acceleration benefit, this makes the transmission in the acceleration pre - credit phase not cause additional queue establishment, that is, perform safe pre - credit phase acceleration.

[0105] In this embodiment, as Figure 7As shown, there are 4 hosts under a ToR switch, and each host has a connection F1, F2, F3, F4 corresponding to the connections from the sender S to the receivers R with corresponding numbers respectively. That is, F1 is sent from the sender S1 to the receiver R1, F2 is sent from the sender S2 to the receiver R2, F3 is sent from the sender S3 to the receiver R3, and F4 is sent from the sender S4 to the receiver R4. Now, a connection F4 ends, and a new connection F5 is added from the sender S4 to the receiver R4. At this time, the connection status table of the egress switch records the entries of F1~F5, and the status bit of F4 indicates that F4 is in the end state. The egress switch sets the credit in the post-credit stage received from F4 to idle credit. This embodiment uses the acceleration mechanism of the idle credit security to accelerate the pre-credit stage, and rewrites the key fields (IP and port information of the credit) in the idle credit header as the connection information corresponding to F5, because according to the calculation of the acceleration benefit coefficient , F5 should be accelerated the most, and the destination hosts belong to the same remote ToR switch. By this method of using idle credit security to accelerate the active transmission flow, the basic delay of the active transmission is reduced, and the conflict problem of unscheduled packets and credit-scheduled packets is not introduced, greatly improving the superiority of the credit-based active transmission in the long-distance scenario. This embodiment has the following advantages:

[0106] Optimizes the delay problem of active transmission and avoids relying on unscheduled data packets to reduce interference with the active transmission bandwidth. This method is completely implemented based on credit scheduling and does not require the introduction of loss detection and retransmission mechanisms caused by active packet discarding. Through the explicit path marking function supported by ordinary switches and the host granularity expected bandwidth table maintenance and some connection status tables supported by the egress switch, the credit transmission in the pre-credit stage is accelerated. This embodiment utilizes the characteristics supported by ordinary switches, and the table entry maintenance function only requires the support of a minimum number of egress switches, minimizing the network deployment cost.

[0107] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An active transmission method for long-distance data centers that accelerates the transmission in the pre-credit stage, characterized in that, The method is executed on a cross-data center network, and the method includes: The active transfer communication between the sender and the receiver includes a pre-credit phase, in which the sending host sends a credit request to the receiving host. Until the first credit from the receiving host is received, in the pre-credit phase, the sending host is by default in a waiting state and does not transmit any data packets. The switch set between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding credit packets, and randomly forwards the credit packets to the next hop in the routing table. Maintain the connection status table of the local data center in the cross-data center egress switch, and maintain the host granularity expected bandwidth table of the remote data center. The host granularity expected bandwidth table records the number of credit packets belonging to each remote host and is updated at fixed intervals Time update; Use the connection state table to judge the state of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, determine the sender of the request established most recently in time under the same ToR switch as the idle credit according to the host-granularity expected bandwidth table and the connection state table as the destination host, and forward the idle credit to the destination host to realize the secure acceleration of the transmission in the pre-credit phase by using the idle credit.

2. The long-distance data center active transmission method for accelerating the pre-credit stage transmission according to claim 1, wherein, There are three phases in the active transfer communication between the sender and the receiver: the pre-credit phase, the credit phase, and the post-credit phase; The credit phase is that the sending host receives the first credit from the receiving host, and upon reaching the credit, a data packet of one MTU is sent. Until all the data corresponding to this connection of the sending host is completely transmitted, the MTU is the maximum transmission unit; The post-credit phase is that the sending host sends all the data and sends a credit stop control message, declaring that there is no traffic to send on this connection, and the receiving host stops sending credits at the latest after receiving the credit stop control message.

3. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit phase according to claim 1, characterized in that, When the switch uses the random packet spraying scheme to forward credit packets, it complies with the symmetric path based on credit active transfer, and achieves the result of a shallow queue through packet-level symmetric path constraints.

4. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit stage according to claim 3, wherein The symmetric path is implemented based on the explicit path marking method. The explicit path marking method is specifically: Explicitly record the forwarding path of the switch for the credit in the credit packet header; In the explicit path marking policy, when the switch receives a credit packet, it determines the ingress port number of the received credit packet, adds the ingress port number to the explicit path marking field of the credit packet header, and at the same time increments the hop counter by one. The counter records the number of forwarding times experienced by the credit packet; After receiving the credit packet, the sender copies the explicit path marking field to the data packet, and the data packet reaches the receiver hop by hop along the symmetric path according to the ingress port number recorded in the explicit path marking field.

5. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1, wherein The host granularity expected bandwidth table constraint The number of credits for each host within a time period is less than , where is the uplink bandwidth of the host, and MTU is the maximum transmission unit.

6. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit stage according to claim 2, characterized in that In judging the state of the credit packet received by the egress switch from the remote data center by using the connection state table, two types of states of the connection are identified from the three phases of the active transfer communication: the sending state and the ending state; The sending state represents the sendable characteristic of the connection sender, and it can send upon receiving a credit; The ending state represents that the transmission of this connection is completed, declaring the idle characteristic of subsequent credits.

7. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit phase according to claim 1, wherein The connection state table includes: Basic tuple information of the connection: source IP address, destination IP address, source port, destination port; Connection status flag bit , the sending state and the end state have opposite bit positions; Traffic volume , which is the volume of this traffic in bytes; Establish request time , record the timestamp when the request is established, indicating that the latest connection needs to be accelerated first.

8. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit stage according to claim 1, characterized in that, Determine the sender of the connection with the largest acceleration benefit coefficient under the same ToR switch as the idle credit according to the expected bandwidth table and connection status table of the host granularity, that is, in the transmission of the accelerated pre-credit stage using the idle credit, use the acceleration benefit coefficient to measure the effectiveness: Measure effectiveness Among them, is the current moment during the pre-credit stage transmission process, is the establishment request time, is the end-to-end transmission time, is the host bandwidth, is the traffic size.

9. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 8, characterized in that, During the transmission process of accelerating pre-credit stage using idle credit security, the maximum data queue of the downlink port under the same ToR switch is as follows: Among them, is the long-distance link bandwidth, is the number of hosts under the same ToR switch.

10. The active transmission method for long-distance data centers that accelerates the transmission in the pre-credit stage according to claim 9, characterized in that, When there is no queue risk; When , that is, when the long-distance link bandwidth is greater than the host bandwidth, there is a risk of queue establishment, and it is constrained by using the expected bandwidth table at the host granularity: wherein, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the expected bandwidth table at the host granularity, is the source IP address of the host.

Citation Information

Patent Citations

  • Credit packet-based active transmission method for data center

    CN112437019A

  • Data stream transmission scheduling control method and system in cloud data center application

    CN115022249A