Long-distance data center active transmission method for accelerating transmission in pre-credit stage
By adopting random packet spraying and explicit path marking technology in the data center network, combined with the use of host granularity expected bandwidth tables, the problems of high latency and low bandwidth utilization between long-distance data centers are solved, achieving more efficient transmission performance and lower network deployment costs.
Patent Information
- Application Number
- CN202510499784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
When existing active transmission solutions transmit between long-distance data centers, the delay is high and the idle bandwidth cannot be effectively utilized, resulting in waste and performance degradation in the pre-credit stage.
By using random packet spraying and explicit path marking techniques on the switch, the forwarding path of credit packets is optimized and the host granularity expected bandwidth table is maintained on the egress switch, using idle credit to accelerate transmission of the pre-credit phase.
Reduced latency and bandwidth utilization during long-distance data centers are achieved, avoiding performance degradation caused by unscheduled packets and reducing network deployment costs.
Smart Images

Figure CN120050242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network transmission control, and in particular to a long-distance data center active transmission method for accelerating pre-credit stage transmission. Background Art
[0002] In the past decade, with the continuous development of business, data centers have become key infrastructure supporting the digital economy. Computing, storage, and networking are the three elements of data centers, and they promote each other's common development: in the past decade, computing power has increased hundreds of times, and the access performance of solid-state drives has increased a hundred times compared to mechanical hard drives. The development of computing and storage has put forward new requirements for networks, requiring greater bandwidth and lower latency. More and more different types of traditional services and big data processing, high-performance computing, and artificial intelligence tasks are running on data centers, requiring more efficient and reliable data centers. For data center network transmission, the TCP / IP system of traditional wide area networks is not used in modern data centers because it is designed for general complex and highly dynamic environments and requires high CPU overhead. A transmission method that is more suitable for data center business needs and hardware facilities is needed.
[0003] At the same time, due to the energy, scalability and disaster recovery considerations of data center networking, a single data center cannot be expanded infinitely, and geographically separated data centers need to be considered. Traditional geographically separated cloud data centers are interconnected via a wide area network to carry traditional services such as mail services and web services. The performance requirements of the network are not high, and the use of wide area network links from network service providers can meet the requirements, but there are problems such as protocol conversion and heterogeneous network environments.
[0004] However, today's artificial intelligence, such as the ultra-large-scale cluster networking training required for large language models, requires higher network requirements and reliable latency bandwidth conditions. More and more data center operators have begun to build dedicated cross-data center optical fibers to connect data centers. This shift from wide area network links to dedicated links gives greater freedom in the design of cross-data center transmission, allowing for transmission solutions designed specifically for long-distance interconnection between data centers.
[0005] At present, the data center transmission scheme mainly adopted by the industry is passive, relying on passive congestion control to control traffic and avoid high network latency. Priority-based Flow Control (PFC) technology is also needed to provide a lossless network with zero packet loss. The basic design intention of the passive transmission scheme is to achieve low latency. In order to enable the data center mouse flow to complete the transmission as quickly as possible, the method of sending directly at the maximum rate of line speed is adopted when sending. According to the network status caused by the traffic injected into the network, the congestion signal is returned to the sender through different congestion signals. The sender then passively performs congestion control to adjust the sending rate (or window) according to the received congestion signal. Typical congestion control schemes such as DCQCN use Explicit Congestion Notification (ECN), Swift (congestion control mechanism for data center-level networks) uses delay, and HPCC (congestion control protocol) uses more advanced in-band network telemetry signals. These passive methods require feedback delays in the end-to-end round-trip time and take longer to converge, which poses a great challenge to the switch queue. Therefore, PFC hop-by-hop flow control technology is usually used to design data center lossless networks. However, PFC will bring problems such as deadlock and head-of-line blocking. Among them, DCQCN (Data Center Quantized Congestion Notification) is a lossless data transmission congestion control algorithm applied to RoCEv2 networks. RoCEv2 (RDMA over Converged Ethernet version 2) is a high-performance network protocol based on Ethernet that implements low-latency, high-throughput remote direct memory access functions through lossless Ethernet.
[0006] Compared with the traditional passive transmission scheme, active transmission works in a different working paradigm such as "request-allocation". Compared with the "try-backoff" method of the passive scheme, active transmission has a natural advantage in avoiding congestion due to its characteristic of allocating bandwidth resources in advance. By reasonably scheduling and allocating bandwidth resources in advance, the shallow queues in the network can be controlled, and the requirements of low latency and high bandwidth can be met at the same time. Active transmission is divided into two categories: centralized and distributed. The centralized scheme uses a central arbitrator to schedule all requests and data packets, optimizing the behavior of data packets in the entire network. The advantage is that packet-level scheduling is performed, and the disadvantage is that a central node is required, lacking flexibility and scalability; the distributed scheme allocates resources for the entire network through an end-to-end request-allocation mechanism, avoiding the disadvantages of the centralized scheme. The distributed scheme usually uses a minimum Ethernet size frame when allocating bandwidth resources, usually called a credit or token, and the credit packet and the data packet are in a one-to-one correspondence.
[0007] Although the active transmission solution has the advantages of strong congestion avoidance and good handling of concurrency (incast), its biggest disadvantage is that it requires an end-to-end request mechanism in advance and needs to notify the receiver to start the resource allocation process. This non-data transmission notification process consumes a basic end-to-end transmission time (Round-Trip Time, RTT). During this period of time, no data is sent because no credit packet is received. It is believed that bandwidth may be wasted during the RTT waiting time, and from the perspective of delay characteristics, the final delay characteristics cannot be improved compared to the passive solution. From the perspective of the final task effect, it cannot bring advantages to the data center business.
[0008] Some work, such as Aeolus, adopts the method of allowing the transmission of the number of unscheduled packets of the bandwidth delay product (BDP) before waiting for the arrival of credits, so as to utilize the RTT idle time of active transmission by occupying idle bandwidth. However, the existence of such unscheduled packets will affect the effectiveness of the original scheduled packets, reduce the congestion avoidance characteristics of active transmission, and cause performance degradation. In order to avoid these problems, Aeolus adopts the method of discarding unscheduled packets to ensure the scheduled packets that meet the active transmission working mode, which inevitably introduces the detection and retransmission of discarded packets, increasing the complexity of the protocol. As the bandwidth of data center networks is evolving to 400Gbps or even 800Gbps, the BDP that grows linearly with bandwidth and delay makes the direct transmission of unscheduled packets face more and more severe problems. The effectiveness of Aeolus's solution is increasingly challenged, and it is more difficult to cope with the increasingly tight switch buffer resources. Among them, Aeolus is a research paper on data center network congestion control published by Alibaba Cloud at the 2020 SIGCOMM conference.
[0009] Under long-distance conditions, effective bandwidth utilization of long-distance links is a key indicator, and the impact of the basic long-distance delay introduced by active transmission is also greater. Although active transmission can effectively utilize shallow buffer switches to avoid the long-distance link-level BDP buffer requirements introduced by the PFC mechanism, effective methods to reduce the delay of active transmission are still needed. Summary of the invention
[0010] Based on the technical problems existing in the background technology, the present invention proposes a long-distance data center active transmission method for accelerating pre-credit stage transmission, accelerating pre-credit stage credit transmission, and minimizing network deployment costs.
[0011] The present invention proposes a long-distance data center active transmission method for accelerating pre-credit stage transmission, the method is performed on an inter-data center network, and the method includes: The switch set up between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding credit packets, and randomly forwards the credit packets to the next hop in the routing table; The cross-data center egress switch maintains a connection status table of the local data center and a host-granular expected bandwidth table of the remote data center. The host-granular expected bandwidth table records the number of credit packets from each remote host and reports the number of credit packets to the remote host at fixed intervals. Time update; The connection status table is used to determine the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, the sender of the request that was most recently established under the same ToR switch as the idle credit is determined as the destination host based on the host granularity expected bandwidth table and the connection status table, and the idle credit is forwarded to the destination host, thereby utilizing the idle credit to safely accelerate the transmission in the pre-credit stage.
[0012] Further, there are three phases of active transmission communication between the sender and the receiver: pre-credit phase, credit phase, and post-credit phase; The pre-credit stage is when the sending host sends a credit request to the receiving host until the first credit from the receiving host is received. In the pre-credit stage, the sending host is in a waiting state by default and does not transmit any data packets. The credit stage is when the sending host receives the first credit from the receiving host, and the arrival of the credit triggers the sending of a data packet of MTU until all the corresponding data of this connection to the sending host is transmitted. The MTU is the maximum transmission unit; The post-credit stage is when the sending host sends all data and sends a credit stop control message, declaring that there is no traffic to be sent in this connection. The receiving host stops sending credit after receiving the credit stop control message at the latest.
[0013] Furthermore, when the switch uses a random packet spraying scheme when forwarding credit packets, it actively adheres to the symmetric path based on credit transmission and achieves the result of shallow queue through packet-level symmetric path constraints.
[0014] Furthermore, the symmetric path is realized based on the explicit path marking method, and the explicit path marking method is specifically: The switch's forwarding path for the credit is explicitly recorded in the credit packet header; In the explicit path marking strategy, when a switch receives a credit packet, it determines the ingress port number of the received credit packet, adds the ingress port number to the explicit path marking field of the credit packet header, and simultaneously increases the hop count counter by one, which records the number of forwarding times experienced by the credit packet; After receiving the credit packet, the sender copies the explicit path tag field to the data packet, and the data packet reaches the receiver along the symmetric path hop by hop according to the ingress port number recorded in the explicit path tag field.
[0015] Furthermore, the host granularity expected bandwidth table constraints The number of credits per host during the time is less than ,in is the uplink bandwidth of the host, and MTU is the maximum transmission unit.
[0016] Further, in determining the state of the credit packet received by the egress switch from the remote data center using the connection state table, two types of connection states are identified from three stages of active transmission communication: a sending state and an ending state; The sending state indicates that the sender of the connection can send, and can send after receiving credit; The end state indicates that the connection has been sent and declares the idle nature of subsequent credits.
[0017] Furthermore, the connection status table includes: Basic tuple information of the connection: source IP address, destination IP address, source port, destination port; Connection status flag , the sending state and the ending state have opposite bits; Traffic size , used for the volume of this traffic, in bytes; Create request time , records the timestamp when the request is established, indicating that the latest connection needs to be accelerated first.
[0018] Further, the acceleration benefit coefficient of the same ToR switch as the idle credit is determined according to the host granularity expected bandwidth table and the connection status table. The sender of the largest connection is used as the destination host, that is, in the transmission of the pre-credit stage using idle credit security acceleration, the acceleration benefit coefficient is used Measuring effectiveness: ; in, is the current moment in the transmission process of the pre-credit phase, To establish the request time, is the end-to-end transmission time, is the host bandwidth, The flow rate.
[0019] Furthermore, in the process of using idle credit to safely accelerate the pre-credit stage transmission, the maximum data queue of the downlink port under the same ToR switch is for: ; in, is the long-distance link bandwidth, The number of hosts on the same ToR switch.
[0020] Furthermore, when There is no queue risk; when , that is, when the long-distance link bandwidth is greater than the host bandwidth, there is a risk of queue establishment, and the host granularity expected bandwidth table is used for constraint: ; in, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the expected bandwidth table at the host granularity, is the source IP address of the host.
[0021] The advantage of the long-distance data center active transmission method for accelerating pre-credit stage transmission provided by the present invention is that it realizes active transmission completely relying on credit scheduling, so there is no need to introduce loss detection and retransmission mechanism caused by active discarding. The pre-credit stage credit transmission is accelerated by the explicit path marking function supported by ordinary switches and the host granularity expected bandwidth table maintenance and some connection status tables supported by the egress switch. This implementation utilizes the characteristics that ordinary switches can support, and the table maintenance function only requires the support of a minimum number of egress switches, minimizing the network deployment cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 It is a three-stage schematic diagram of active transmission; Figure 3 Schematic diagram of the work of random spraying for credit; Figure 4 This is a schematic diagram of the explicit path marking principle of a long-distance data center; Figure 5 Schematic diagrams of message structures of credit packets and data packets; (a) is a schematic diagram of the message structure of a credit packet, and (b) is a schematic diagram of the message structure of a data packet; Figure 6 Schematic diagram of the egress switch architecture to accelerate pre-credit stage transmission; Figure 7 This is a topology diagram and connection diagram for long-distance active transmission. DETAILED DESCRIPTION
[0023] Below, the technical solution of the present invention is described in detail through specific embodiments. Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific implementation disclosed below.
[0024] like Figures 1 to 7 As shown, the present invention proposes a long-distance data center active transmission method for accelerating pre-credit stage transmission, the method is performed on an inter-data center network, and the method includes: Step 1: The switch set between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding the credit packet, and randomly forwards the credit packet to the next hop in the routing table; Step 2: Maintain the connection status table of the local data center and the host granularity expected bandwidth table of the remote data center at the cross-data center egress switch. The host granularity expected bandwidth table records the number of credit packets from each remote host and reports the number of credit packets at fixed intervals. Time update; Step 3: Use the connection status table to determine the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, determine the sender of the request that was most recently established under the same ToR switch as the idle credit according to the host granularity expected bandwidth table and the connection status table as the destination host, and forward the idle credit to the destination host, thereby utilizing the idle credit to safely accelerate the transmission in the pre-credit stage.
[0025] Active transmission can avoid congestion, so that long-distance interconnection only requires a general shallow buffer switch to achieve data center interconnection, and can achieve a lossless network with zero packet loss. Compared with the passive solution, it can greatly reduce the cost of deep buffer switching equipment. However, due to the basic delay characteristics of active transmission, the final delay performance is poor, which degrades the performance of cross-data center applications. In this context, the purpose of this embodiment is to propose a method for accelerating pre-credit stage transmission based on the existing credit-based active transmission mechanism between data centers, effectively reduce the basic delay of active transmission across data centers, and do not affect the original active allocation mechanism, so there is no need to introduce loss detection and retransmission mechanism. This method is easy to deploy, puts forward the minimum requirements for the network side switch, only adds part of the message header, and does not involve complex processing; only the minimum number of long-distance export switches is required to participate in the specific function implementation, and only the minimum equipment replacement is required. Through the connection status table recorded by the export switch, the idle credit is used to safely accelerate the sending of the recently created connection, because the existing idle credit is used for sending, there is no unscheduled packet and there is no interference bandwidth. The problem of increasing the switch queue causes degradation to the advantage of active transmission.
[0026] In the long-distance data center interconnection scenario, in order to make full use of the strong congestion avoidance and shallow queue characteristics of active transmission, this embodiment is based on the data center active transmission solution, and the cross-data center egress switch maintains the connection status table of the local data center, and maintains the host granularity expected bandwidth table of the remote data center, which plays a key role in reducing the basic delay of the active transmission mechanism: when the egress switch detects the idle credit indicated by the connection status table, it detects the host connection with idle bandwidth indicated by the expected bandwidth table, and modifies the necessary fields of the packet header of the idle credit packet so that it is finally forwarded to the available host to trigger data packet transmission. The pre-credit acceleration effect of active transmission of long-distance links is finally obtained through the operation of the egress switch according to these states.
[0027] (1) In one embodiment, this embodiment is based on active credit transmission, specifically: The working paradigm of active transmission determines that it has the disadvantage of basic end-to-end delay, but in exchange for advantages that are difficult to achieve simultaneously with other passive solutions: when initiating a connection, the sender first initiates a connection establishment request to the receiver, carrying the necessary connection information, and the requested frame size is the minimum Ethernet frame size; after receiving the request, the receiver starts to send credit packets in the reverse direction, which are also set to the minimum Ethernet frame size. Each port of the switch requires two queue overheads, a credit queue and a data queue. The entire network limits the speed of the credit queue, and the interval is the maximum Ethernet size (Maximum Transmission Unit, MTU) transmission time. From the sender's perspective, there are three stages in active transmission communication: pre-credit stage, credit stage, and post-credit stage: The pre-credit stage is when the sending host sends a credit request to the receiving host until the first credit from the receiving host is received. In the pre-credit stage, the sending host is in a waiting state by default and does not transmit any data packets. The credit stage is when the sending host receives the first credit from the receiving host, and the arrival of the credit triggers the sending of a data packet of MTU until all the corresponding data of this connection to the sending host is transmitted. The MTU is the maximum transmission unit; The post-credit stage is when the sending host sends all data and sends a credit stop control message, declaring that there is no traffic to be sent in this connection. The receiving host stops sending credit after receiving the credit stop control message at the latest.
[0028] The most significant impact of long-distance data center interconnection on transmission is the huge increase in latency (5us / km) caused by long-distance links. This first brings millisecond-level basic latency to active transmission, making the advantages of active transmission masked by high latency under long-distance conditions. At the same time, long latency also brings a large number of idle credits in the post-credit stage, which is difficult to control. However, the existence of idle credits has brought a turning point in solving the basic high latency.
[0029] This embodiment is based on the credit active congestion following the hop-by-hop credit design of the credit mechanism. The three stages of active transmission are as follows: Figure 2 As shown, the sending end first sends a request or a continuous simulation packet to the receiving end, assuming the size information of the unknown flow, which is more practical and also makes this embodiment have better scalability.
[0030] After receiving the request, the receiver returns credit to the sender or returns credit for each simulated packet received. The request packet and the credit packet are the minimum Ethernet frame of MinTU bytes, and the data packet is the maximum Ethernet frame of MTU bytes. The host network card and the switch limit the speed of the credit packet so that the forwarding interval of the credit packet is 1MTU transmission time, so as to constrain the triggered data packet queue. Each switch port has a credit queue and a data queue. The credit queue has a shallow threshold. , credit packets exceeding the threshold are randomly discarded, and credit packets sharing the bottleneck can obtain bandwidth shares fairly. The sender immediately sends data packets to the receiver after receiving the credit packet. In the ideal case (super large flow, each credit triggers a data packet of 1MTU size), the active control overhead is: .
[0031] (2) In one embodiment, a credit packet forwarding mechanism with explicit path marking; To accelerate long-distance active transmission in the pre-credit phase, a forwarding strategy of explicit path marking is first adopted. Credit-based active transmission requires that the packets triggered by reverse credit follow the same forwarding path to achieve advantages such as strong congestion avoidance. Symmetric forwarding paths are easy to implement in the network and have no impact on performance. To cope with the potential core load imbalance caused by the use of hash load balancing and to fully reduce the basic latency of active transmission, a forwarding strategy called explicit path marking is adopted.
[0032] The switch uses a random packet spraying scheme when forwarding credit packets. Credit random spraying is as follows Figure 3 , The switch randomly forwards the credit packet to the next hop in the routing table. The benefit of random packet spraying is that it has an almost optimal core load balancing effect. And because there is redundancy between the actual and explicit path labels of the equal-cost multi-path mechanism, it is also necessary to introduce multi-path load balancing. The disorder reordering problem involved in random packet spraying itself is minimized due to the queue drop characteristics of the credit, and only the minimum data buffer reordering overhead is required on the receiving end, which is much smaller than the disorder overhead in the passive solution.
[0033] The basis for using random packet spraying is that credit active transmission still complies with the requirements of symmetric paths, and the effect of shallow queues can also be achieved through packet-level symmetric path constraints. In order to achieve symmetric paths and subsequent design needs, the implementation solution of explicit path marking is: Figure 4 As shown, all switches in the long-distance interconnection topology of the data center need to spray credits randomly and explicitly mark the paths. The core of this operation is to explicitly record the forwarding path of the switch for the credits in the credit packet header for the forwarded credit packets. By marking the path of each credit packet, the path information missing due to random credit packet spraying is compensated. In order to obtain a reverse data symmetric path, in the explicit path marking strategy, when receiving a credit packet, the switch in the network needs to determine the ingress port number of the received credit packet (obtained directly from the data plane, which can be supported by almost all modern commercial switches), and add the ingress port number to the explicit path marking field in the credit packet header, and at the same time increase the hop counter by one. The counter records the number of forwarding times the credit message has experienced, such as the attached Figure 4 In the credit packet sent by the receiving direction to the sender, it enters the same TOR switch through sequence number link ①, and explicitly records the forwarding path of the TOR switch for the credit in the credit packet header, and obtains the credit packet corresponding to sequence number ②. Then the credit packet enters the export switch 2 through link ③, and explicitly records the forwarding path of the export switch 2 for the credit in the credit packet header, and obtains the credit packet corresponding to sequence number ④. The subsequent credit packet enters the export switch 1 through the long-distance link ⑤, and explicitly records the forwarding path of the export switch 1 for the credit in the credit packet header, and obtains the credit packet corresponding to sequence number ⑥. The credit packet enters the same TOR switch through link ⑦, and explicitly records the forwarding path of the same TOR switch for the credit in the credit packet header, and obtains the credit packet corresponding to sequence number ⑧. Finally, the sender receives the credit packet corresponding to sequence number ⑧.
[0034] The message structure of credit packet and data packet is as follows Figure 5. After receiving the credit packet, the sender copies the explicit path tag field to the data packet through the copy field method of sequence number ⑨, and obtains the data packet corresponding to sequence number ⑩ with the explicit path tag field in the header. When the switch forwards the data packet, it forwards it from the specified port according to the port number recorded by the hop counter, and reduces the hop counter by one, so that the data packet reaches the receiver along the symmetric path hop by hop according to the ingress port number recorded in the field. This function involves the write and read operations of the message header across all switches in the data center environment, which is easy to implement. The explicit path mechanism only takes advantage of the multi-path. In order to reduce the basic delay of the active transmission flow, the following key mechanisms are also required: connection status table and host granularity expected bandwidth table.
[0035] (3) in one embodiment, maintenance of a connection state table and a host-granular expected bandwidth table; In order to effectively utilize idle credits to reduce the basic delay of active transmission, the host throughput table (HTT) of the remote data center is maintained at the egress switch. Because modern data centers are oversubscribed to varying degrees, the core bandwidth is usually smaller than the uplink bandwidth of all server nodes, so that the sending hosts under the active transmission scheduling characteristics will not be fully occupied at the same time. The passive solution causes the oversubscribed core to have traffic exceeding the carrying capacity due to uncontrolled line speed startup, resulting in the accumulation of queues to buffer excessive traffic, which is also the main congestion problem that passive transmission attempts to solve; while active transmission consumes time to allocate bandwidth and has the advantage of limited shallow data queues, so that the network core of active transmission is fully utilized in ideal situations. However, due to the long delay caused by long distances, the link may be in a state of low utilization, which is caused by the discardable nature of credits and the post-credit stage.
[0036] Therefore, there may be idle bandwidth under long-distance active transmission, which makes it possible to accelerate active transmission in the pre-credit stage. Compared with Aeolus' heuristic occupation of idle bandwidth, the key point of this embodiment is to accurately discover idle host bandwidth through the host-granularity expected bandwidth table to avoid conflicts between the core occupied bandwidth and the bandwidth allocation of credits, and use the idle credits that cause idle bandwidth to accelerate the transmission in the pre-credit stage. In order to effectively judge the idle bandwidth, the expected bandwidth is maintained at the host granularity of the remote data center at the egress switch, and the host-granularity expected bandwidth table of the remote host is maintained according to the IP address of the credit sent by the egress switch to the local. The host-granularity expected bandwidth table of the egress switch records the number of credit packets from each remote host and sends it at fixed intervals. Time update, because it is updated at a fixed time interval, does not require division operations, only simple counting operations are required, avoiding the overhead of switch division. Host granularity expected bandwidth table constraints The number of credits per host during the time is less than ,in is the uplink bandwidth of the host, and MTU is the maximum transmission unit.
[0037] The egress switch manages the cross-data center connections from the local data center and maintains the connection status table shown in Table 1: Table 1
[0038] The status table maintenance information can be divided into the following three sections: Basic tuple information for the connection: source IP address ( )、Destination IP address( )、Source Port( )、Destination Port( ); Connection status flag , the sending state and the ending state have opposite bits; Traffic size , used for the volume of this traffic, in bytes; Create request time , records the timestamp when the request is established, indicating that the latest connection needs to be accelerated first.
[0039] Therefore, the egress switch of this embodiment needs to perform two additional operations compared to ordinary switches. One is to maintain the throughput of the remote data center host granularity (i.e., maintain the host granularity expected bandwidth table). The egress switch counts the host bandwidth corresponding to the header IP address of the credit packet sent to the local data center (i.e., maintains the connection status table). This requires a fixed amount of resource overhead for the egress switch and can be performed on the control plane. The egress switch counts the number of credit packets of each host (IP address) at fixed time intervals. The host granularity expected bandwidth table is actually a counter because it counts the number of credit packets of each host (IP address) at fixed intervals. The time update feature does not require specific division operations when calculating bandwidth, reducing the computational overhead of the switch.
[0040] (4) The egress switch accelerates the transmission in the pre-credit phase; According to the aforementioned host-granular expected throughput table maintenance, the egress switch has observed hosts with idle bandwidth receiving capabilities and can receive data packets. Since the network core is occupied by credit packets, the network core bandwidth can be planned based on the idle credits by explicit path constraint scheduling, which can safely accelerate the transmission of flows in the pre-credit stage. Idle credits can ensure that the core bandwidth will not be overused because of the bandwidth scheduling of the network core. At the same time, the host's uplink can be effectively accelerated because of the topological oversubscription characteristics of modern data centers. Consider such a typical scenario, a certain number of hosts are connected to a Top of Rack (ToR) switch. When these hosts are receivers, the credits sent by the same ToR switch are forwarded through the long-distance data center core layer to the local data center egress switch. A large number of hosts are connected to the same ToR switch. If the idle credits are converted into connections belonging to other hosts under the same ToR switch, since the implementation of path marking does not rely on the common hash load balancing based on five-tuple information in the data center, it can ensure that the path of the core layer remains consistent after the IP and port are changed.
[0041] For hosts under the same ToR switch with inconsistent first hops, the host-granular expected bandwidth table is used to handle it. If there is idle bandwidth, the credit can be safely converted. Since the existing path information is retained, the only different path from the host to the same ToR switch can also be safely converted because the egress switch detects that there is idle bandwidth.
[0042] In order to realize the knowledge of the connection level of the local data center by the egress switch, it is necessary to maintain a connection status table at the egress switch to record the connection status initiated by the local switch, including basic connection tuple information (source IP address, destination IP address, source port number, destination port number, etc.), request establishment time, and status flag. The two states of the connection are identified from the three stages of active transmission: sending state and ending state. The sending state indicates that the sender of the connection can send, and can send after receiving credit; the ending state indicates that the connection has been sent and declares the idle characteristics of subsequent credits.
[0043] For the credit received by the egress switch from the remote end, the connection status table of the egress switch is checked and found to be in the ended state. The credit is then forwarded to the sender of the request that was most recently established under the same ToR switch to which the host belongs. Under optimal conditions, the basic delay of active transmission can be reduced from the long-distance end-to-end millisecond delay to the microsecond delay of two or three hops within the data center, greatly reducing the basic delay problem.
[0044] Specifically: Figure 6 This is the architecture diagram of the egress switch. According to the connection status table of the egress switch, the transmission process of the pre-credit stage is accelerated by using idle credit security: For the credit packets received by the egress switch from the remote data center, the connection status table is queried. If the current connection is already in the ended state, it indicates that the credit is idle credit and idle bandwidth has been allocated. If it is not used, it will damage the core bandwidth across data centers. Then, based on the host granularity expected bandwidth table and the connection status table, which records the connection of the destination host under the same ToR switch to which the idle credit belongs, for the destination host with idle bandwidth, the connection with the most recent establishment time sent to the idle host is selected to obtain the necessary IP and port information of the connection. The egress switch rewrites the IP and port information of the credit, converting it from an idle credit to a credit sent to a connection in the sending state, and retains the existing explicit path tag header. Subsequent forwarding continues with explicit path tagging, which does not affect the implementation of the symmetric path.
[0045] In step 3, the acceleration benefit coefficient of the same ToR switch as the idle credit is determined according to the host granularity expected bandwidth table and the connection status table. The sender of the largest connection is used as the destination host, that is, in the use of idle credit security to accelerate the transmission of the pre-credit stage, in order to clarify the effectiveness of accelerating the transmission of the pre-credit stage, the acceleration benefit coefficient is used To measure: ; (1) in, is the current moment in the transmission process of the pre-credit phase, To establish the request time, is the end-to-end transmission time, is the host bandwidth, The flow rate.
[0046] This embodiment adds Constraints: , so that The calculation range is 0 to 50%, where The larger the value, the more effective it is to accelerate the transmission in the pre-credit phase. The core goal of formula (1) is to quantify the effect of early transmission in the pre-credit phase on reducing latency. Represents the total time required to complete credit allocation and transmission in the traditional protocol (two RTTs + transmission time); Indicates the estimated transmission time of the connection acceleration (including the time after the request is established, RTT and theoretical transmission time). Close to 50%, indicating that the security acceleration pre-credit phase of this embodiment almost completely avoids the waiting time in the traditional protocol. , traverse the connection status flag in the connection status table for the connection in the sending state, and select The parameters in formula (1) jointly characterize the transmission efficiency of the pre-credit stage, which is achieved by maximizing the acceleration benefit coefficient. , which can significantly reduce the latency of short flows and improve the overall performance of data center networks.
[0047] Formula (1) approximates The reduction ratio of latency brought by early transmission of traffic in the pre-credit stage. To accelerate transmission in the pre-credit stage, it is necessary to select the maximum acceleration benefit coefficient. At the same time, since the idle credits under the same ToR switch may be used to accelerate the traffic of a host, resulting in possible bandwidth oversubscription and queue establishment, the maximum data queue of the downlink port under the same ToR switch is for: ; (2) in, is the long-distance link bandwidth, The number of hosts on the same ToR switch.
[0048] when There is no queue risk; But more general scenarios include In the scenario where the bandwidth of the long-distance link is greater than the host bandwidth, a large queue is established and the aforementioned host-granular expected bandwidth table needs to be used for constraints: ; (3) in, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the expected bandwidth table at the host granularity, is the source IP address of the host.
[0049] In formula (3), The core goal is to ensure that the transmission of the accelerated flow does not cause additional queue accumulation. The optimization goal of formula (3) can be divided into two levels: (a1) maximizing the acceleration benefit coefficient (Right now ): That is, maximize the reduction ratio of the delay in the pre-credit phase of the traffic, which will reduce the waiting time in the credit negotiation phase and minimize the overall delay. (a2) Satisfy the bandwidth constraint of the destination host : Controlled The traffic bandwidth received by the host within the time does not exceed the capacity of the link, avoiding the establishment of additional queues, so that the accelerated pre-credit stage transmission will not bring additional queue accumulation risks and increase queuing delays.
[0050] By imposing the host bandwidth not exceeding the limit condition and then maximizing the delay acceleration benefit, the accelerated pre-credit stage transmission does not lead to the establishment of additional queues, that is, safe pre-credit stage acceleration is performed.
[0051] In this embodiment, if Figure 7 As shown, there are 4 hosts under a ToR switch, and each host has a connection F1, F2, F3, and F4 corresponding to the connection from the sender S to the corresponding numbered receiver R, that is, F1 is sent from the sender S1 to the receiver R1, F2 is sent from the sender S2 to the receiver R2, F3 is sent from the sender S3 to the receiver R3, and F4 is sent from the sender S4 to the receiver R4. Now a connection F4 ends, and a new connection F5 is added from the sender S4 to the receiver R4. At this time, the connection status table of the export switch records the entries of F1~F5, and the status bit of F4 indicates that F4 is in the end state. The export switch sets the credit of the post-credit stage received from F4 to idle credit. This embodiment uses the acceleration mechanism of the idle credit security acceleration pre-credit stage to rewrite the key fields (IP and port information of the credit) in the idle credit header as the connection information corresponding to F5, because according to the acceleration efficiency coefficient According to the calculation, F5 should be accelerated the most, and the destination host belongs to the same remote ToR switch. By using this method of safely accelerating active transmission flows using idle credits, the basic delay of active transmission is reduced, and the conflict problem between unscheduled packets and credit-scheduled packets is not introduced, which greatly improves the superiority of credit-based active transmission in long-distance scenarios. This embodiment has the following advantages: The delay problem of active transmission is optimized, and reliance on unscheduled data packets is avoided to reduce interference with active transmission bandwidth. This method is completely based on credit scheduling, and there is no need to introduce loss detection and retransmission mechanisms caused by active packet discard. The credit transmission in the pre-credit stage is accelerated through the explicit path marking function supported by ordinary switches and the host-granular expected bandwidth table maintenance and some connection status tables supported by egress switches. This implementation utilizes the features that ordinary switches can support, and the table maintenance function only requires the support of a minimum number of egress switches, minimizing the network deployment cost.
[0052] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A long-distance data center active transmission method for accelerating pre-credit stage transmission, characterized in that: The method is performed on a cross-data center network, and the method includes: The switch set up between the sender and the receiver uses a random packet spraying scheme and explicit path marking when forwarding credit packets, and randomly forwards the credit packets to the next hop in the routing table; The cross-data center egress switch maintains a connection status table of the local data center and a host-granular expected bandwidth table of the remote data center. The host-granular expected bandwidth table records the number of credit packets from each remote host and reports the number of credit packets to the remote host at fixed intervals. Time update; The connection status table is used to determine the status of the credit packet received by the egress switch from the remote data center. If the credit packet is an idle credit, the sender of the request that was most recently established under the same ToR switch as the idle credit is determined as the destination host based on the host granularity expected bandwidth table and the connection status table, and the idle credit is forwarded to the destination host, thereby utilizing the idle credit to safely accelerate the transmission in the pre-credit stage.
2. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1 is characterized in that: There are three phases of active transmission communication between sender and receiver: pre-credit phase, credit phase and post-credit phase; The pre-credit stage is when the sending host sends a credit request to the receiving host until the first credit from the receiving host is received. In the pre-credit stage, the sending host is in a waiting state by default and does not transmit any data packets. The credit stage is when the sending host receives the first credit from the receiving host, and the arrival of the credit triggers the sending of a data packet of MTU until all the corresponding data of this connection to the sending host is transmitted. The MTU is the maximum transmission unit; The post-credit stage is when the sending host sends all data and sends a credit stop control message, declaring that there is no traffic to be sent in this connection. The receiving host stops sending credit after receiving the credit stop control message at the latest.
3. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1 is characterized in that: When the switch uses a random packet spraying scheme when forwarding credit packets, it actively adheres to the symmetric path based on credit transmission and achieves the result of shallow queue through packet-level symmetric path constraints.
4. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 3 is characterized in that: The symmetric path is realized based on the explicit path marking method. The explicit path marking method is as follows: The switch's forwarding path for the credit is explicitly recorded in the credit packet header; In the explicit path marking strategy, when a switch receives a credit packet, it determines the ingress port number of the received credit packet, adds the ingress port number to the explicit path marking field of the credit packet header, and simultaneously increases the hop count counter by one, which records the number of forwarding times experienced by the credit packet; After receiving the credit packet, the sender copies the explicit path tag field to the data packet, and the data packet reaches the receiver along the symmetric path hop by hop according to the ingress port number recorded in the explicit path tag field.
5. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1 is characterized in that: The host-granular expected bandwidth table constraints The number of credits per host during the time is less than ,in is the uplink bandwidth of the host, and MTU is the maximum transmission unit.
6. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 2 is characterized in that: In determining the state of the credit packet received by the egress switch from the remote data center using the connection state table, two types of connection states are identified from the three stages of active transmission communication: a sending state and an ending state; The sending state indicates that the sender of the connection can send, and can send after receiving credit; The end state indicates that the connection has been sent and declares the idle nature of subsequent credits.
7. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1 is characterized in that: The connection status table includes: Basic tuple information of the connection: source IP address, destination IP address, source port, destination port; Connection status flag , the sending state and the ending state have opposite bits; Traffic size , used for the volume of this traffic, in bytes; Create request time , records the timestamp when the request is established, indicating that the latest connection needs to be accelerated first.
8. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 1 is characterized in that: Determine the acceleration benefit coefficient of the same ToR switch as the idle credit according to the host granularity expected bandwidth table and the connection status table The sender of the largest connection is used as the destination host, that is, in the transmission of the pre-credit stage using idle credit security acceleration, the acceleration benefit coefficient is used Measuring effectiveness: in, is the current moment in the transmission process of the pre-credit phase, To establish the request time, is the end-to-end transmission time, is the host bandwidth, The flow rate.
9. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 8, characterized in that: The maximum data queue of the downlink port under the same ToR switch during the pre-credit stage transmission process using idle credit safety acceleration for: in, is the long-distance link bandwidth, The number of hosts on the same ToR switch.
10. The long-distance data center active transmission method for accelerating pre-credit stage transmission according to claim 9, characterized in that: when There is no queue risk; when , that is, when the long-distance link bandwidth is greater than the host bandwidth, there is a risk of queue establishment, and the host granularity expected bandwidth table is used for constraint: in, is the uplink bandwidth of the host, MTU is the maximum transmission unit, is the expected bandwidth table at the host granularity, is the source IP address of the host.
Citation Information
Patent Citations
Credit packet-based active transmission method for data center
CN112437019A
Data stream transmission scheduling control method and system in cloud data center application
CN115022249A
Credit-based multipath transmission method for datacenter network load balancing
KR101932138B1