TCP acceleration method for cross-data center network

By deploying the TCP congestion control algorithm and congestion state judgment algorithm on the switch, the congestion coefficient is calculated in real time and the appropriate data stream processing mode is selected, the problems of untimely feedback of congestion information and inaccurate judgment of congestion state in cross-data center networks are solved, and the accurate slowdown and growth of data flow is achieved, network throughput is improved and queue delay is reduced.

CN119996318APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178746.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The problems of throughput drop and queue delay increase in cross-data center networks due to untimely feedback on congestion information and inaccurate judgment of congestion status.

Method used

By deploying the TCP congestion control algorithm and congestion state judgment algorithm on the switch, the congestion coefficient is calculated in real time, the release mode, speed limit release mode and non-release mode are selected for data stream processing, and the ACK message header is parsed and modified on the data plane to achieve accurate speed reduction operation.

Benefits of technology

Effectively eliminate the problems of long congestion control loops and untimely adjustment of transmission rates caused by long-distance links, ensure the accurate slowdown and growth behavior of data flows, improve network throughput and reduce queue delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996318A_ABST
    Figure CN119996318A_ABST
Patent Text Reader

Abstract

The invention discloses a TCP acceleration method for a cross-data center network, and relates to the technical field of cross-data center network congestion control. The invention discloses a TCP acceleration method oriented to a cross-data center network. A relay switch device is utilized to delay forwarding of congestion information of a long-distance transmission data stream and temporarily store part of congestion messages at the same time, so that the problems of long congestion control loop and untimely sending rate adjustment caused by a long-distance link are solved, accurate speed reduction and speed increase behaviors of the data stream are kept, and the transmission efficiency of the data stream is improved. When the cross-data center network is congested, the relay device can calculate the congestion degree through various congestion signals and accurately judge the type of the switch which is congested in the cross-data center network so as to select a corresponding speed regulation scheme and eliminate network congestion, and the method has the advantages of being convenient to deploy, high in compatibility and high in reliability. And meanwhile, normal use can be realized without modifying system network configuration by the sending end or the receiving end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-data center network congestion control, and in particular to a TCP acceleration method for a cross-data center network. Background Art

[0002] An inter-data center network refers to a network in which data center networks in different regions are connected via a wide area network. It is designed to provide high-throughput transmission capabilities across long-distance wide area network links for different data centers. Business scenarios include redundant backup of business data between two data centers and maintaining data consistency. Currently, the TCP transmission protocol is mainly used in wide area networks, but in inter-data center network transmission scenarios, TCP traffic often faces problems of low throughput and high transmission latency. The specific problem details are as follows:

[0003] 1. Data centers in different geographical locations are connected through a wide area network. The communication link distance is long and the round-trip delay (RTT) is high. When congestion occurs, the congestion information cannot be transmitted to the sender in time, resulting in the sender rate not being adjusted in time and unable to effectively handle the congestion in the network, resulting in low throughput and unable to effectively meet the communication needs between data center networks.

[0004] 2. The buffer sizes of routers / switches in data center networks vary greatly. The buffer size of the egress router can reach GB level, while other switches such as core layer switches use a shallow buffer design with only tens of MB buffers. The buffer capacity of the two differs by nearly 1,000 times. When the current TCP congestion control algorithm cannot accurately perceive specific congestion information, inaccurate and untimely congestion control will cause unnecessary throughput loss.

[0005] There are existing improved TCP congestion control algorithms and software-based TCP acceleration solutions, but the performance of the improved TCP congestion control algorithm in real scenarios has not been widely verified, and the deployment is relatively complex, requiring changes to the switch device configuration. The software-based TCP solution does not require changes to the device configuration, but is incompatible with the current mainstream TCP congestion control algorithm and has obvious performance bottlenecks. Therefore, a TCP acceleration method for cross-data center networks is proposed. Summary of the invention

[0006] In view of the deficiencies in the prior art, the present invention provides a TCP acceleration method for cross-data center networks, which solves the problem of decreased throughput and increased queuing delay caused by untimely feedback of congestion information and inaccurate congestion status judgment.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A TCP acceleration method for a cross-data center network specifically includes the following steps:

[0008] S1. Connect different data center networks through a wide area network to form an inter-data center network. Deploy a TCP congestion control algorithm in the inter-data center network. Deploy a switch at the near end of the data center network as a relay device to access the inter-data center network. The switch maintains an information table of each data flow on each thread and records the information table and ECN mark bit of the data flow passing through the switch.

[0009] S2. The switch receives a congestion signal from the destination. The switch selects one of the release mode, speed-limited release mode, and non-release mode according to the congestion signal and the switch buffer occupancy. For non-congested data flows, the switch can forward them normally. For congested data flows, the switch selects speed-limited forwarding or suspends forwarding of the data packets of the data flow according to the specific congestion signal and the current switch buffer occupancy. At this time, the data packets will be temporarily cached in a special queue of the switch.

[0010] S3. Deploy a congestion status judgment algorithm on the switch to calculate the congestion coefficients of different data flows in the cross-data center network in real time, and divide the data flows into shallow buffer congested data flows and deep buffer congested data flows according to the congestion coefficients. For deep buffer congested data flows, the switch parses and modifies the ACK message header on the data plane, and the source port performs a speed reduction operation after receiving the ACK message.

[0011] The present invention is further configured as follows: the information table of the data flow in S1 includes a source-destination four-tuple, a congestion two-tuple and a round-trip delay table;

[0012] The source-destination quadruple includes a source IP address, a destination IP address, a source port, and a destination port, and is used to add a unique identifier to the data flow passing through the switch;

[0013] The congestion tuple includes a message sequence number and a timestamp, which is used to record the congestion status of the data flow corresponding to the message;

[0014] The round trip delay table is used to record the latest round trip delay of the data flow passing through the switch.

[0015] The present invention is further configured as follows: the source IP address, destination IP address, source port, destination port, message sequence number, timestamp and latest round-trip delay are respectively stored in a register in the switch.

[0016] The present invention is further configured as follows: the method in which the switch in S1 maintains each data flow information table includes:

[0017] A1. Define a hash function to map the source-destination quadruple and the message sequence number to a unique index position in the register, which is used to maintain consistency when storing and retrieving data in the source-destination quadruple, congestion binary, and round-trip delay tables;

[0018] A2. The switch parses the Ethernet header field, IP header field, and TCP header field of the TCP message at the ingress at line speed, and assigns values ​​to the source-destination 4-tuple and the congestion 2-tuple respectively.

[0019] A3. While parsing the TCP message header field, the switch obtains the timestamp carried by the current TCP message and calculates the round-trip delay at the outbound port. After obtaining the ACK message sequence number, the index in the congestion tuple is obtained by calculating the hash value, the sending timestamp of the corresponding TCP message is obtained, and the latest round-trip delay is obtained and written into the round-trip delay table. The calculation formula includes:

[0020] lastRTT=rttTable[index,seq]

[0021] curRTT=(1-α)*lastRTT+α*(curTime-flowsTable[index,TS])

[0022] Where rttTable is the round-trip delay table, index is the index of the congestion tuple, lastRTT is the last calculated round-trip delay, seq is the message sequence number, curRTT is the latest round-trip delay, α is an empirical setting parameter, and TS is a constant, which represents the register storing the timestamp.

[0023] The present invention is further configured as follows: after the hash function is defined in A1, the method of calculating the hash value includes:

[0024] indexHash=(srcIP^dstIP&srcPort^dstPort)%M+K%1024

[0025] Where srcIP is the source IP address, dstIP is the destination IP address, srcPort is the source port, dstPort is the destination port, ^ is the XOR operation, & is the AND operation, % is the modulo operation, M is 1024, and K is a parameter used to prevent hash conflicts.

[0026] The present invention is further configured as follows: the switch in S2 selects one of the release mode, the speed-limited release mode and the non-release mode to operate according to the congestion signal and the switch buffer occupancy, including:

[0027] B1. After parsing the TCP message at the ingress, the switch checks the ACK flag bit and SACK field in the data packet header. When the ACK flag bit is 1 and the SACK field shows packet loss information, the data flow enters the non-forwarding mode: the egress port of subsequent messages of the data flow is corrected to the switch self-circulation port recycPort, and subsequent messages are temporarily cached in a dedicated queue in the switch;

[0028] B2. For data flows in non-forwarding mode, a register is used to store the SACK information of the message with the lowest sequence number in the buffer queue. When the buffer queue length is greater than the congestion threshold, subsequent messages are discarded.

[0029] B3. There is no SACK packet loss information in the TCP message header field, and the ECN mark shows that it is effective. The data flow enters the speed-limited forwarding mode: the forwarding speed of the data flow message is halved. The selection rule is that the message whose last 1 bit of the TCP message sequence number modulo 2 equals 0 can be forwarded normally. The outbound port of the message that does not meet the selection rule is changed to a self-loop port, and these messages are temporarily cached in the special queue of the switch;

[0030] B4. When the length of the buffer queue is greater than the congestion threshold, the switch sets the ECN field in the message header corresponding to the data flow in the speed-limited forwarding mode to congestion, generates a pseudo ACK packet, and notifies the source port TCP congestion control algorithm to reduce the packet sending window. All subsequent messages of the data flow are forwarded normally, and the operation of sending messages to the special queue is stopped until the length of the switch special queue is 0;

[0031] B5. When the buffer queue length is not 0 and is less than the congestion threshold, it means that the source port is slowing down or the switch congestion is improved, but the data flow is still in a congested state. The subsequent message forwarding speed of the data flow in the speed-limited forwarding mode is halved according to B3 until the buffer queue length is greater than the congestion threshold again. The switch cyclically controls the data flow messages in the speed-limited forwarding mode until the data flow is released or not released;

[0032] B6. If no packet loss information is detected in the SACK field and the ECN flag does not indicate congestion, the data flow is judged to be in release mode;

[0033] B7. When the data flow is judged to be in release mode and there are messages cached in the special queue of the switch at the same time, the messages in the special queue of the switch are sent to the corresponding destination port with the highest priority.

[0034] When the accumulation of any message exceeds the upper limit of the switch buffer, the message at the tail of the queue will be discarded.

[0035] The present invention is further configured as follows: the calculation method of the congestion coefficient in S3 includes:

[0036] C=ΔRTT*(qDepth+1)*(RTT>>β+1)

[0037] In the formula, C is the congestion coefficient, ΔRTT is the change in the round-trip delay of the data flow, qDepth is the accumulation length of the buffer queue corresponding to the data flow in the switch, RTT is the round-trip delay of the data flow, >> is a right shift operation, and β is the weight value representing RTT.

[0038] The present invention is further configured as follows: the processing method of the shallow cache congested data flow and the deep buffer congested data flow in S3 includes:

[0039] C1. At the ingress, the switch checks the ECN flag bit in the TCP packet header every N packets. The calculation method of interval N includes:

[0040] N=RTT%T+1

[0041] Where RTT is the round trip delay of the data flow to which the data packet belongs, % is the modulo operation, and T is a constant, indicating an empirical parameter setting;

[0042] C2. When the ECN value is 0b10 or 0b01 and the buffer queue length is less than the congestion threshold, the information table is retrieved to obtain the current round-trip delay change value of the data flow corresponding to the TCP message, and the congestion coefficient is calculated;

[0043] C3. Reverse the source-destination quadruple obtained by parsing the TCP message header, swap the source IP address and the destination IP address, swap the source port and the destination port, form a new source-destination quadruple value and write it into the source-destination quadruple;

[0044] C4. When the congestion coefficient C is less than T_C, T_C is the set threshold parameter, indicating that congestion occurs in the shallow buffer switch of the data center network. The round-trip delay change value of the data flow is calculated twice every N data packets, and the average value of the two round-trip delay change values ​​is calculated;

[0045] C5. When the average value in C4 is less than T_C+M, M is an empirical parameter, indicating that the congestion level on the shallow cache switch of the data center network is low, and the source port does not use ECN speed reduction. The switch is used to identify the ACK message of the data flow, delete the ECN mark bit in the message header, and eliminate the perception of the ECN mark by the TCP congestion control algorithm of the source port;

[0046] C6. In C4, the average value is greater than T_C+M, indicating that the congestion level on the shallow cache switch of the data center network is high, and the source port needs to be decelerated to retain the ECN mark bit on the ACK message of the data flow. The TCP congestion control algorithm of the source port adjusts the message sending rate through the ECN mark bit;

[0047] C7. When the congestion coefficient is greater than T_C, it indicates that the deep buffer switch of the data center network is congested. The switch is used to identify the ACK message of the data flow. For all subsequent ACK data messages, the ECN mark bit is displayed as congestion. At the same time, the three latest consecutive ACK messages are selected, and their SACK fields are modified to indicate that the message corresponding to the current message header sequence number is lost. The source port TCP congestion control algorithm performs the corresponding packet loss and speed reduction operation.

[0048] The present invention provides a TCP acceleration method for a cross-data center network. It has the following beneficial effects:

[0049] The present invention utilizes a relay switch device to delay forwarding the congestion information of a long-distance transmission data flow and temporarily store some congested messages, thereby eliminating the problems of long congestion control loops and untimely transmission rate adjustment caused by long-distance links, so that the data flow maintains accurate deceleration and acceleration behaviors. Moreover, when congestion occurs in the inter-data center network, the relay device can calculate the degree of congestion through a variety of congestion signals, accurately determine the type of switch that is congested in the inter-data center network, and select the corresponding speed regulation scheme to eliminate network congestion. The present invention has the advantages of convenient deployment and high compatibility, and can be used normally without the need for the sender or the receiver to modify the system network configuration. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic diagram of a local congestion processing flow in an embodiment of the present invention;

[0051] Figure 2 A schematic diagram of storing messages in a special queue in an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of remote congestion notification in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0054] See also Figure 1-3 The embodiment of the present invention provides the following technical solution: a TCP acceleration method for a cross-data center network, the steps of performing local congestion processing include:

[0055] S1. Connect different data center networks through a wide area network to form an inter-data center network. Deploy a TCP congestion control algorithm in the inter-data center network. The TCP congestion control algorithm includes but is not limited to the Cubic algorithm and the BBR algorithm. The data center network deploys a switch at the near end as a relay device to access the inter-data center network. The switch maintains an information table for each data flow on each thread and records the information table and ECN mark bit of the data flow passing through the switch.

[0056] The information table of the data flow includes a source-destination four-tuple, a congestion two-tuple, and a round-trip delay table;

[0057] The source-destination quadruple includes the source IP address, destination IP address, source port, and destination port, which are used to add a unique identifier to the data flow passing through the switch;

[0058] The congestion tuple includes a message sequence number and a timestamp, which is used to record the congestion status of the data flow corresponding to the message;

[0059] The round-trip delay table is used to record the latest round-trip delay of the data flow passing through the switch.

[0060] Furthermore, four registers are defined in the switch for storing source-destination quadruple, and each register is used to store source IP address, destination IP address, source port and destination port respectively.

[0061] Two registers are defined in the switch for storing congestion tuple information, and each register stores a message sequence number and a timestamp respectively.

[0062] A register is defined in the switch to store the latest round-trip delay of a data flow.

[0063] To further explain, the switch maintains each data flow information table in the following ways:

[0064] A1. Define a hash function to map the source-destination quadruple and the message sequence number to a unique index position in the register, which is used to maintain consistency when storing and retrieving data in the source-destination quadruple, congestion binary, and round-trip delay table. The hash value is calculated in the following ways:

[0065] indexHash=(srcIP^dstIP&srcPort^dstPort)%M+K%1024

[0066] Where srcIP is the source IP address, dstIP is the destination IP address, srcPort is the source port, dstPort is the destination port, ^ is the XOR operation, & is the AND operation, % is the modulus operation, M is 1024, and K is a parameter used to prevent hash conflicts.

[0067] A2. The switch parses the Ethernet header field, IP header field, and TCP header field of the TCP message at the ingress at line speed, and assigns values ​​to the source-destination 4-tuple and the congestion 2-tuple respectively.

[0068] A3. While parsing the TCP message header field, the switch obtains the timestamp carried by the current TCP message and calculates the round-trip delay at the outbound port. After obtaining the ACK message sequence number, the index in the congestion tuple is obtained by calculating the hash value, the sending timestamp of the corresponding TCP message is obtained, and the latest round-trip delay is obtained and written into the round-trip delay table. The calculation formula includes:

[0069] lastRTT=rttTable[index,seq]

[0070] curRTT=(1-α)*lastRTT+α*(curTime-flowsTable[index,TS])

[0071] Where rttTable is the round-trip delay table, index is the index of the congestion tuple, lastRTT is the last calculated round-trip delay, seq is the message sequence number, curRTT is the latest round-trip delay, α is an empirical setting parameter, and TS is a constant, which represents the register storing the timestamp.

[0072] S2. The switch receives a congestion signal from the destination end. The switch selects one of the release mode, speed-limited release mode, and non-release mode according to the congestion signal and the switch buffer occupancy. Specifically, for non-congested data flows, the switch can forward them normally. For congested data flows, the switch selects speed-limited forwarding or suspends forwarding of the data packets of the data flow according to the specific congestion signal and the current switch buffer occupancy. At this time, the data packets will be temporarily cached in a special queue of the switch.

[0073] As a detailed description, the manner in which the switch selects one of the release mode, the rate-limited release mode, and the non-release mode to operate according to the congestion signal and the switch buffer occupancy includes:

[0074] B1. After parsing the TCP message at the ingress, the switch checks the ACK flag bit and SACK field in the data packet header. When the ACK flag bit is 1 and the SACK field shows packet loss information, the data flow enters the non-forwarding mode: the egress port of subsequent messages of the data flow is corrected to the switch self-circulation port recycPort, and subsequent messages are temporarily cached in a dedicated queue in the switch;

[0075] B2. For data flows in non-forwarding mode, a register is used to store the SACK information of the message with the lowest sequence number in the buffer queue. When the buffer queue length is greater than the congestion threshold, subsequent messages are discarded.

[0076] B3. There is no SACK packet loss information in the TCP message header field, and the ECN mark shows that it is effective. The data flow enters the speed-limited forwarding mode: the forwarding speed of the data flow message is halved. The selection rule is: ipv4.hdr.seq[0:0]%2==0. ipv4.hdr.seq is a data structure defined on the programmable switch. The rule means that the message whose last 1 bit of the TCP message sequence number is equal to 0 modulo 2 can be forwarded normally. The outgoing port of the message that does not meet the rule is changed to the self-loop port, and these messages are temporarily cached in the special queue of the switch;

[0077] B4. When the length of the buffer queue is greater than the congestion threshold, the switch sets the ECN field in the message header corresponding to the data flow in the speed-limited forwarding mode to congestion, generates a pseudo ACK packet, and notifies the source port TCP congestion control algorithm to reduce the packet sending window. All subsequent messages of the data flow are forwarded normally, and the operation of sending messages to the special queue is stopped until the length of the switch special queue is 0;

[0078] B5. When the buffer queue length is not 0 and is less than the congestion threshold, it means that the source port is slowing down or the switch congestion is improved, but the data flow is still in a congested state. The subsequent message forwarding speed of the data flow in the speed-limited forwarding mode is halved according to B3 until the buffer queue length is greater than the congestion threshold again. The switch cyclically controls the data flow messages in the speed-limited forwarding mode until the data flow is released or not released;

[0079] B6. If no packet loss information is detected in the SACK field and the ECN flag does not indicate congestion, the data flow is judged to be in release mode;

[0080] B7. When the data flow is judged to be in release mode and there are messages cached in the special queue of the switch at the same time, the messages in the special queue of the switch are sent to the corresponding destination port with the highest priority.

[0081] It should be noted that for the above release mode, speed-limited release mode and non-release mode, when the accumulation of any message exceeds the upper limit of the switch buffer, the message at the tail of the queue will be discarded.

[0082] The remote congestion notification module includes the following specific steps:

[0083] Deploy a congestion status judgment algorithm on the switch to calculate the congestion coefficient of different data flows in the cross-data center network in real time. Specifically, the congestion coefficient is calculated in the following ways:

[0084] C=ΔRTT*(qDepth+1)*(RTT>>β+1)

[0085] In the formula, C is the congestion coefficient, ΔRTT is the round-trip delay change of the data flow, qDepth is the buffer queue accumulation length corresponding to the data flow in the switch, RTT is the round-trip delay of the data flow, >> is a right shift operation, and β is the weight value representing RTT;

[0086] According to the congestion coefficient, the data flow is divided into shallow buffer congestion data flow and deep buffer congestion data flow. For the deep buffer congestion data flow, the switch parses and modifies the ACK message header on the data plane, and the source port performs a speed reduction operation after receiving the ACK message.

[0087] As a detailed description, the shallow buffer congested data flow and the deep buffer congested data flow are handled as follows:

[0088] C1. At the ingress, the switch checks the ECN flag bit in the TCP packet header every N packets. The calculation method of interval N includes:

[0089] N=RTT%T+1

[0090] Where RTT is the round trip delay of the data flow to which the data packet belongs, % is the modulo operation, and T is a constant, indicating an empirical parameter setting;

[0091] C2. When the ECN value is 0b10 or 0b01 and the buffer queue length is less than the congestion threshold, the information table is retrieved to obtain the current round-trip delay change value of the data flow corresponding to the TCP message, and the congestion coefficient is calculated;

[0092] C3. Reverse the source-destination quadruple obtained by parsing the TCP message header, swap the source IP address and the destination IP address, swap the source port and the destination port, form a new source-destination quadruple value and write it into the source-destination quadruple;

[0093] C4. When the congestion coefficient C is less than T_C, T_C is the set threshold parameter, indicating that congestion occurs in the shallow buffer switch of the data center network. The round-trip delay change value of the data flow is calculated twice every N data packets, and the average value of the two round-trip delay change values ​​is calculated;

[0094] C5. When the average value in C4 is less than T_C+M, M is an empirical parameter, indicating that the congestion level on the shallow cache switch of the data center network is low, and the source port does not use ECN speed reduction. The switch is used to identify the ACK message of the data flow, delete the ECN mark bit in the message header, and eliminate the perception of the ECN mark by the TCP congestion control algorithm of the source port;

[0095] C6. In C4, the average value is greater than T_C+M, indicating that the congestion level on the shallow cache switch of the data center network is high, and the source port needs to be decelerated to retain the ECN mark bit on the ACK message of the data flow. The TCP congestion control algorithm of the source port adjusts the message sending rate through the ECN mark bit;

[0096] C7. When the congestion coefficient is greater than T_C, it indicates that the deep buffer switch of the data center network is congested. The switch is used to identify the ACK message of the data flow. For all subsequent ACK data messages, the ECN mark bit is displayed as congestion. At the same time, the three latest consecutive ACK messages are selected, and their SACK fields are modified to indicate that the message corresponding to the current message header sequence number is lost. The source port TCP congestion control algorithm performs the corresponding packet loss and speed reduction operation.

[0097] It can be seen that the above TCP acceleration method for cross-data center networks includes three parts: switch calculation point, switch storage point and switch notification point. Specifically:

[0098] When the switch calculation point receives a message, it parses the message header information, calculates the latest data at line speed, and writes it into the switch register in real time;

[0099] When forwarding a message, the switch storage point calculates and defines the message forwarding logic according to the data of the switch calculation point. The switch storage point forwards the message according to a certain process or temporarily stores the message in a special queue;

[0100] When forwarding messages, the switch notification point determines the cross-data center network congestion situation based on the data from the switch calculation point. If the congestion occurs in a shallow buffer switch, the switch notification point adjusts the message ECN mark according to a certain logic; if the congestion occurs in a deep buffer switch, the switch notification point generates a packet loss signal and notifies the source end to adjust the sending rate.

[0101] When performing local congestion handling, such as Figure 1 As shown, the specific operations include the following:

[0102] Each time the switch receives a message, it parses the message header information to update the information flow table on the switch. Then the switch determines the forwarding mode of the data flow message through the ECN and SACK fields in the message header. If the data flow is in the non-release mode, the buffer queue length is queried. If it is greater than the given threshold, the subsequent messages are discarded, otherwise the subsequent messages are temporarily cached; if the data flow is in the speed-limited release mode, the buffer queue length is queried. If it is greater than the given threshold, the source end is notified to reduce the speed and the special queue is emptied until the buffer queue length is less than the given threshold; otherwise, the message forwarding rate is halved; if the data flow is in the release mode, the special queue length is checked, the messages in the special queue are forwarded first, and then all messages are forwarded normally.

[0103] When a message is stored in a special queue, such as Figure 2 As shown, the specific operations include the following:

[0104] The switch calculates and stores the source port, source IP address, destination port, destination IP address, ECN flag, message sequence number, timestamp, and latest round-trip delay of each message by defining registers. When the message arrives at the switch entrance, the corresponding information is parsed and written into the register. The switch reads and calculates the values ​​in the register, and caches the number of messages that are not released or released at a limited speed in a special queue. The remaining messages are forwarded through the in / out queues. If the data flow meets the remote congestion notification conditions, the switch will forge a packet loss signal to notify the source end that the sending rate needs to be reduced.

[0105] When the remote end notifies of congestion, Figure 3 As shown, the specific operations include the following:

[0106] After the switch recognizes that the ECN mark in the message header is effective, it calculates the congestion coefficient of the data flow to which the message belongs. If the congestion coefficient is greater than the given threshold, it actively modifies the ACK message header and then forwards the message normally; otherwise, it calculates the average round-trip delay of the data flow in two consecutive time periods, that is, the average RTT, and compares the result with the given threshold; if it is less than the threshold, it directly forwards the message, otherwise it deletes the ECN mark in the message header to avoid sending congestion signals to the source end by mistake.

[0107] In summary, it can be seen that the present invention is based on the background of the need for high throughput in cross-data center networks, and provides a timely, high-throughput acceleration method mechanism for congestion control in cross-data center networks. In long-distance and high-RTT network links, the method of optimizing the TCP congestion control algorithm cannot respond to the remote bottleneck congested switch in time, which will lead to packet loss and cause the problem of decreased link throughput. At the same time, the current Cubic congestion control algorithm cannot identify the exact congestion points in the cross-data center network, and adjusting the sending rate too aggressively or too gently may cause unnecessary throughput loss.

[0108] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A TCP acceleration method for a cross-data center network, characterized by: The specific steps include: S1. Connect different data center networks through a wide area network to form an inter-data center network. Deploy a TCP congestion control algorithm in the inter-data center network. Deploy a switch at the near end of the data center network as a relay device to access the inter-data center network. The switch maintains an information table of each data flow on each thread and records the information table and ECN mark bit of the data flow passing through the switch. S2. The switch receives a congestion signal from the destination. The switch selects one of the following modes: release mode, speed-limited release mode, and non-release mode according to the congestion signal and the switch buffer occupancy. For non-congested data flows, the switch can forward them normally. For congested data flows, the switch selects speed-limited forwarding or suspends forwarding of the data packets of the data flow according to the specific congestion signal and the current switch buffer occupancy. At this time, the data packets will be temporarily cached in a special queue of the switch. S3. Deploy a congestion status judgment algorithm on the switch to calculate the congestion coefficients of different data flows in the cross-data center network in real time, and divide the data flows into shallow buffer congested data flows and deep buffer congested data flows according to the congestion coefficients. For deep buffer congested data flows, the switch parses and modifies the ACK message header on the data plane, and the source port performs a speed reduction operation after receiving the ACK message.

2. According to claim 1, a TCP acceleration method for a cross-data center network is characterized in that: The information table of the data flow in S1 includes a source-destination four-tuple, a congestion two-tuple and a round-trip delay table; The source-destination quadruple includes a source IP address, a destination IP address, a source port, and a destination port, and is used to add a unique identifier to the data flow passing through the switch; The congestion tuple includes a message sequence number and a timestamp, which is used to record the congestion status of the data flow corresponding to the message; The round trip delay table is used to record the latest round trip delay of the data flow passing through the switch.

3. According to claim 2, a TCP acceleration method for a cross-data center network is characterized in that: The source IP address, destination IP address, source port, destination port, message sequence number, timestamp and latest round-trip delay are respectively stored in a register in the switch.

4. The TCP acceleration method for a cross-data center network according to claim 3, characterized in that: The method in which the switch in S1 maintains each data flow information table includes: A1. Define a hash function to map the source-destination quadruple and the message sequence number to a unique index position in the register, which is used to maintain consistency when storing and retrieving data in the source-destination quadruple, congestion binary, and round-trip delay tables; A2. The switch parses the Ethernet header field, IP header field, and TCP header field of the TCP message at the inbound port at line speed, and assigns values ​​to the source-destination 4-tuple and the congestion 2-tuple respectively. A3. While parsing the TCP message header field, the switch obtains the timestamp carried by the current TCP message and calculates the round-trip delay at the outbound port. After obtaining the ACK message sequence number, the index in the congestion tuple is obtained by calculating the hash value, the sending timestamp of the corresponding TCP message is obtained, and the latest round-trip delay is obtained and written into the round-trip delay table. The calculation formula includes: lastRTT=rttTable[index,seq] curRTT=(1-α)*lastRTT+α*(curTime-flowsTable[index,TS]), where rttTable is the round-trip delay table, index is the index of the congestion tuple, lastRTT is the last calculated round-trip delay, seq is the message sequence number, curRTT is the latest round-trip delay, α is the empirical setting parameter, and TS is a constant representing the register for storing timestamps.

5. The TCP acceleration method for a cross-data center network according to claim 4, characterized in that: After the hash function is defined in A1, the method of calculating the hash value includes: indexHash=(srcIP^dstIP&srcPort^dstPort)%M+K%1024 Where srcIP is the source IP address, dstIP is the destination IP address, srcPort is the source port, dstPort is the destination port, ^ is the XOR operation, & is the AND operation, % is the modulo operation, M is 1024, and K is a parameter used to prevent hash conflicts.

6. A TCP acceleration method for a cross-data center network according to claim 5, characterized in that: The switch in S2 selects one of the release mode, the speed-limited release mode and the non-release mode to operate according to the congestion signal and the switch buffer occupancy, including: B1. After parsing the TCP message at the inbound port, the switch checks the ACK flag bit and SACK field in the data packet header. When the ACK flag bit is 1 and the SACK field shows packet loss information, the data flow enters the non-forwarding mode: the outbound port of subsequent messages of the data flow is corrected to the switch self-circulation port recycPort, and subsequent messages are temporarily cached in a dedicated queue in the switch; B2. For data flows in non-forwarding mode, a register is used to store the SACK information of the message with the lowest sequence number in the buffer queue. When the buffer queue length is greater than the congestion threshold, subsequent messages are discarded. B3. There is no SACK packet loss information in the TCP message header field, and the ECN mark shows that it is effective. The data flow enters the speed-limited forwarding mode: the forwarding speed of the data flow message is halved. The selection rule is that the message whose last 1 bit of the TCP message sequence number modulo 2 equals 0 can be forwarded normally. The outbound port of the message that does not meet the selection rule is changed to a self-loop port, and these messages are temporarily cached in the special queue of the switch; B4. When the length of the buffer queue is greater than the congestion threshold, the switch sets the ECN field in the message header corresponding to the data flow in the speed-limited forwarding mode to congestion, generates a pseudo ACK packet, and notifies the source port TCP congestion control algorithm to reduce the packet sending window. All subsequent messages of the data flow are forwarded normally, and the operation of sending messages to the special queue is stopped until the length of the switch special queue is 0; B5. When the buffer queue length is not 0 and is less than the congestion threshold, it means that the source port is slowing down or the switch congestion is improved, but the data flow is still in a congested state. The subsequent message forwarding speed of the data flow in the speed-limited forwarding mode is halved according to B3 until the buffer queue length is greater than the congestion threshold again. The switch cyclically controls the data flow messages in the speed-limited forwarding mode until the data flow is released or not released; B6. If no packet loss information is detected in the SACK field and the ECN flag does not indicate congestion, the data flow is judged to be in release mode; B7. When the data flow is judged to be in release mode and there are messages cached in the special queue of the switch at the same time, the messages in the special queue of the switch are sent to the corresponding destination port with the highest priority.

7. The TCP acceleration method for a cross-data center network according to claim 6, characterized in that: The calculation method of the congestion coefficient in S3 includes: C=ΔRTT*(qDepth+1)*(RTT>>β+1) In the formula, C is the congestion coefficient, ΔRTT is the change in the round-trip delay of the data flow, qDepth is the accumulation length of the buffer queue corresponding to the data flow in the switch, RTT is the round-trip delay of the data flow, >> is a right shift operation, and β is the weight value representing RTT.

8. The TCP acceleration method for a cross-data center network according to claim 7, characterized in that: The processing method of the shallow buffer congested data flow and the deep buffer congested data flow in S3 includes: C1. At the ingress, the switch checks the ECN flag bit in the TCP packet header every N packets. The calculation method of interval N includes: N=RTT%T+1 Where RTT is the round trip delay of the data flow to which the data packet belongs, % is the modulo operation, and T is a constant, indicating an empirical parameter setting; C2. When the ECN value is 0b10 or 0b01 and the buffer queue length is less than the congestion threshold, the information table is retrieved to obtain the current round-trip delay change value of the data flow corresponding to the TCP message, and the congestion coefficient is calculated; C3. Reverse the source-destination quadruple obtained by parsing the TCP message header, swap the source IP address and the destination IP address, swap the source port and the destination port, form a new source-destination quadruple value and write it into the source-destination quadruple; C4. When the congestion coefficient C is less than T_C, T_C is the set threshold parameter, indicating that congestion occurs in the shallow buffer switch of the data center network. The round-trip delay change value of the data flow is calculated twice every N data packets, and the average value of the two round-trip delay change values ​​is calculated; C5. When the average value in C4 is less than T_C+M, M is an empirical parameter, indicating that the congestion level on the shallow cache switch of the data center network is low, and the source port does not use ECN speed reduction. The switch is used to identify the ACK message of the data flow, delete the ECN mark bit in the message header, and eliminate the perception of the ECN mark by the TCP congestion control algorithm of the source port; C6. In C4, the average value is greater than T_C+M, indicating that the congestion level on the shallow cache switch of the data center network is high, and the source port needs to be decelerated to retain the ECN mark bit on the ACK message of the data flow. The TCP congestion control algorithm of the source port adjusts the message sending rate through the ECN mark bit; C7. When the congestion coefficient is greater than T_C, it indicates that the deep buffer switch of the data center network is congested. The switch is used to identify the ACK message of the data flow. For all subsequent ACK data messages, the ECN mark bit is displayed as congestion. At the same time, the three latest consecutive ACK messages are selected, and their SACK fields are modified to indicate that the message corresponding to the current message header sequence number is lost. The source port TCP congestion control algorithm performs the corresponding packet loss and speed reduction operation.