A data processing and transmission method and related devices
By dynamically configuring the Flowlet detection time interval on the host side and combining network path delay feedback, the problem of difficulty in adapting to dynamic network load in the existing technology is solved, and more effective load balancing is achieved and the risk of data packet out of order is reduced.
Patent Information
- Application Number
- CN202080105542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-09-30
AI Technical Summary
In the prior art, the method of detecting Flowlets based on fixed time intervals is difficult to adapt to dynamically changing network loads, resulting in less significant load balancing effects or increasing the risk of data packets being out of order.
On the host side, combined with the delay feedback information of the network path, dynamically configures to detect the time interval of the partitioning of the Flowlets, so that the granularity of the Flowlet matches the state of the network path. The specific method includes generating a message segment, determining the target TCP stream to which it belongs, obtaining the target stream information, including the time threshold and the timestamp of adjacent message segments, and determining whether to divide it into the same Flowlet by comparing the time stamp difference value of the message segment with the time threshold.
By dynamically adjusting the detection time interval of the Flowlet, it can adapt to dynamic network load changes, improve the effect of load balancing, and reduce the risk of out-of-order data packets.
Smart Images

Figure CN116325708B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a data processing and transmission method and related devices. Background Art
[0002] Currently, with the high-speed and real-time requirements of data transmission services, data transmission devices are required to perform load balancing of traffic quickly and accurately to improve their forwarding performance and enhance network reliability, so as to better serve users.
[0003] Currently, Equal-cost multi-path routing (ECMP) is a relatively common load balancing processing method. ECMP technology includes a path selection method based on packets and a path selection method based on flows. Among them, the path selection method based on packets can achieve load balancing. There are differences in the delays of different paths in the multi-path, resulting in out-of-order packets received at the receiving end, and packet reordering is required; in the path selection method based on flows, the outgoing interface (packet forwarding path) for forwarding packets can be determined according to the hash algorithm, and packet reordering is not required at the receiving end. However, the rates of different flows may vary (for example, large flows (ElephantFlow) that occupy a large bandwidth and small flows (Mice Flow) that occupy a small bandwidth), and the flows transmitted in different paths are also different. When the rates of the flows transmitted in different paths are not equal, load imbalance may occur.
[0004] In order to achieve better load balancing, a load balancing processing method based on the flowlet mechanism has been proposed. In the load balancing processing method based on the flowlet mechanism, as Figure 1 shown Figure 1Schematic diagram of dividing a TCP flow in the prior art into Flowlets. For example, after the data packets of the current TCP flow arrive at the switch, they are detected to contain 5 Flowlets. The time difference between the arrival times of the data packets within the same Flowlet at the switch is generally small, while the time difference between the arrival times of the data packets between different Flowlets is relatively obvious. Among them, the time difference (timegap) between the first data packet and the second data packet of Flowlet 1 is less than the established threshold (timeout), so the switch regards these two data packets as the same Flowlet. For another example, the time difference between the last data packet of Flowlet 1 and the first data packet of Flowlet 2 is greater than the established threshold, so the switch regards the first data packet of Flowlet 2 as a new Flowlet. That is, if the time difference between two adjacent data packets of the same TCP flow arriving at the switch is less than the established time interval (timeout), the switch regards these two data packets as the same Flowlet.
[0005] In summary, the existing load balancing scheme with Flowlet granularity detects Flowlets based on a fixed time interval at the switch. However, the load situation within the data transmission network (such as the data center network) changes dynamically and unpredictably, and the fixed time interval is difficult to adapt to the dynamically changing network load. When the time interval is too small, the number of detected Flowlets will increase, the processing granularity of load balancing is finer, and the risk of out-of-order data packets is likely to increase; when the time interval is too large, the number of detected Flowlets will decrease, the processing granularity of load balancing is too rough, and the effect of load balancing is not significant. Summary of the Invention
[0006] Embodiments of the present invention provide a data processing, transmission method and related devices to improve the efficiency and accuracy during data transmission.
[0007] In a first aspect, embodiments of the present invention provide a data processing method applied to a host, which may include:
[0008] Generate a first packet segment, and determine the target TCP flow to which the first packet segment belongs; obtain the timestamp of the first packet segment, and obtain target flow information matching the target TCP flow, where the target flow information includes the time threshold corresponding to the target TCP flow and the timestamp of a second packet segment in the target TCP flow; wherein, the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow, the time threshold is the difference between a first path delay and a second path delay, the first path delay is the delay of the uplink path with the largest delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the smallest delay in the multi-path set of the target TCP flow; compare the difference between the timestamp of the first packet segment and the timestamp of the second packet segment with the time threshold; according to the comparison result, determine whether to divide the first packet segment and the second packet segment into the same Flowlet.
[0009] In the embodiment of the present invention, on the host side, for the currently generated first packet segment to be sent, first determine which TCP flow the first packet segment specifically belongs to (for example, determined by the source port), and then obtain the target flow information matching the TCP flow. Among them, the target flow information contains various information of the target TCP flow, such as the timestamp of the previous adjacent packet segment (i.e., the second packet segment), and the time threshold for dividing Flowlets in the target TCP flow; further, the host side compares the difference between the timestamp value between the currently to-be-sent first packet segment and the previous adjacent second packet segment with the time threshold, so as to decide whether to divide the first packet segment and the previous adjacent packet segment into the same Flowlet; and the time threshold among them is dynamically calculated by the delays of the paths in the multi-path set of the target TCP flow. For example, it is calculated based on the difference between the maximum delay and the minimum delay updated in real time according to the historical packet segments received by the host (ACK packets with the same triple or quintuple information as the target TCP flow); that is, the time threshold is a value that changes and is dynamically adjusted in real time according to the network transmission load situation. That is to say, in the embodiment of the present invention, for different TCP flows or data packets of the same TCP flow in different states, the time threshold for dividing Flowlets is dynamically changed and is dynamically adjusted according to the real-time transmission delay of the data in the corresponding TCP flow. Therefore, it can always adapt to the dynamic network load change, avoiding the problem in the prior art that it is difficult to adapt to the dynamic network load caused by the switch side detecting Flowlets based on a fixed time interval. In summary, the embodiment of the present invention combines the delay feedback information of the network path on the host side, dynamically configures the time interval for detecting and dividing Flowlets, makes the Flowlet granularity match the network path state, reduces the risk of packet out-of-order, and ensures the effect of load balancing.
[0010] In a possible implementation, determining the target TCP flow to which the first packet segment belongs includes: determining the target TCP flow to which the first packet segment belongs according to the source port number of the first packet segment.
[0011] In the embodiment of the present invention, by identifying the source port number in the five-tuple information of the packet segment, it is identified which TCP flow the packet segment to be currently sent (i.e., the first packet segment) belongs to, so as to further obtain the flow information of the TCP flow to which the packet segment belongs (including the timestamp value of the previous adjacent packet segment and the time threshold for dividing Flowlet), so as to further determine whether the currently to-be-sent data packet segment and the previous data packet segment belong to the same Flowlet or divide it into a new Flowlet based on the relevant information in the flow information.
[0012] In a possible implementation, the host maintains a flow information table, and the flow information table includes the flow information of N TCP flows, where N is an integer greater than or equal to 1. Each TCP flow's flow information includes the flow index of the corresponding TCP flow; obtaining the target flow information matching the target TCP flow includes: searching for the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow.
[0013] In the embodiment of the present invention, the host side maintains a flow information table of one or more TCP flows (such as currently active TCP flows). The flow information table includes the flow information of one or more TCP flows, and each TCP flow information can further include the index of the TCP flow, as well as the time threshold involved in the first aspect and the timestamp of the previous latest packet segment of the currently to-be-sent data packet segment. That is, the host can maintain the flow information of all currently active TCP flows, so that when a packet segment needs to be sent, the target flow information matching the flow index of the TCP flow to which the packet segment belongs can be found from the flow information table, so as to perform subsequent Flowlet division.
[0014] In a possible implementation, determining whether to divide the first packet segment and the second packet segment into the same Flowlet according to the comparison result includes: if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, dividing the first packet segment and the second packet segment into the same Flowlet; if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, dividing the first packet segment into a new Flowlet.
[0015] In an embodiment of the present invention, if the difference between the timestamps of the current first packet segment to be sent and the previous adjacent second packet segment in the target TCP flow to which it belongs is less than the time threshold corresponding to the target TCP flow (the time threshold is dynamically variable), it is considered that the condition for sending the first packet segment and the previous adjacent second packet segment in the same Flowlet is met, that is, the first packet segment can be determined to be divided into the same Flowlet as the previous second packet segment; similarly, if the difference between the timestamps of the current first packet segment to be sent and the previous adjacent second packet segment in the target TCP flow to which it belongs is greater than the time threshold corresponding to the target TCP flow, it is considered that the condition for sending the first packet segment and the previous adjacent second packet segment in the same Flowlet is not met, that is, the first packet segment is divided into a new Flowlet.
[0016] In a possible implementation manner, the target flow information further includes a reference Flowlet identifier of the target TCP flow, and the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment; the method further includes: generating a first data packet, where the first data packet includes the first packet segment and the Flowlet identifier of the first packet segment; where, if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier; if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
[0017] In an embodiment of the present invention, when further encapsulating the first packet segment to transmit data over a network, a Flowlet identifier corresponding to the packet segment can be set during the encapsulation process, so that after the packet segment is encapsulated into a data packet, the switch side can identify which Flowlet the data packet belongs to through the Flowlet identifier, and thus determine which path to use for transmission. For example, when the first packet segment is to enter the data link layer where the switch is located, the first packet segment needs to be further encapsulated. At this time, an identifier bit for the switch to identify which Flowlet the packet segment belongs to is set in the encapsulated data packet. When the Flowlet identifiers of the first packet segment and the second packet segment are the same, the first data packet corresponding to the first packet segment and the second data packet corresponding to the second packet segment are forwarded through the same path on the switch side. In summary, in the embodiment of the present invention, Flowlets are partitioned on the host side, and the bit in the reserved field of the transport layer header (for example, 1 bit) can be used to mark the Flowlet. The switch can identify the Flowlet only relying on the header field, with high efficiency and low hardware overhead. At the same time, it is ensured that the same Flowlet will not be split again no matter how many hops of switches it experiences in the network, reducing the risk of out-of-order data packets.
[0018] In a possible implementation manner, the method further includes: if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, updating the reference Flowlet identifier to the second Flowlet identifier.
[0019] In the embodiments of the present invention, the flow information of each TCP flow maintained on the host side further includes the reference Flowlet identifier of each TCP flow, that is, the identifier of the current Flowlet of each TCP flow is maintained in the flow information table, so as to set the corresponding Flowlet identifier for the packet segment to be sent. For example, assuming that the reference Flowlet identifier is the first Flowlet identifier (that is, the Flowlet identifier corresponding to the second packet segment), then when the first packet segment and the second packet segment are divided into the same Flowlet, the Flowlet identifier of the first packet segment is also marked as the first Flowlet identifier, that is, the reference Flowlet identifier remains unchanged as the first Flowlet identifier; if the reference Flowlet identifier is the first Flowlet identifier, and when the first packet segment and the second packet segment are divided into different Flowlets (that is, the first packet segment is divided into a new Flowlet), then the Flowlet identifier of the first packet segment is marked as the second Flowlet identifier, and at this time the reference Flowlet identifier needs to be updated to the second Flowlet identifier. Optionally, the reference Flowlet identifier can be switched between 0 or 1, that is, between two adjacent Flowlets, the Flowlet identifiers take values at intervals between 0 or 1, so that only 1 bit can accurately indicate whether different data packets belong to the same Flowlet.
[0020] In a possible implementation manner, the method further includes: receiving a target ACK packet, where the target ACK packet is an ACK packet with the same destination port number or the same destination address as that of the target TCP flow; determining the uplink path delay of the target ACK packet, where the uplink path delay is the difference between the timestamp value of the target ACK packet and the timestamp echo reply value; comparing the uplink path delay of the target ACK packet with the delay of the uplink path in the multi-path set of the target TCP flow; if the uplink path delay of the target ACK packet is greater than the delay of the uplink path with the largest current delay in the multi-path set, then updating the first path delay to the uplink path delay of the target ACK packet; if the uplink path delay of the target ACK packet is less than the delay of the uplink path with the smallest current delay in the multi-path set, then updating the second path delay to the uplink path delay of the target ACK packet.
[0021] In the embodiments of the present invention, the time threshold for dividing Flowlets can be calculated from the difference between the maximum uplink path delay and the minimum uplink path delay of the historical packet segments received in the target TCP flow (or the TCP flows in the same network session as the target TCP flow). That is, the time threshold is a value that changes and is dynamically adjusted in real time according to the network transmission load. Specifically, each time the host side receives an ACK packet belonging to the target TCP flow (i.e., the destination port numbers are the same), or receives an ACK packet of a TCP flow in the same network session as the target TCP flow (i.e., the destination addresses are the same or the destination network segments are the same), the transmission delay of the uplink path of the ACK packet is calculated by the difference between the timestamp value of the target ACK packet and the timestamp echo reply value. Based on the historical values of the uplink path transmission delays of all the received ACK packets, a current minimum uplink path delay is determined and used as the first path delay, and a current maximum uplink path delay is determined and used as the second path delay. Finally, the time threshold for dividing Flowlets in the target TCP flow is calculated using the difference between the first path delay and the second path delay. Thus, for different TCP flows or the data of the same TCP flow in different states, the time threshold for dividing Flowlets is dynamically changing and is dynamically adjusted according to the real-time transmission delay of the data in the corresponding TCP flow. Therefore, it can always adapt to the dynamic network load changes.
[0022] In a possible implementation, the multi-path set includes multiple equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple equivalent transmission paths and non-equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple non-equivalent transmission paths of the target TCP flow.
[0023] In the embodiments of the present invention, when the network accessed by the host is an equivalent multi-path model, the multiple transmission paths in the multi-path set of the target TCP flow are all equivalent paths. At this time, the first path delay is the delay of the uplink path with the maximum delay among these equivalent paths, and the second path delay is the delay of the uplink path with the minimum delay among these equivalent paths. When the network accessed by the host is a conventional multi-path model, the multiple transmission paths in the multi-path set of the target TCP flow may include equivalent paths or non-equivalent paths. At this time, the first path delay is the delay of the uplink path with the maximum delay among these equivalent or non-equivalent paths, and the second path delay is the delay of the uplink path with the minimum delay among these equivalent or non-equivalent paths. In summary, whether the multiple paths in the above multi-path set are equivalent depends on the type of the network topology structure of the network accessed by the host. The embodiments of the present invention can be applied to all network types with multi-path transmission.
[0024] Second aspect, an embodiment of the present invention provides a data transmission method, which is applied to a switch and may include:
[0025] Receiving a first data packet, where the first data packet includes a first packet segment and a Flowlet identifier of the first packet segment; determining a target TCP flow to which the first data packet belongs, and obtaining forwarding information matching the target TCP flow; the forwarding information includes a reference Flowlet identifier of the target TCP flow and a reference forwarding path; wherein, the reference Flowlet identifier is currently the first Flowlet identifier corresponding to a second packet segment, and the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow; the reference forwarding path is the first forwarding path of a second data packet, and the second data packet includes the second packet segment and the first Flowlet identifier; comparing the Flowlet identifier of the first packet segment with the first Flowlet identifier; and determining whether to forward the first packet segment through the first forwarding path according to the comparison result.
[0026] In an embodiment of the present invention, after receiving a data packet on the switch side, by identifying the Flowlet identifier in the data packet, and based on this Flowlet identifier, it is determined whether the Flowlet identifier of the first data packet is the same as that of the adjacent data packet in the target TCP flow to which it belongs, and based on this, it is decided whether the first data packet needs to be forwarded through the forwarding path corresponding to the second data packet. That is, on the switch side, it is not necessary to divide data packets into Flowlets according to the time interval of received data packets, but directly identify whether the currently to-be-sent data packet belongs to the same Flowlet as the previous adjacent data packet in the same TCP flow according to the Flowlet identifier bits included in the received data packet, so as to decide whether to continue forwarding through the forwarding path of the adjacent data packet, or divide a new Flowlet for the data packet and determine a new forwarding path for it.
[0027] In a possible implementation manner, the switch maintains a forwarding information table, and the forwarding information table includes forwarding information of M TCP flows, where M is an integer greater than or equal to 1, and the forwarding information of each TCP flow includes a five-tuple hash value of the corresponding TCP flow; the determining the target TCP flow to which the first data packet belongs and obtaining the forwarding information matching the target TCP flow includes: calculating a five-tuple hash value of the first data packet according to the five-tuple information of the first data packet; and searching for the forwarding information matching the target TCP flow from the forwarding information table according to the five-tuple hash value of the first data packet.
[0028] In an embodiment of the present invention, the switch maintains a forwarding information table, which includes forwarding information of one or more TCP flows (e.g., currently active TCP flows) on the hosts connected thereto, and the forwarding information of each TCP flow may include the five-tuple hash value of the TCP flow. That is, the switch can maintain the forwarding information of all currently active TCP flows, so that when a data packet needs to be sent, the forwarding information (including reference Flowlet identifier, forwarding path, etc.) matching the five-tuple hash value of the data packet can be found in the forwarding information table according to the five-tuple hash value of the data packet, thereby forwarding the data packet to be sent.
[0029] In a possible implementation manner, determining whether to forward the first packet segment through the forwarding path according to the comparison result includes: if the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, forwarding the first data packet through the first forwarding path; if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, determining a second forwarding path for the first data packet and forwarding it through the second forwarding path.
[0030] In an embodiment of the present invention, when the switch identifies that the Flowlet identifiers of the first data packet and the adjacent data packet in the target TCP flow to which the first data packet belongs are the same, the first data packet and the second data packet are forwarded on the same path; when the switch identifies that the Flowlet identifiers of the first data packet and the adjacent data packet in the target TCP flow to which the first data packet belongs are different, a new forwarding path is determined for the first data packet and it is forwarded through the new forwarding path. It should be noted that the second forwarding path may be the same as or different from the first forwarding path, depending on the decision of the switch.
[0031] In a possible implementation manner, the method further includes: if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is the second Flowlet identifier, updating the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and updating the reference forwarding path to the second forwarding path.
[0032] In an embodiment of the present invention, when the Flowlet identifiers of the first data packet and the second data packet are different, it indicates that the first data packet and the previous adjacent second data packet in the TCP flow to which the first data packet belongs do not belong to the same Flowlet. Therefore, the switch needs to divide the first data packet into a new Flowlet and update the reference Flowlet identifier of the TCP flow to which the first data packet belongs to the Flowlet identifier corresponding to the current latest data packet, that is, the second Flowlet identifier.
[0033] In a third aspect, an embodiment of the present invention provides a data processing device, which may include:
[0034] A first generating unit, configured to generate a first packet segment;
[0035] A first determining unit, configured to determine a target TCP flow to which the first packet segment belongs;
[0036] An obtaining unit, configured to obtain a timestamp of the first packet segment, and obtain target flow information matching the target TCP flow, where the target flow information includes a time threshold corresponding to the target TCP flow, and a timestamp of a second packet segment in the target TCP flow; wherein, the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow, the time threshold is the difference between a first path delay and a second path delay, the first path delay is the delay of the uplink path with the largest delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the smallest delay in the multi-path set of the target TCP flow;
[0037] A first comparing unit, configured to compare the difference between the timestamp of the first packet segment and the timestamp of the second packet segment with the time threshold;
[0038] A Flowlet partitioning unit, configured to determine whether to partition the first packet segment and the second packet segment into the same Flowlet according to the comparison result.
[0039] In a possible implementation manner, the first determining unit is specifically configured to:
[0040] Determine the target TCP flow to which the first packet segment belongs according to the source port number of the first packet segment.
[0041] In a possible implementation manner, the device further includes:
[0042] A maintenance unit, configured to maintain a flow information table, where the flow information table includes flow information of N TCP flows, N is an integer greater than or equal to 1, and the flow information of each TCP flow includes a flow index corresponding to the TCP flow;
[0043] The obtaining unit is specifically configured to: look up the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow.
[0044] In a possible implementation manner, the Flowlet partitioning unit is specifically configured to:
[0045] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the first packet segment and the second packet segment are classified into the same Flowlet;
[0046] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the first packet segment is classified into a new Flowlet.
[0047] In a possible implementation, the target flow information further includes a reference Flowlet identifier of the target TCP flow, and the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment; the apparatus further includes:
[0048] A second generating unit, configured to generate a first data packet, where the first data packet includes the first packet segment and the Flowlet identifier of the first packet segment; where
[0049] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier;
[0050] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
[0051] In a possible implementation, the apparatus further includes:
[0052] A first updating unit, configured to update the reference Flowlet identifier to the second Flowlet identifier if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold.
[0053] In a possible implementation, the apparatus further includes:
[0054] A receiving unit, configured to receive a target ACK packet, where the target ACK packet is an ACK packet with the same destination port number or the same destination address as that of the target TCP flow;
[0055] A second determining unit, configured to determine an uplink path delay of the target ACK packet, where the uplink path delay is a difference between a timestamp value of the target ACK packet and a timestamp echo reply value;
[0056] A second comparing unit, configured to compare the uplink path delay of the target ACK packet with the delay of the uplink path in the multi-path set of the target TCP flow;
[0057] A second update unit, configured to update the first path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is greater than the delay of the uplink path with the maximum current delay in the multipath set;
[0058] A third update unit, configured to update the second path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is less than the delay of the uplink path with the minimum current delay in the multipath set.
[0059] In a possible implementation, the multipath set includes multiple equivalent transmission paths of the target TCP flow; or, the multipath set includes multiple equivalent transmission paths and non-equivalent transmission paths of the target TCP flow; or, the multipath set includes multiple non-equivalent transmission paths of the target TCP flow.
[0060] Fourthly, an embodiment of the present invention provides a data processing apparatus, which may include:
[0061] A receiving unit, configured to receive a first data packet, where the first data packet includes a first segment and a Flowlet identifier of the first segment, and the first data packet belongs to a target TCP flow;
[0062] A determining unit, configured to determine the target TCP flow to which the first data packet belongs, and obtain forwarding information matching the target TCP flow; the forwarding information includes a reference Flowlet identifier of the target TCP flow and a reference forwarding path; where the reference Flowlet identifier is currently the first Flowlet identifier corresponding to a second segment, and the second segment is the previous segment adjacent to the first segment in the target TCP flow; the reference forwarding path is the first forwarding path of a second data packet, and the second data packet includes the second segment and the first Flowlet identifier;
[0063] A comparing unit, configured to compare the Flowlet identifier of the first segment with the first Flowlet identifier;
[0064] A forwarding unit, configured to determine whether to forward the first segment through the first forwarding path according to the comparison result.
[0065] In a possible implementation, the switch maintains a forwarding information table, and the forwarding information table includes the forwarding information of M TCP flows, where M is an integer greater than or equal to 1, and the forwarding information of each TCP flow includes the five-tuple hash value of the corresponding TCP flow; the determining unit is specifically configured to:
[0066] Calculate the five-tuple hash value of the first data packet according to the five-tuple information of the first data packet;
[0067] Look up the forwarding information matching the target TCP flow in the forwarding information table according to the five-tuple hash value of the first data packet.
[0068] In a possible implementation, the forwarding unit is specifically configured to:
[0069] If the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, forward the first data packet through the first forwarding path;
[0070] If the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, determine a second forwarding path for the first data packet and forward it through the second forwarding path.
[0071] In a possible implementation, the apparatus further includes:
[0072] An update unit, if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is a second Flowlet identifier, update the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and update the reference forwarding path to the second forwarding path.
[0073] In a fifth aspect, the present application provides a semiconductor chip, which may include the data processing apparatus provided in any one of the implementations in the third aspect above.
[0074] In a sixth aspect, the present application provides a semiconductor chip, which may include the data processing apparatus provided in any one of the implementations in the fourth aspect above.
[0075] In a seventh aspect, the present application provides a semiconductor chip, which may include: the data processing apparatus provided in any one of the implementations in the third aspect above, an internal memory coupled to the data processing apparatus, and an external memory.
[0076] In an eighth aspect, the present application provides a semiconductor chip, which may include: the data transmission apparatus provided in any one of the implementations in the fourth aspect above, an internal memory coupled to the data processing apparatus, and an external memory.
[0077] In a ninth aspect, the present application provides a system-on-chip (SoC) chip, which includes a data processing device provided by any one of the implementation manners in the above third aspect, an internal memory and an external memory coupled to the data processing device. The SoC chip may be composed of chips or may include chips and other discrete devices.
[0078] In a tenth aspect, the present application provides a system-on-chip (SoC) chip, which includes a data transmission device provided by any one of the implementation manners in the above fourth aspect, an internal memory and an external memory coupled to the data transmission device. The SoC chip may be composed of chips or may include chips and other discrete devices.
[0079] In an eleventh aspect, the present application provides a chip system, which includes a data processing device provided by any one of the implementation manners in the above third aspect. In a possible design, the chip system further includes a memory for storing program instructions and data necessary or related to the operation of the data processing device. The chip system may be composed of chips or may include chips and other discrete devices.
[0080] In a twelfth aspect, the present application provides a chip system, which includes a data transmission device provided by any one of the implementation manners in the above fourth aspect. In a possible design, the chip system further includes a memory for storing program instructions and data necessary or related to the operation of the data transmission device. The chip system may be composed of chips or may include chips and other discrete devices.
[0081] In a thirteenth aspect, the present application provides a data processing device, which has the function of implementing any one of the data processing methods in the above first aspect. This function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above function.
[0082] In a fourteenth aspect, the present application provides a data transmission device, which has the function of implementing any one of the data transmission methods in the above second aspect. This function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above function.
[0083] Fifteenth aspect, the present application provides a host, which includes a processor for executing the data processing method provided by any one of the implementations in the above first aspect. The host may further include a memory for coupling with the processor and storing necessary program instructions and data of the host. The host may further include a communication interface for the host to communicate with other devices or communication networks.
[0084] Sixteenth aspect, the present application provides a switch, which includes a processor for executing the data transmission method provided by any one of the implementations in the above first aspect. The switch may further include a memory for coupling with the processor and storing necessary program instructions and data of the switch. The switch may further include a communication interface for the switch to communicate with other devices or communication networks.
[0085] Seventeenth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a host, implements the processing method flow of the multi-core processor described in any one of the above second aspects.
[0086] Eighteenth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a switch, implements the processing method flow of the multi-core processor described in any one of the above fourth aspects.
[0087] Nineteenth aspect, an embodiment of the present invention provides a computer program including instructions, which, when executed by a multi-core processor, enable a host to execute the processing method flow of the multi-core processor described in any one of the above second aspects.
[0088] Twentieth aspect, an embodiment of the present invention provides a computer program including instructions, which, when executed by a multi-core processor, enable a switch to execute the processing method flow of the multi-core processor described in any one of the above fourth aspects. Description of the Drawings
[0089] Figure 1 Schematic diagram of dividing a TCP flow into Flowlets in the prior art.
[0090] Figure 2 Schematic diagram of a network transmission system architecture provided by an embodiment of the present application.
[0091] Figure 3 Schematic diagram of a data center network topology structure provided by an embodiment of the present application.
[0092] Figure 4It is a schematic diagram of a computer network OSI model and a TCP / IP model provided by an embodiment of the present application.
[0093] Figure 5 It is a schematic flowchart of a data transmission method provided by an embodiment of the present invention.
[0094] Figure 6A It is a schematic diagram of a first data packet and a second data packet in the same Flowlet provided by an embodiment of the present invention.
[0095] Figure 6B It is a schematic diagram of a first data packet and a second data packet in different Flowlets provided by an embodiment of the present invention.
[0096] Figure 6C It is a schematic flowchart of a process for an additional layer protocol to divide and label Flowlets provided by an embodiment of the present invention.
[0097] Figure 6D It is a schematic flowchart of a process for an additional layer protocol to dynamically update the splitting threshold of Flowlets provided by an embodiment of the present invention.
[0098] Figure 7 It is a schematic flowchart of a data transmission method provided by an embodiment of the present invention.
[0099] Figure 8 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present invention.
[0100] Figure 9 It is a schematic diagram of the structure of a data transmission device provided by an embodiment of the present invention. Detailed implementation manners
[0101] The embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention. Terms such as "first", "second", "third", and "fourth" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices. The mention of "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0102] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components can reside in a process and / or an execution thread, and the components can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media on which various data structures are stored. Components can communicate, for example, through local and / or remote processes according to signals having one or more data packets (such as data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems through signals).
[0103] First, some terms in this application are explained to facilitate the understanding of those skilled in the art.
[0104] (1) Equal-Cost Multipath Routing (ECMP) means that there are multiple paths with the same cost to reach the same destination address. Here, the same cost means that the number of hops (i.e., the number) of the switches passed through is the same. For example, in the network topology of the equal-cost multipath model (such as a data center network), all possible transmission paths between the same pair of source and destination hosts are equal-cost paths. When a device supports equal-cost routing, the traffic destined for this destination IP or destination network segment can be shared through different paths to achieve network load balancing. And when some of these paths fail, other paths can replace them to complete the forwarding process, realizing the routing redundancy backup function. If traditional routing technology is used, the data packets destined for this destination address can only utilize one of the links, and the other links are in a backup state or invalid state, and the mutual switching in a dynamic routing environment takes a certain amount of time. However, the equal-cost multipath routing protocol can use multiple links simultaneously in this network environment, not only increasing the transmission bandwidth but also backing up the data transmission of the failed link without delay and packet loss.
[0105] (2) Transmission Control Protocol (TCP) is a connection-oriented, reliable, byte-stream-based transport layer communication protocol. TCP aims to adapt to the hierarchical protocol architecture that supports multiple network applications. Reliable communication services are provided between pairs of processes in the main computers connected to different but interconnected computer communication networks relying on TCP. TCP assumes that it can obtain a simple, possibly unreliable datagram service from a lower-level protocol. In principle, TCP should be able to operate over various communication systems ranging from hard-wired connections to packet-switching or circuit-switching networks.
[0106] (3) A network flow (Flow), also simply referred to as a net flow, is a set of data packets with the same five-tuple within a certain period of time. The five-tuple includes the source IP address, source port number, destination IP address, destination port number of the communication parties, and the transport layer protocol.
[0107] (4) A network session is a collection of multiple network flows, and multiple network flows have the same triple (source address, destination address, transport layer protocol).
[0108] (5) Flowlet, which can be understood as a packet group composed of multiple packets continuously sent in a Flow. Each Flow includes multiple Flowlets. When forwarding packets based on the Flowlet mechanism, the forwarding of multiple packets included in a Flowlet can be implemented based on Flowlet flow table entries. Different Flowlets correspond to different Flowlet flow table entries. The Flowlet flow table entries are used to indicate the packet forwarding paths of multiple packets included in each Flowlet.
[0109] (6) Transmission Control Protocol / Internet Protocol (TCP / IP) refers to a protocol suite that can implement information transmission among multiple different networks. The TCP / IP protocol does not merely refer to the TCP and IP protocols, but rather refers to a protocol suite composed of protocols such as FTP, SMTP, TCP, UDP, and IP. Only because the TCP protocol and the IP protocol are the most representative in the TCP / IP protocol, it is called the TCP / IP protocol. Among them, TCP is a connection-oriented, reliable, byte-stream-based transport layer communication protocol.
[0110] (7) Internet Service Provider (ISP) network, that is, a telecommunications operator that comprehensively provides Internet access services, information services, and value-added services to the general public.
[0111] To facilitate the understanding of the embodiments of the present application, the network transmission system architecture on which the embodiments of the present application are based will be described first below. Figure 2 It is a schematic diagram of a network transmission system architecture provided by an embodiment of the present application. Please refer to Figure 2 , this network transmission system architecture mainly includes: host 10, switch (SWITCH) 20, and the Internet. The host 10 can be further divided into a source host or a destination host according to whether it is a sending end or a receiving end. The source host can be connected to the Internet through the switch 20 to communicate with the destination host.
[0112] The host 10 can be any computing device that generates data and has network access capabilities. For example, any computer connected to the Internet can be called a host, and each host has a unique IP address. Among them, the host 10 can specifically be various devices such as servers, personal computers, tablets, mobile phones, personal digital assistants, smart wearable devices, driverless terminals, etc. When two hosts (such as the source host and the destination host) need to communicate and transfer data, the source host needs to encapsulate the application data into data packets (such as TCP / IP packets), and then hand them over to the next layer, the data link layer (such as a switch), to continue encapsulating them into frames; then the switch, etc., transfers the data from the source host to the destination host accurately according to the MAC address. In the embodiments of the present invention, the host 10 also has functions such as dividing Flowlets for the message segments, Flowlet identification, and dynamically configuring the time threshold for dividing Flowlets. For specific details, please refer to the description of the subsequent related embodiments, and details will not be elaborated here.
[0113] The switch 20 is a network device that identifies based on the MAC (hardware address of the network card) and completes the function of encapsulating and forwarding data packets. It can "learn" the MAC address and store it in the internal address table, and establish a temporary switching path between the sending end and the receiving end of the data frame, so that the data frame directly reaches the destination address from the source address. The functions of the switch 20 can include physical addressing, network topology structure, error checking, frame sequence, and flow control, etc. In the embodiments of the present invention, the switch 20 also has functions such as dividing Flowlets for the message segments and Flowlet identification on the host 10 side, and then based on the already divided and identified Flowlet identification, it forwards the same Flowlets or different Flowlets in the data stream. For specific details, please refer to the description of the subsequent related embodiments, and details will not be elaborated here.
[0114] For example, the source host 10 uses, such as the Transmission Control Protocol (TCP), and processes the data through the data processing method in this application, and then sends a message to the message forwarding device in the routing and switching network. The message forwarding devices (such as switches, routers, etc.) in the routing and switching network use the ECMP technology and forward the message through the data transmission method in this application, and finally forward it to the destination host 10, thereby achieving the effect of load balancing processing.
[0115] The data processing method or data transmission method in the embodiments of the present invention can be applicable to a transmission mechanism based on TCP / IP. The application scope of the data transmission method in the present invention is not limited to the data center network, but also applicable to any network with multiple paths, such as an ISP network (Internet Service Provider), where the network topology provides multiple network paths for any two source-destination communication nodes (i.e., source host and destination host). Therefore, the technical solution in this application can be applied to perform dynamic load balancing at the Flowlet granularity.
[0116] It should be noted that for the data center network, the characteristics of its network topology structure determine that in the same network session of the data center network, that is, the TCP flows with the same triple (source address, destination address, transport layer protocol) information, one or more paths included in the corresponding multi-path set are all equivalent paths; while for other types of networks, for example, in the same network session of the ISP network, one or more paths included in the corresponding multi-path set may be equivalent or may not be equivalent. Therefore, depending on the type of network topology structure to which the host is connected, one or more paths included in the multi-path set of the target TCP flow described in this application can be equivalent or can be non-equivalent.
[0117] It can be understood that the above Figure 2 network architecture is only an exemplary implementation manner in the embodiments of this application. The network architecture in the embodiments of the present invention includes but is not limited to the above network architecture.
[0118] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the topology structure of a data center network provided by the embodiments of this application. The data center network mainly includes: a Core core layer, an Aggregation aggregation layer, an Access access layer, and a POD aggregation area layer. The source host can communicate with the destination host through the switch and the core network via the TCP protocol. Among them,
[0119] The aggregation area (Point of delivery, POD) layer consists of multiple PODs, and each POD can include servers, storage, and network devices. Among them, the Top of Rack (ToR) is a way of wiring the server cabinets in the data center. When wiring in the TOR way, 1 to 2 access switches are deployed at the upper end of each server cabinet.
[0120] Access Layer: Physically connects to servers, usually placed at the top of the cabinet, also known as ToR switches, or Edge Access Layer (Edge Layer). Access switches are usually located at the top of the rack, so they are also called ToR (Top of Rack) switches, and they physically connect to servers.
[0121] Aggregation Layer: Aggregation switches that aggregate and connect access switches, and also provide other services such as Fire Wall (FW), Server Load Balancer (SLB), Secure Sockets Layer offload (SSL offload), intrusion detection, network analysis, etc.
[0122] Core Layer: Core switches that provide high-speed forwarding and connectivity to multiple aggregation layers. Core switches provide high-speed forwarding for packets entering and leaving the data center, and connectivity to multiple aggregation layers. Core switches usually provide a flexible L3 routing network for the entire network.
[0123] For example, in Figure 3 , for TOR1 in Pod1, it has at least 4 equivalent paths (ECMP) to access the Internet, as shown in Figure 3 , TOR1 in Pod1 can access the Internet through at least the equivalent paths: Path 1, Path 2, Path 3, and Path 4.
[0124] It should be noted that in the embodiments of the present invention, the data transmission method applied to the switch side can be applied to the switches of the above layers (Access Layer, Aggregation Layer, or Core Layer), that is, in the entire forwarding path of the data packet from the source host to the destination host, all switches participating in the forwarding can implement any one of the data transmission methods provided in this application.
[0125] It can be understood that the above Figure 3 data center network topology structure is only an exemplary implementation manner in the embodiments of this application. The data center network topology structure in the embodiments of the present invention includes but is not limited to the above network architectures.
[0126] Please refer to Figure 4 , Figure 4It is a schematic diagram of a computer network OSI model and a TCP / IP model provided by an embodiment of the present application. In the embodiment of the present application, in the existing computer network OSI model or TCP / IP model, an additional layer is added between the transport layer and the network layer. This additional layer is mainly used for dividing, marking, and setting relevant parameters of Flowlets in the TCP stream. Specifically, the OSI eight-layer network model provided by the embodiment of the present invention consists of layers 1 to 8 from bottom to top, namely the physical layer, data link layer, network layer, additional layer, transport layer, session layer, presentation layer, and application layer; the TCP / IP model provided by the embodiment of the present invention can be simplified to layers 1 to 5 from bottom to top, mainly including the network interface layer, network layer, additional layer, transport layer, and application layer. Among them,
[0127] (1) Application layer
[0128] The layer closest to the user in the OSI reference model, which provides an application interface for computer users and also directly provides various network services for users. It provides rich system application interfaces for user application software. Common network service protocols in the application layer include: Hyper Text Transfer Protocol (HTTP), Hyper Text Transfer Protocol over Secure Socket Layer (HTTPS), File Transfer Protocol (FTP), Post Office Protocol-Version 3 (POP3), Simple Mail Transfer Protocol (SMTP), etc.
[0129] (2) Presentation layer
[0130] Responsible for data encoding and conversion to ensure the normal operation of the application layer. Perform data format conversion to ensure that the application layer data generated by one system can be recognized and understood by the application layer of another system. Computers on the network may use different data representations, so data format conversion is required during data transmission. In order for computers using different data representation methods to communicate with each other and exchange data, abstract data structures need to be used to represent the transmitted data during the communication process. However, the internal machine still uses its own standard encoding. Managing these abstract data structures, converting the internal encoding of the machine into a transfer syntax suitable for network transmission at the sender, and performing the reverse conversion at the receiver are all tasks completed by the presentation layer.
[0131] (3) Session layer
[0132] Responsible for establishing, maintaining, and controlling sessions, differentiating different sessions, and providing services in three communication modes: simplex, half duplex, and full duplex. For example, establish, manage, and terminate sessions between the two communication parties, and determine whether both parties should start a communication initiated by one party.
[0133] (4) Transport layer
[0134] Responsible for splitting and combining data to achieve end-to-end logical connections. The transport layer establishes end-to-end connections between hosts. The role of the transport layer is to provide reliable and transparent data transmission services from end to end for upper-layer protocols, including handling issues such as error control and flow control. This layer shields the details of lower-layer data communication from the upper layer, so that the upper-layer users only see a host-to-host, user-controllable and settable, reliable data path between two transport entities. TCP / UDP is at this layer.
[0135] (5) Additional layer
[0136] The additional layer in the embodiments of the present invention is used for Flowet partitioning and marking of TCP flows and setting of related parameters. The additional layer protocol dynamically configures the splitting threshold of Flowet according to the delay feedback of the network path, and divides the segments of the TCP flow into Flowets based on the dynamic splitting threshold. Since the partitioning of Flowets is completed on the host side, the present invention can use the 1-bit reserved field in the transport layer header to mark adjacent Flowets of the same TCP flow (the present invention names this 1-bit field FL_Tag), and transfer the partitioning result of Flowets to the in-network switch. The switch then identifies Flowets based on the header flag bits of the data packets. In the embodiments of the present invention, the functions on the host side mainly involve the above-mentioned application layer, presentation layer, session layer, transport layer, and additional layer.
[0137] It should be noted that the additional layer described in this application can be deployed as a single layer or deployed to the existing transport layer mentioned above. That is, the functions implemented by the additional layer are combined and implemented in the transport layer. The embodiments of the present invention do not make specific limitations in this regard.
[0138] (6) Network layer
[0139] Responsible for managing network addresses, locating devices, and determining routes. This layer establishes a connection between two nodes through IP addressing, selects appropriate routes and switching nodes for the packets sent by the transport layer at the source end, and correctly delivers them to the transport layer at the destination end according to the address. That is, it is usually the so-called IP protocol layer. Specifically, the network layer realizes the entire transmission process of data from any node to any other node according to the network layer address information contained in the data. That is, the main function is to complete the message transmission between hosts in the network and use the services of the data link layer to transmit each message from the source end to the destination end. The functions involved in the switch in the embodiments of the present invention correspond to this network layer.
[0140] (7) Data link layer
[0141] Responsible for preparing physical transmission, cyclic redundancy check (CRC), error notification, network topology, flow control, etc. Combine bits into bytes, then combine bytes into frames, use the link layer address (Ethernet uses MAC address) to access the medium, and perform error detection. Establish a logically meaningful data link between adjacent nodes connected by a physical link, and realize point-to-point or point-to-multipoint direct communication of data on the data link. In a wide area network, the data link layer is responsible for the reliable transmission of data between the host's interface message processor (IMP) and IMP-IMP. In a local area network, the data link layer is responsible for the reliable transmission of data between nodes.
[0142] (8) Physical layer
[0143] Complete the conversion of logically "0" and "1" to physical (optical / electrical signals) suitable for transmission medium bearing; realize the sending, receiving of physical signals, and the transmission process in the medium. The main function of the physical layer is to complete the transmission of the original bit stream between adjacent nodes. That is, it is responsible for sending and receiving data in the form of a bit stream. In fact, the final signal transmission is realized through the physical layer. Common transmission media of the physical layer include (various physical devices) hubs, repeaters, modems, network cables, twisted pairs, coaxial cables, etc.
[0144] It should be noted that in the simplified TCP / IP model, when application layer data is sent to the network through the protocol stack, each protocol layer adds a data header, which is called encapsulation. Different protocol layers have different names for data packets. For example, it is called a message at the application layer, a segment at the transport layer, a datagram or a packet at the network layer, and a frame at the link layer, etc.
[0145] It can be understood that the above Figure 4 related network models and functions are only an exemplary implementation in the embodiments of this application. The network models and functions involved in the embodiments of the present invention include but are not limited to the above models and functions.
[0146] First, in order to better understand the embodiments of the present invention, a further description of Flow and Flowlet involved in this application is given. As described above Figure 1 shown, a Flowlet is actually a micro-Flow. A Flow can be divided into many Flowlets. The same Flowlet has the same five-tuple information, that is, the source IP, destination IP, source port, destination port, and transport layer protocol are all the same. Multiple consecutive Packets sent in a certain Flow are regarded as a Flowlet, and the Flowlet mechanism is applied for path selection to forward the multiple Packets included in the Flowlet based on the selected path. In this application, for different data packets in the same Flowlet, the exact same forwarding path (excluding equivalent paths) is used for forwarding, while for different Flowlets in the same TCP Flow, different but equivalent paths (i.e., equivalent multi-paths) can be used for forwarding, or non-equivalent paths can be used for forwarding, depending on the type of the network topology structure of the network to which the host is connected. Among them, equivalent multi-paths can include paths with equal switch hops in the forwarding path of the data packet; while non-equivalent multi-paths refer to paths with unequal switch hops in the forwarding path of the data packet.
[0147] In other words, a Flow can be regarded as composed of multiple Flowlets. Load balancing introduces an intermediate layer based on Flowlets. It is neither a data packet nor a Flow, but a Flowlet that is larger than a packet and smaller than a Flow. That is, a Flowlet can be considered as a micro-Flow composed of one or more Packets in the same Flow.
[0148] Based on the above Figure 2 or Figure 3The provided network architecture, and Figure 4 The provided computer network model, combined with the data transmission method provided in this application, specifically analyzes and solves the technical problems proposed in this application.
[0149] See Figure 5 , Figure 5 is a schematic flowchart of a data transmission method provided by an embodiment of the present invention. This method can be applied to the above Figure 2 or Figure 3 in the described network architecture, where the host 10 can be used to support and execute Figure 5 the method flow steps S501 - step S504 shown in. The following will be described from the host 10 (source host) side in conjunction with the attached Figure 3 This method may include the following steps S501 - step S504. Optionally, it may further include steps S505 - step S506.
[0150] Step S501: Generate a first packet segment and determine the target TCP flow to which the first packet segment belongs.
[0151] Specifically, at the sending end, when a host (which can be called the source host) needs to send a message to another host (which can be called the destination host), the source host first generates a data packet that conforms to the relevant protocol standards locally, and then sends it to the destination host through a switch or the like. Among them, on the host side, the process of generating a data packet mainly involves the application layer (including the application layer, presentation layer, and session layer), the transport layer, and the network layer. For example, when a certain application in the source host needs to send a message to the destination host, after the message is encapsulated by the application layer on the source host side, it enters the transport layer and generates a packet segment that conforms to the transport layer protocol (i.e., the first packet segment), such as a TCP packet segment (segment) that conforms to the TCP protocol. That is, on the host side, when the TCP packet segment completes the encapsulation of the transport layer header fields (i.e., generates the first packet segment in the embodiment of the present invention), the function of the additional layer in this application (as Figure 2 described herein and will not be elaborated further) is triggered and subsequent Flowlet division and marking of the first packet segment, as well as dynamic configuration of the time threshold for dividing Flowlet, etc. Specifically, the source host first determines the TCP flow to which the first packet segment belongs through the additional layer, and then obtains the relevant information for dividing and identifying Flowlet corresponding to the first packet segment according to the TCP flow to which it belongs. In this application, a TCP flow (Flow) represents the data transmission process in a certain business process, that is, from TCP three-way handshake → data transmission end → connection release; and the five-tuple information of the same TCP flow is the same, where the five-tuple information includes source IP, destination IP, source port, destination port, and transport layer protocol.
[0152] Optionally, the host determines the target TCP flow to which the first packet segment belongs according to the source port number in the first packet segment. That is, the source port numbers corresponding to different TCP flows must be different. Therefore, it is possible to determine whether different packet segments belong to the same TCP flow through the source port in the packet segment. For example, the source host can determine which target TCP flow the first packet segment belongs to according to the source port information in the first packet segment.
[0153] Step S502: Obtain the timestamp of the first packet segment, and obtain the target flow information matching the target TCP flow.
[0154] Specifically, in this application, the division and identification of Flowlet for the first packet segment depend on the relationship between the difference in timestamps between the first packet segment and the previous adjacent packet segment and the corresponding time threshold. Therefore, it is necessary to obtain the relevant flow information for matching. After the source host determines the TCP flow to which the currently to-be-sent first packet segment belongs, it further obtains the timestamp of the first packet segment and the target flow information matching the target TCP flow to further perform subsequent division, identification, etc. of Flowlet. It should be noted that in the embodiments of the present invention, the timestamp of the packet segment is usually the time information added when encapsulating the packet segment in the transport layer, and this time information represents the moment when the packet segment is generated (for the destination host, it can also be understood as the moment when the source host sends the packet segment). For example, when the host needs to send data to the destination host, it will encapsulate the sending time into the timestamp item of the data. For both the source host and the destination host, they can both know the moment when the data is sent through this timestamp, so as to calculate (or measure) network latency, calculate service processing time consumption, etc. Optionally, the host can obtain the timestamp of the first packet segment by obtaining the timestamp value carried in the first packet segment, or can obtain the timestamp of the first packet segment according to the current system timestamp, and this obtaining step can be completed after generating the first packet segment and before step S503, and there is no limitation on the specific execution time point. The target flow information includes the time threshold corresponding to the target TCP flow and the timestamp of the second packet segment in the target TCP flow. Among them,
[0155] The second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow (that is, the packet segment in the target TCP flow whose timestamp value is earlier than that of the first packet segment and is adjacent to the first packet segment); the timestamp of the second packet segment can be obtained by the host according to the timestamp value carried in the second packet segment, or can be obtained by the host according to the moment recorded by the system at that time, that is, the timestamps of the first packet segment and the second packet segment are obtained using the same standard.
[0156] The time threshold is the difference between the first path delay and the second path delay. The first path delay is the delay of the uplink path with the maximum delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the minimum delay in the multi-path set of the target TCP flow. Among them, the multi-path set of the target TCP flow may include multiple transmission paths corresponding to the target TCP flow, that is, multiple possible uplink transmission paths between the source IP, source port, destination IP, and destination port of the target TCP flow. Optionally, the multi-path set of the target TCP flow may further include multiple transmission paths corresponding to the TCP flows that are in the same network session as the target TCP flow, that is, multiple possible uplink transmission paths between the source IP and the destination IP. In other words, the multi-path set of the target TCP flow may include multiple transmission paths corresponding to the target TCP flow itself, or may further include multiple transmission paths corresponding to the TCP flows with the same triple information (source address, destination address, transport layer protocol) as the target TCP flow. That is to say, it may be that one TCP flow corresponds to one multi-path set, or it may be that multiple TCP flows in the same network session correspond to the same multi-path set. Therefore, each TCP flow may maintain a time threshold for dividing Flowlets separately, or multiple TCP flows may jointly maintain a time threshold for dividing Flowlets. Among them, the uplink path refers to the path from the sending end (i.e., the source host) to the receiving end (destination host); and the uplink path delay refers to the total delay experienced between the time when the packet segment is sent from the source host and reaches the destination host, and the time when the source host receives the acknowledgment from the destination host (the destination host sends the acknowledgment immediately after receiving the data).
[0157] Further optionally, when the network accessed by the host is an Equal-Cost Multi-Path (ECMP) model, the multiple transmission paths in the multi-path set of the target TCP flow are all equal-cost paths. At this time, the first path delay is the delay of the uplink path with the maximum delay among these equal-cost paths, and the second path delay is the delay of the uplink path with the minimum delay among these equal-cost paths. When the network accessed by the host is a conventional multi-path model, the multiple transmission paths in the multi-path set of the target TCP flow may include equal-cost paths or non-equal-cost paths. At this time, the first path delay is the delay of the uplink path with the maximum delay among these equal-cost or non-equal-cost paths, and the second path delay is the delay of the uplink path with the minimum delay among these equal-cost or non-equal-cost paths. It should be noted that since the source addresses of the data packets sent from the source host are necessarily the same, if the destination host address, i.e., the destination address, is the same or the destination network segment is the same, the multi-path sets between different TCP flows between the sending end and the receiving end in the network of the equal-cost multi-path model (such as a data center network) are actually the same. Therefore, the uplink path delay of the historical packet segment with the same five-tuple or three-tuple as the target TCP flow can be used to calculate the time threshold.
[0158] For example, as Figure 3 shown in Figure 3 , since the network topology structure in Figure 3 is a data center network, which belongs to the equal-cost multi-path model, all paths in the multi-path set of the target TCP flow under this network are equal-cost paths, such as the paths 1, 2, 3, and
[0159] In a possible implementation manner, the host maintains a flow information table, which includes the flow information of N TCP flows, where N is an integer greater than or equal to 1. Among them, the flow information of each TCP flow includes the flow index of the corresponding TCP flow. The obtaining of the target flow information matching the target TCP flow includes: searching for the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow. For example, to implement the above functions of the embodiments of the present invention, the additional layer protocol needs to maintain a flow information table FlowInfoTable for recording the flow information when each TCP flow is divided into Flowlets. Each TCP flow occupies an entry in the FlowInfoTable table, as shown in Table 1,
[0160] Table 1
[0161]
[0162] In the above Table 1, each entry may include six items: SrcPort, LstFLTag, LstTS, TTDiff, TripTime_max, and TripTime_min. Among them,
[0163] (1) The TCP flow index (SrcPort) item is used to index each TCP flow. That is, the label of the TCP flow. Among them, this item corresponds to the flow index of the TCP flow described in this application.
[0164] (2) The field value of the previous packet (LstFLTag) item is the FL_Tag field value of the previous segment of this TCP flow. That is, whether the Flowlet identifier corresponding to the previous just-sent packet is 0 or 1. Among them, this item corresponds to the reference Flowlet identifier described in this application.
[0165] (3) The LstTS item is the timestamp value of the previous segment of this TCP flow. That is, the timestamp value of the previous just-sent packet (the timestamp added at the transport layer), and its unit is usually at the microsecond (us) level. Among them, this item corresponds to the timestamp value of the second segment described in this application.
[0166] (4) The TTDiff item is the time threshold for dividing Flowlets for this TCP flow. That is, each flow maintains a separate time threshold, and its unit is usually at the microsecond (us) level. For example, it is the difference between the TripTime_max item and the TripTime_min item shown in Table 1. For example, if the value of the TripTime_max item is 58 and the value of the TripTime_min item is 31, then the value of the TTDiff item is 27. Another example, if the value of the TripTime_max item is 49 and the value of the TripTime_min item is 36, then the value of the TTDiff item is 13. Among them, this item corresponds to the time threshold described in this application.
[0167] (5) The TripTime_max item is the delay of the uplink path with the maximum delay in the multi-path set corresponding to this TCP flow (the uplink path refers to the path from the sender to the receiver), and its unit is usually at the microsecond (us) level. Among them, this item corresponds to the first path delay described in this application.
[0168] (6) The TripTime_min item is the delay of the uplink path with the minimum delay in the multi-path set corresponding to this TCP flow, and its unit is usually at the microsecond (us) level. Among them, this item corresponds to the second path delay described in this application.
[0169] Step S503: Compare the difference between the timestamp of the first segment and the timestamp of the second segment with the time threshold.
[0170] Specifically, after determining the target TCP flow to which the first packet segment belongs and finding the target flow information corresponding to the target TCP flow (for example, one of the table entry contents in Table 1 above), the timestamp of the second packet segment can be determined therefrom, and the difference between the timestamp of the first packet segment (which is determined from the first packet segment) and the timestamp of the second packet segment is compared with the time threshold in the above target flow information; thereby comparing whether the time difference between the first packet segment and the second packet segment exceeds the maximum time interval corresponding to the target TCP flow to which the packet segment belongs, that is, the time threshold.
[0171] Step S504: According to the comparison result, determine whether to divide the first packet segment and the second packet segment into the same Flowlet.
[0172] Specifically, according to the comparison result in step S303, it is determined whether to divide the first packet segment and the second packet segment into the same Flowlet. For the same Flowlet in a TCP flow, it is forwarded through the same path, and for different Flowlets, the forwarding path needs to be re-determined.
[0173] In a possible implementation manner, if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the first packet segment and the second packet segment are divided into the same Flowlet; if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the first packet segment is divided into a new Flowlet. In the embodiments of the present invention, if the difference between the timestamps of the currently to-be-sent first packet segment and the previous adjacent second packet segment belonging to the target TCP flow is less than the time threshold corresponding to the target TCP flow (the time threshold is dynamically changed), it is considered that the first packet segment and the previous adjacent second packet segment meet the condition of being sent in the same Flowlet, that is, the first packet segment can be determined to be divided into the same Flowlet as the previous second packet segment; similarly, if the difference between the timestamps of the currently to-be-sent first packet segment and the previous adjacent second packet segment belonging to the target TCP flow is greater than the time threshold corresponding to the target TCP flow, it is considered that the first packet segment and the previous adjacent second packet segment do not meet the condition of being sent in the same Flowlet, that is, the first packet segment is divided into a new Flowlet.
[0174] As Figure 6A shown, Figure 6A is a schematic diagram of the first data packet and the second data packet in the same Flowlet provided by the embodiments of the present invention; in Figure 6AAmong them, the first data packet is the data packet after the first message segment is encapsulated by the additional layer, and the second data packet is the data packet after the second message segment is encapsulated by the additional layer. In Figure 6A Among them, assuming that the difference between the timestamp values of the first message segment and the second message segment is less than or equal to the time threshold, then the host side divides the first message segment and the second message segment into the same Flowlet (Flowlet5 in the figure), that is, the corresponding first data packet and the second data packet are divided into the same Flowlet5.
[0175] As Figure 6B shown, Figure 6B is a schematic diagram of the first data packet and the second data packet in different Flowlets provided by the embodiment of the present invention. In Figure 6B Among them, assuming that the difference between the timestamp values of the first message segment and the second message segment is greater than the time threshold, then the host side divides the first message segment and the second message segment into different Flowlets (Flowlet5 and Flowlet4 in the figure respectively), that is, the corresponding first data packet and the second data packet are divided into Flowlet5 and Flowlet4. It can be understood that at this time, it is equivalent to the first data packet being the first data packet in the new Flowlet.
[0176] In view of the problem that the fixed detection interval in the existing Flowlet granularity load balancing scheme is difficult to adapt to the dynamic network load, the embodiment of the present invention combines the delay feedback information of the network path and dynamically configures the time interval for detecting Flowlets to ensure that the Flowlet granularity matches the network path state.
[0177] Optionally, the embodiment of the present invention may further include the following method steps S505-S506.
[0178] Step S505: Generate a first data packet, where the first data packet includes the first message segment and the Flowlet identifier of the first message segment.
[0179] Specifically, the source host side further generates a first data packet after encapsulating the first data message segment through the additional layer and the network layer. The first data packet includes the first message segment and the Flowlet identifier of the first message segment. That is, in the process of generating the first data packet, in addition to encapsulating the headers of relevant protocols, the Flowlet identifier of the message segment also needs to be encapsulated. Optionally, the Flowlet identifier of the first message segment can be encapsulated in the Flowlet identifier bit on the header.
[0180] The target flow information further includes a reference Flowlet identifier of the target TCP flow (i.e., corresponding to the LstFLTag field in Table 1 above). Assume that the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment. That is, when the previous most recently sent packet segment is the second packet segment, this reference Flowlet identifier actually refers to the first Flowlet identifier corresponding to the second packet segment. If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier; if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
[0181] In an embodiment of the present invention, when further encapsulating the first packet segment to transmit data over the network, the Flowlet identifier corresponding to the packet segment can be set during the encapsulation process. After the packet segment is encapsulated into a data packet, the switch can identify which Flowlet the data packet belongs to through this Flowlet identifier, so as to determine which path to use for sending. For example, when the first packet segment enters the data link layer where the switch is located, the first packet segment needs to be further encapsulated. At this time, a flag bit for the switch to identify which Flowlet the packet segment belongs to is set in the encapsulated data packet. When the Flowlet identifiers of the first packet segment and the second packet segment are the same, the first data packet and the second data packet corresponding to the second packet segment are forwarded through the same path on the switch side. In summary, in the embodiment of the present invention, Flowlets are partitioned on the host side, and bits in the reserved field of the transport layer header (for example, 1 bit) can be used to mark Flowlets. The switch can identify Flowlets only relying on the header field, with high efficiency and low hardware overhead. At the same time, it also ensures that the same Flowlet will not be split again no matter how many hops of switches it goes through in the network, reducing the risk of out-of-order data packets.
[0182] Step S506: If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, update the reference Flowlet identifier to the second Flowlet identifier.
[0183] Specifically, the flow information of each TCP flow maintained on the host side further includes the reference Flowlet identifier of each TCP flow, that is, the identifier of the current Flowlet of each TCP flow is maintained in the flow information table, so as to set the corresponding Flowlet identifier for the packet segment to be sent. For example, assume that the reference Flowlet identifier is the first Flowlet identifier (that is, the Flowlet identifier corresponding to the second packet segment). Then, when the first packet segment and the second packet segment are divided into the same Flowlet, the Flowlet identifier of the first packet segment is also marked as the first Flowlet identifier, that is, the reference Flowlet identifier remains unchanged as the first Flowlet identifier; if the reference Flowlet identifier is the first Flowlet identifier, and when the first packet segment and the second packet segment are divided into different Flowlets (that is, the first packet segment is divided into a new Flowlet), the Flowlet identifier of the first packet segment is marked as the second Flowlet identifier, and at this time the reference Flowlet identifier needs to be updated to the second Flowlet identifier. Optionally, the reference Flowlet identifier can be switched between 0 or 1, that is, between two adjacent Flowlets, the Flowlet identifier takes values at intervals between 0 or 1. Therefore, only 1 bit can accurately indicate whether different data packets belong to the same Flowlet.
[0184] In a possible implementation manner, the host also updates the first path delay or the second path delay according to the received ACK packet to update the time threshold of the target TCP flow in real time. Specifically, the host receives the target ACK packet, where the target ACK packet is an ACK packet with the same destination port number or the same destination address as the target TCP flow; determines the uplink path delay of the target ACK packet, where the uplink path delay is the difference between the timestamp value and the timestamp echo reply value of the target ACK packet; compares the uplink path delay of the target ACK packet with the uplink path delays of the historical ACK packets in the target TCP flow; if the uplink path delay of the target ACK packet is greater than the maximum value of the uplink path delays of the historical ACK packets, updates the first path delay to the uplink path delay of the target ACK packet; if the uplink path delay of the target ACK packet is less than the minimum value of the uplink path delays of the historical ACK packets, updates the second path delay to the uplink path delay of the target ACK packet.
[0185] In the embodiment of the present invention, the time threshold for dividing Flowlets can be calculated from the difference between the maximum uplink path delay and the minimum uplink path delay of the historical packet segments received in the target TCP flow (or the TCP flows in the same network session as the target TCP flow). That is, the time threshold is a value that changes and adjusts dynamically in real time according to the network transmission load. Specifically, each time the host side receives an ACK packet belonging to the target TCP flow (i.e., the destination port numbers are the same), or receives an ACK packet of a TCP flow in the same network session as the target TCP flow (i.e., the destination addresses are the same or the destination network segments are the same), the transmission delay of the uplink path of the ACK packet is calculated by the difference between the timestamp value of the target ACK packet and the timestamp echo reply value. Based on the historical values of the uplink path transmission delays of all the received ACK packets, a current minimum uplink path delay is determined and used as the first path delay, and a current maximum uplink path delay is determined and used as the second path delay. Finally, the time threshold for dividing Flowlets in the target TCP flow is calculated using the difference between the first path delay and the second path delay. That is, after each receipt of the target ACK packet, it is necessary to detect whether the first path delay or the second path delay needs to be updated, so that the time thresholds for dividing Flowlets for different TCP flows or the data of the same TCP flow in different states are dynamically changing and are dynamically adjusted according to the real-time transmission delay of the data in the corresponding TCP flow. Therefore, it can always adapt to the dynamic network load changes.
[0186] Please refer to Figure 6C , Figure 6C FIG. is a schematic flow chart of dividing and marking Flowlets by an additional layer protocol provided by an embodiment of the present invention. Based on the flow information table maintained by the host in Table 1 above, the following exemplarily describes the implementation process of triggering the function of the additional layer and dividing and marking Flowlets when the TCP packet segment completes the encapsulation of the transport layer header fields, which may specifically include the following steps:
[0187] 1. After the TCP packet segment (such as the first packet segment) enters the transport layer, first index the corresponding entry of the TCP flow (denoted as [SrcPort]) in the FlowInfoTable information table (the flow information table described in Table 1) according to the source port information carried by the packet segment.
[0188] 2. Further obtain the timestamp value carried by the packet segment (such as the first packet segment) (denoted as CruTS).
[0189] 3. Next, it will be determined whether the packet segment (such as the first packet segment) meets the splitting conditions of Flowlet, that is, to determine the size relationship between the difference between the timestamp value of the packet segment (such as the timestamp value of the first packet segment) and the timestamp value of the previous packet segment recorded in the corresponding entry of the TCP flow and the splitting threshold of Flowlet (i.e., CurTS–[SrcPort].LstTS≥[SrcPort].TTDiff).
[0190] 4. If the matching decision condition is met, the current packet segment (such as the first packet segment) is regarded as the first packet segment of a new Flowlet and a mark is made on the FL_Tag bit in the header. The method of making the mark is to set the value of the FL_Tag bit in the header of the packet segment to the opposite value of the LstFLTag item in the entry, and then update the value of the LstFLTag item in the entry;
[0191] 5. If the matching decision condition is not met, the current packet segment (such as the first packet segment) is regarded as a subsequent packet segment of the previous Flowlet and a mark is made on the FL_Tag bit in the header. The method of making the mark is to set the value of the FL_Tag bit in the header of the packet segment to the same value as the LstFLTag item in the entry. When the additional layer completes the division and marking of Flowlet, the TCP packet segment is passed to the network layer.
[0192] Please refer to Figure 6D , Figure 6D which is a schematic flowchart of the process for the additional layer protocol to dynamically update the splitting threshold of Flowlet provided by the embodiment of the present invention. Based on the flow information table maintained by the host in Table 1 above, first of all, it should be noted that since the data center network topology provides multiple equivalent paths for the same pair of source and destination hosts, the embodiment of the present invention first continuously obtains the uplink path delay of the equivalent paths (including the equivalent paths within the same TCP flow, or the equivalent paths that can include within the same network session) according to the timestamp carried in the ACK packet, and records the maximum and minimum values of the uplink path delay. The difference between the maximum and minimum values of the uplink path delay is used to characterize the maximum delay difference between the equivalent paths, and then the TTDiff parameter is periodically configured based on the maximum delay difference. The following is an exemplary description of the implementation process of the additional layer protocol for dynamically configuring the TTDiff parameter used to indicate the division of Flowlet, which specifically may include the following steps:
[0193] 1. As Figure 6D shown, when the ACK packet (such as the target ACK packet) returned from the destination host side enters the additional layer protocol of the source host, the host side indexes the entry corresponding to the ACK packet (denoted as [SrcPort]) in the FlowInfoTable information table according to the carried destination port information. This entry is the same as the entry corresponding to the TCP flow associated with the ACK packet.
[0194] 2. Read the timestamp values carried in the ACK packet, including the timestamp value field value (denoted as Timesatmp) and the timestamp echo reply field value (denoted as TimesatmpEcho), and use the difference between these two timestamps to characterize the uplink path delay (denoted as TripTime).
[0195] 3. Compare the calculated uplink path delay with the maximum uplink path delay (i.e., [SrcPort].TrpTime_max) and the minimum uplink path delay (i.e., [SrcPort].TrpTime_min) recorded in the entry.
[0196] 4. If the uplink path delay is greater than the maximum uplink path delay recorded in the information table entry, update the [SrcPort].TrpTime_max item in the entry to the value of this uplink path delay;
[0197] 5. If the uplink path delay is less than the minimum uplink path delay recorded in the information table entry, update the [SrcPort].TripTime_min item in the entry to the value of this uplink path delay.
[0198] Optionally, the additional layer protocol can also periodically update the TTDiff item of all entries in the information table. In the embodiment of the present invention, the value of the TTDiff item is configured as the difference between the TripTime_max item and the TripTime_min item of each entry, and at the same time, the values of the TripTime_max item and the TripTime_min item are reset to avoid the continuous existence of invalid maximum or minimum values.
[0199] In the embodiment of the present invention, the update period can also be set to the time order of the network round-trip delay (about 100 - 200 microseconds), the reset value of the TripTime_max item is configured as zero, and the reset value of the TripTime_min item is configured as a relatively large value (such as the maximum value that can be represented by 4 bytes).
[0200] The Flowlet technology can well solve problems such as hash collision, mouse flow blocking, and asymmetry faced by data center network load balancing. Existing Flowlet-level solutions mostly detect and forward Flowlets at the switch based on a fixed time interval, but the fixed time interval cannot always match the dynamically changing traffic load of the data center network, which will lead to uneven distribution of in-network load. This application proposes to pre-divide the traffic at the terminal host based on a time threshold that adaptively changes with the path load, and then spread the fine-grained Flowlets into the network. After the switch identifies the Flowlets, it can execute any routing algorithm to further balance the load.
[0201] See Figure 7, Figure 7 is a schematic flowchart of a data transmission method provided by an embodiment of the present invention. This method can be applied to the switch in the network architecture described above Figure 2 or Figure 3 The switch 20 in the network architecture can be used to support and execute the method flow steps S701 - S704 shown in Figure 7 The following will describe from the switch side. This method may include the following steps S701 - S704 Figure 3 Step S701: Receive a first data packet
[0202] Specifically, the first data packet includes a first message segment and the Flowlet identifier of the first message segment
[0203] Step S702: Determine the target TCP flow to which the first data packet belongs, and obtain the forwarding information matching the target TCP flow
[0204] Specifically, the forwarding information includes the reference Flowlet identifier of the target TCP flow and the reference forwarding path; wherein, the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second message segment, and the second message segment is the previous message segment adjacent to the first message segment in the target TCP flow; the reference forwarding path is the first forwarding path of the second data packet, and the second data packet includes the second message segment and the first Flowlet identifier
[0205]
[0206] In a possible implementation, the switch maintains a forwarding information table, and the forwarding information table includes the forwarding information of M TCP flows, where M is an integer greater than or equal to 1. Among them, the forwarding information of each TCP flow includes the five-tuple hash value of the corresponding TCP flow; determining the target TCP flow to which the first data packet belongs and obtaining the forwarding information matching the target TCP flow includes: calculating the five-tuple hash value of the first data packet according to the five-tuple information of the first data packet; and searching for the forwarding information matching the target TCP flow from the forwarding information table according to the five-tuple hash value of the first data packet. In the embodiment of the present invention, the switch side maintains a forwarding information table, and the forwarding information table includes the forwarding information of one or more TCP flows (such as currently active TCP flows) on the hosts connected thereto, and the forwarding information of each TCP flow can include the five-tuple hash value of the TCP flow. That is, the switch can maintain the forwarding information of all currently active TCP flows, so that when there is a data packet to be sent, the forwarding information (including reference Flowlet identifier, forwarding path, etc.) matching the five-tuple hash value in the forwarding information table can be found according to the five-tuple hash value of the data packet, so as to forward the data packet to be sent.
[0207] Step S703: Compare the Flowlet identifier of the first segment with the first Flowlet identifier;
[0208] Step S704: Determine whether to forward the first segment through the first forwarding path according to the comparison result.
[0209] Specifically, after receiving a data packet, the switch side determines whether the Flowlet identifier in the data packet is the same as the Flowlet identifier of the adjacent data packet in the target TCP flow to which the first data packet belongs by identifying the Flowlet identifier in the data packet, and based on this, determines whether the first data packet needs to be forwarded through the forwarding path corresponding to the second data packet. That is, the switch side does not need to divide the data packets into Flowlets according to the time interval of the received data packets, but directly identifies whether the currently to-be-sent data packet belongs to the same Flowlet as the previous adjacent data packet in the same TCP flow according to the Flowlet identifier bit included in the received data packet, so as to decide whether to continue forwarding through the forwarding path of the adjacent data packet, or to divide a new Flowlet for the data packet and make a new forwarding path decision for it.
[0210] It should be noted that the forwarding path referred to in the embodiments of the present invention refers to the forwarding port that each switch can currently determine. That is to say, the complete forwarding path of a data packet may actually be jointly determined by the forwarding ports separately decided by multiple-hop switches. Therefore, in the embodiments of the present invention, on the switch side, it actually means that each switch in the multiple-hop switches executes the above data transmission method, and then finally determines the complete forwarding path of the first data packet.
[0211] In a possible implementation manner, if the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, then forward the first data packet through the first forwarding path; if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, then determine a second forwarding path for the first data packet and forward it through the second forwarding path. In the embodiments of the present invention, when a switch identifies that the Flowlet identifiers of the first data packet and the adjacent data packets in the target TCP flow to which it belongs are the same, then forward the first data packet and the second data packet on the same path; when a switch identifies that the Flowlet identifiers of the first data packet and the adjacent data packets in the target TCP flow to which it belongs are different, then determine a new forwarding path for the first data packet and forward it through the new forwarding path. It should be noted that the second forwarding path may be the same as or different from the first forwarding path, depending on the decision result of the switch.
[0212] In a possible implementation manner, the method further includes: if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is a second Flowlet identifier, then update the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and update the reference forwarding path to the second forwarding path. In the embodiments of the present invention, when the Flowlet identifiers of the first data packet and the second data packet are different, it means that the first data packet and the previous adjacent second data packet in the TCP flow to which it belongs do not belong to the same Flowlet. Therefore, the switch needs to divide the first data packet into a new Flowlet, and needs to update the reference Flowlet identifier of the TCP flow to which it belongs to the Flowlet identifier corresponding to the current latest data packet, that is, the second Flowlet identifier.
[0213] In summary, after dividing and marking Flowlets on the host side in the embodiments of the present invention, the switch identifies the Flowlets according to the identification bits of the data packets. The switch realizes the above functions through the Flowlet forwarding information table (Flowlet Table), where the format of Table 2 is as follows:
[0214] Table 2
[0215]
[0216] In the above Table 2, each entry of the Flowlet forwarding information table may include three items: Entry, FLTag, and Port. Among them,
[0217] (1) The Entry item records the hash value of the five-tuple of the data packet (source IP, destination IP, source port, destination port, protocol number), and indexes the corresponding entry of the TCP flow in the forwarding table based on this hash value. Among them, this item corresponds to the five-tuple hash value described in this application.
[0218] (2) The FLTag item records the Flowlet flag bit information for the identification of adjacent Flowlets. Among them, this item corresponds to the reference Flowlet identifier described in this application.
[0219] (3) The Port item records the information of the forwarding port. Among them, this item corresponds to the forwarding information described in this application.
[0220] Based on the forwarding information table maintained by the switch in the above Table 2, the following exemplary description is about the implementation process of identifying Flowlets according to the identification bits of data packets on the switch side. This implementation process may include the following main steps:
[0221] 1. For each arriving data packet, the switch first needs to identify which TCP flow the data packet belongs to, and then identify which Flowlet of the TCP flow the data packet belongs to.
[0222] 2. The switch first performs a hash operation on the five-tuple of the data packet to obtain a hash value, and then looks up the corresponding forwarding table entry of the Entry item equal to this hash value in the Flowlet forwarding table to determine which TCP flow the arriving data packet belongs to. For example, the hash values are key1, key2, etc. in the above Table 1.
[0223] 3. Then read the value of the FL_Tag bit in the data packet header and compare it with the current value of the FLTag item in this forwarding table entry, that is, compare the Flowlet identifier carried in the data packet with the reference Flowlet identifier recorded in the entry of the corresponding TCP flow.
[0224] 4. If they are equal, the data packet belongs to the current flow burst, and the data packet is forwarded to the output port indicated by the Port item of this entry.
[0225] 5. If they are not equal, it indicates that the data packet is a new flow burst of the data stream. Update the value of the FLTag item in the forwarding table entry to the value of the FL_Tag bit in the data packet header, and then perform a load balancing decision on the new Flowlet. Save the decision result to the Port item in the forwarding table entry.
[0226] 6. When determining which TCP flow the arriving data packet belongs to, if the corresponding forwarding table entry cannot be found based on the hash value, it indicates that the data packet belongs to a new TCP flow and is the first data packet of the new flow. At this time, a new forwarding table entry needs to be created. The value of the FLTag item in the forwarding table entry is the value of the Flowlet flag bit of the data packet. Then perform a load balancing decision and save the decision result to the Port item in the forwarding table entry.
[0227] For example, assume that there is a newly arrived data packet A at the switch. The switch performs a five-tuple hash on it and obtains a hash value of key1. Assume that the Flowlet identification value of data packet A is 0. The switch indexes the corresponding table entry in the forwarding information table based on the hash value key1, and compares the FLTag value in the table entry with the Flowlet identification value of data packet A. Since 0 = 0, data packet A is forwarded to the output port indicated by the Port item in the table entry (i.e., the first forwarding path described in this application).
[0228] Assume that there is a newly arrived data packet B at the switch. The switch performs a five-tuple hash on it and obtains a hash value of key2. Assume that the Flowlet identification value of data packet B is 0. The switch indexes the corresponding table entry in the forwarding information table based on the hash value key2, and compares the FLTag value in the table entry with the Flowlet identification value of data packet B. Since 1 ≠ 0, it indicates that data packet B is the first data packet of a new flowlet of this TCP flow. Therefore, data packet B cannot be forwarded to the output port indicated by the Port item in the table entry. It is necessary to re-perform a routing decision on data packet B, select a new output port, and update the value of the Port item (i.e., the second forwarding path described in this application).
[0229] In summary, the main protection points of this application can include the following:
[0230] 1. Add an additional layer protocol at the transport layer or a newly added additional layer, and maintain an information table. Each flow occupies an exclusive entry in the information table, and the entry includes relevant information when splitting the Flowlet.
[0231] 2. Continuously obtain the one-way delay information of the path at the host according to the timestamp option in the ACK packet, and then calculate the delay of the uplink path with the largest delay in the multi-path set (which can include equivalent paths or non-equivalent paths); and periodically set the maximum delay difference as the time threshold for splitting Flowlets to ensure that the time threshold can dynamically adapt to the path load and update the information table.
[0232] 3. When splitting a flow, once the difference between the timestamp of the current packet segment and the timestamp of the previous packet segment of the flow exceeds the set time threshold, the current packet segment is regarded as the first packet segment of a new Flowlet; use one bit of the reserved field in the TCP header as a flag bit to distinguish different Flowlets of the same flow, where the values of the flag bits of all packet segments within the same Flowlet are the same, and the values of the flag bits of adjacent Flowlets are opposite.
[0233] 4. At the switch, each Flowlet can be identified according to the five-tuple hash and one-bit flag of the transport layer or additional layer header, and any routing algorithm can be used to forward the data packet.
[0234] The method of the embodiment of the present invention is elaborated in detail above, and the related device of the embodiment of the present invention is provided below.
[0235] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data processing device provided by an embodiment of the present invention. The data processing device 80 may include a first generation unit 801, a first determination unit 802, an acquisition unit 803, a first comparison unit 804, and a Flowlet division unit 805. The detailed descriptions of each unit are as follows.
[0236] The first generation unit 801 is used to generate a first packet segment;
[0237] The first determination unit 802 is used to determine the target TCP flow to which the first packet segment belongs;
[0238] The acquisition unit 803 is used to acquire the timestamp of the first packet segment, and acquire target flow information matching the target TCP flow. The target flow information includes the time threshold corresponding to the target TCP flow, and the timestamp of the second packet segment in the target TCP flow; where the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow, and the time threshold is the difference between the first path delay and the second path delay. The first path delay is the delay of the uplink path with the largest delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the smallest delay in the multi-path set of the target TCP flow;
[0239] A first comparison unit 804, configured to compare the difference between the timestamp of the first packet segment and the timestamp of the second packet segment with the time threshold;
[0240] A Flowlet division unit 805, configured to determine whether to divide the first packet segment and the second packet segment into the same Flowlet according to the comparison result.
[0241] In a possible implementation manner, the first determination unit is specifically configured to:
[0242] Determine the target TCP flow to which the first packet segment belongs according to the source port number of the first packet segment.
[0243] In a possible implementation manner, the apparatus further includes:
[0244] A maintenance unit, configured to maintain a flow information table, where the flow information table includes flow information of N TCP flows, N is an integer greater than or equal to 1, and the flow information of each TCP flow includes a flow index of the corresponding TCP flow;
[0245] The obtaining unit is specifically configured to: search for the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow.
[0246] In a possible implementation manner, the Flowlet division unit is specifically configured to:
[0247] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, divide the first packet segment and the second packet segment into the same Flowlet;
[0248] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, divide the first packet segment into a new Flowlet.
[0249] In a possible implementation manner, the target flow information further includes a reference Flowlet identifier of the target TCP flow, and the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment; the apparatus further includes:
[0250] A second generation unit, configured to generate a first data packet, where the first data packet includes the first packet segment and the Flowlet identifier of the first packet segment; where
[0251] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier;
[0252] If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
[0253] In a possible implementation, the device further includes:
[0254] A first update unit, configured to update the reference Flowlet identifier to the second Flowlet identifier if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold.
[0255] In a possible implementation, the device further includes:
[0256] A receiving unit, configured to receive a target ACK packet, where the target ACK packet is an ACK packet with the same destination port number or the same destination address as that of the target TCP flow;
[0257] A second determination unit, configured to determine the uplink path delay of the target ACK packet, where the uplink path delay is the difference between the timestamp value of the target ACK packet and the timestamp echo reply value;
[0258] A second comparison unit, configured to compare the uplink path delay of the target ACK packet with the delay of the uplink path in the multi-path set of the target TCP flow;
[0259] A second update unit, configured to update the first path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is greater than the delay of the uplink path with the maximum current delay in the multi-path set;
[0260] A third update unit, configured to update the second path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is less than the delay of the uplink path with the minimum current delay in the multi-path set.
[0261] In a possible implementation, the multi-path set includes multiple equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple equivalent transmission paths and non-equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple non-equivalent transmission paths of the target TCP flow.
[0262] It should be noted that for the functions of each functional unit in the data processing device 80 described in the embodiments of the present invention, reference may be made to the relevant descriptions of steps S501 - S506 in the method embodiment described above, which will not be elaborated here. Figure 5 The relevant descriptions of steps S501 - S506 in the method embodiment described above will not be elaborated here.
[0263] Please refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of a data transmission device provided by an embodiment of the present invention. The data transmission device 90 may include a receiving unit 901, a determining unit 902, a comparing unit 903, and a forwarding unit 904. The detailed descriptions of each unit are as follows.
[0264] The receiving unit 901 is configured to receive a first data packet, where the first data packet includes a first message segment and a Flowlet identifier of the first message segment, and the first data packet belongs to a target TCP flow;
[0265] The determining unit 902 is configured to determine the target TCP flow to which the first data packet belongs, and obtain forwarding information matching the target TCP flow; the forwarding information includes a reference Flowlet identifier of the target TCP flow and a reference forwarding path; where the reference Flowlet identifier is currently the first Flowlet identifier corresponding to a second message segment, and the second message segment is the previous message segment adjacent to the first message segment in the target TCP flow; the reference forwarding path is the first forwarding path of a second data packet, and the second data packet includes the second message segment and the first Flowlet identifier;
[0266] The comparing unit 903 is configured to compare the Flowlet identifier of the first message segment with the first Flowlet identifier;
[0267] The forwarding unit 904 is configured to determine whether to forward the first message segment through the first forwarding path according to the comparison result.
[0268] In a possible implementation manner, the switch maintains a forwarding information table, and the forwarding information table includes the forwarding information of M TCP flows, where M is an integer greater than or equal to 1. The forwarding information of each TCP flow includes the five - tuple hash value of the corresponding TCP flow. The determining unit is specifically configured to:
[0269] Calculate the five - tuple hash value of the first data packet according to the five - tuple information of the first data packet;
[0270] Search for the forwarding information matching the target TCP flow from the forwarding information table according to the five - tuple hash value of the first data packet.
[0271] In a possible implementation, the forwarding unit is specifically configured to:
[0272] If the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, forward the first data packet through the first forwarding path;
[0273] If the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, determine a second forwarding path for the first data packet and forward it through the second forwarding path.
[0274] In a possible implementation, the apparatus further includes:
[0275] An update unit, if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is the second Flowlet identifier, update the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and update the reference forwarding path to the second forwarding path.
[0276] It should be noted that for the functions of the functional units in the data transmission apparatus 90 described in the embodiments of the present invention, reference may be made to the relevant descriptions of steps S701 - S704 in the method embodiments described above, which will not be elaborated here. Figure 7 in the method embodiments described above, which will not be elaborated here.
[0277] The embodiments of the present invention further provide a host, where the host includes a processor, a memory, and a communication interface. The memory is used to store data processing program codes, and the processor is used to call the data processing program codes to execute some or all of the steps of any one of the data processing methods described in the above method embodiments.
[0278] The embodiments of the present invention further provide a switch, where the host includes a processor, a memory, and a communication interface. The memory is used to store data transmission program codes, and the processor is used to call the data transmission program codes to execute some or all of the steps of any one of the data transmission methods described in the above method embodiments.
[0279] The embodiments of the present invention further provide a computer-readable storage medium, where the computer-readable storage medium may store a program, and when the program is executed by a host, it includes some or all of the steps of any one of the data processing methods described in the above method embodiments.
[0280] The embodiments of the present invention further provide a computer program, the computer program includes instructions, and when the computer program is executed by a switch, it enables the switch to execute some or all of the steps of any one of the data transmission methods.
[0281] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0282] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps may be implemented in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0283] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0284] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0285] In addition, in each embodiment of this application, the various functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0286] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute all or part of the steps of the above methods in various embodiments of this application. Among them, the aforementioned storage medium can include: various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, optical discs, read-only memory (ROM), or random access memory (RAM).
[0287] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application.
Claims
1. A data processing method, characterized in that, Applied to a host, including: Generate a first packet segment and determine the target TCP flow to which the first packet segment belongs; Obtain the timestamp of the first packet segment and obtain target flow information matching the target TCP flow, where the target flow information includes a time threshold corresponding to the target TCP flow and the timestamp of a second packet segment in the target TCP flow; wherein, the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow, the time threshold is the difference between a first path delay and a second path delay, the first path delay is the delay of the uplink path with the largest delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the smallest delay in the multi-path set of the target TCP flow; Compare the difference between the timestamp of the first packet segment and the timestamp of the second packet segment with the time threshold; Determine whether to divide the first packet segment and the second packet segment into the same Flowlet according to the comparison result.
2. The method according to claim 1, wherein The determining the target TCP flow to which the first packet segment belongs includes: Determine the target TCP flow to which the first packet segment belongs according to the source port number of the first packet segment.
3. The method according to claim 1, wherein The host maintains a flow information table, and the flow information table includes the flow information of N TCP flows, where N is an integer greater than or equal to 1. Each TCP flow's flow information includes the flow index of the corresponding TCP flow; the obtaining the target flow information matching the target TCP flow includes: Search for the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow.
4. The method according to any one of claims 1 to 3, characterized in that The determining whether to divide the first packet segment and the second packet segment into the same Flowlet according to the comparison result includes: If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, divide the first packet segment and the second packet segment into the same Flowlet; If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, divide the first packet segment into a new Flowlet.
5. The method according to any one of claims 1 to 3, characterized in that, The target flow information further includes a reference Flowlet identifier of the target TCP flow, and the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment; the method further includes: Generate a first data packet, where the first data packet includes the first packet segment and the Flowlet identifier of the first packet segment; wherein, If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier; If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
6. The method according to claim 5, characterized in that, The method further includes: If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, update the reference Flowlet identifier to the second Flowlet identifier.
7. The method according to any one of claims 1 to 3, characterized in that The method further includes: Receiving a target ACK packet, where the target ACK packet is an ACK packet with the same destination port number or the same destination address as that of the target TCP flow; Determining the uplink path delay of the target ACK packet, where the uplink path delay is the difference between the timestamp value of the target ACK packet and the timestamp return reply value; Comparing the uplink path delay of the target ACK packet with the uplink path delays in the multipath set of the target TCP flow; If the uplink path delay of the target ACK packet is greater than the uplink path delay of the currently largest-delay uplink path in the multipath set, update the first path delay to the uplink path delay of the target ACK packet; If the uplink path delay of the target ACK packet is less than the uplink path delay of the currently smallest-delay uplink path in the multipath set, update the second path delay to the uplink path delay of the target ACK packet.
8. The method according to any one of claims 1 to 3, characterized in that The multipath set includes multiple equivalent transmission paths of the target TCP flow; or, the multipath set includes multiple equivalent transmission paths and non-equivalent transmission paths of the target TCP flow; or, the multipath set includes multiple non-equivalent transmission paths of the target TCP flow.
9. A data transmission method, characterized in that, Applied to a switch, it includes: Receiving a first data packet, where the first data packet includes a first packet segment and the Flowlet identifier of the first packet segment; Determining the target TCP flow to which the first data packet belongs, and obtaining the forwarding information matching the target TCP flow; the forwarding information includes the reference Flowlet identifier and the reference forwarding path of the target TCP flow; where the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment, and the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow; the reference forwarding path is the first forwarding path of the second data packet, and the second data packet includes the second packet segment and the first Flowlet identifier; Comparing the Flowlet identifier of the first packet segment with the first Flowlet identifier; Determining whether to forward the first packet segment through the first forwarding path according to the comparison result.
10. The method according to claim 9, wherein The switch maintains a forwarding information table, and the forwarding information table includes the forwarding information of M TCP flows, where M is an integer greater than or equal to 1, and the forwarding information of each TCP flow includes the five-tuple hash value of the corresponding TCP flow; the determining the target TCP flow to which the first data packet belongs and obtaining the forwarding information matching the target TCP flow includes: Calculating the five-tuple hash value of the first data packet according to the five-tuple information of the first data packet; Searching for the forwarding information matching the target TCP flow from the forwarding information table according to the five-tuple hash value of the first data packet.
11. The method according to claim 9, wherein Determining whether to forward the first packet segment through the forwarding path according to the comparison result includes: If the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, forwarding the first data packet through the first forwarding path; If the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, determining a second forwarding path for the first data packet and forwarding it through the second forwarding path.
12. The method according to claim 11, wherein The method further includes: If the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is the second Flowlet identifier, updating the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and updating the reference forwarding path to the second forwarding path.
13. A data processing device, characterized in that, Including: A first generation unit for generating a first packet segment; A first determination unit for determining the target TCP flow to which the first packet segment belongs; An acquisition unit for acquiring the timestamp of the first packet segment and acquiring target flow information matching the target TCP flow, where the target flow information includes a time threshold corresponding to the target TCP flow, and the timestamp of a second packet segment in the target TCP flow; wherein, the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow, the time threshold is the difference between a first path delay and a second path delay, the first path delay is the delay of the uplink path with the largest delay in the multi-path set of the target TCP flow, and the second path delay is the delay of the uplink path with the smallest delay in the multi-path set of the target TCP flow; A first comparison unit for comparing the difference between the timestamp of the first packet segment and the timestamp of the second packet segment with the time threshold; A Flowlet division unit for determining whether to divide the first packet segment and the second packet segment into the same Flowlet according to the comparison result.
14. The device according to claim 13, characterized in that, The first determination unit is specifically used for: Determining the target TCP flow to which the first packet segment belongs according to the source port number of the first packet segment.
15. The device according to claim 13, characterized in that, The apparatus further includes: A maintenance unit for maintaining a flow information table, where the flow information table includes the flow information of N TCP flows, N is an integer greater than or equal to 1, and the flow information of each TCP flow includes the flow index corresponding to the TCP flow; The acquisition unit is specifically used for: looking up the target flow information matching the target TCP flow from the flow information table according to the flow index of the target TCP flow.
16. The device according to any one of claims 13-15, characterized in that, The Flowlet division unit is specifically used for: If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, dividing the first packet segment and the second packet segment into the same Flowlet; If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, dividing the first packet segment into a new Flowlet.
17. The device according to any one of claims 13 to 15, characterized in that The target flow information further includes a reference Flowlet identifier of the target TCP flow, and the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment; The apparatus further includes: A second generating unit, configured to generate a first data packet, where the first data packet includes the first packet segment and the Flowlet identifier of the first packet segment; Wherein, If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is less than or equal to the time threshold, the Flowlet identifier of the first packet segment is the first Flowlet identifier; If the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold, the Flowlet identifier of the first packet segment is the second Flowlet identifier.
18. The device according to claim 17, characterized in that, The apparatus further includes: A first updating unit, configured to update the reference Flowlet identifier to the second Flowlet identifier if the difference between the timestamp of the first packet segment and the timestamp of the second packet segment is greater than the time threshold.
19. The device according to any one of claims 13 to 15, characterized in that, The apparatus further includes: A receiving unit, configured to receive a target ACK packet, where the target ACK packet is an ACK packet having the same destination port number or the same destination address as the target TCP flow; A second determining unit, configured to determine an uplink path delay of the target ACK packet, where the uplink path delay is the difference between the timestamp value and the timestamp return answer value of the target ACK packet; A second comparing unit, configured to compare the uplink path delay of the target ACK packet with the delay of the uplink path in the multi-path set of the target TCP flow; A second updating unit, configured to update the first path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is greater than the delay of the uplink path with the maximum current delay in the multi-path set; A third updating unit, configured to update the second path delay to the uplink path delay of the target ACK packet if the uplink path delay of the target ACK packet is less than the delay of the uplink path with the minimum current delay in the multi-path set.
20. The device according to any one of claims 13-15, characterized in that, The multi-path set includes multiple equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple equivalent transmission paths and non-equivalent transmission paths of the target TCP flow; or, the multi-path set includes multiple non-equivalent transmission paths of the target TCP flow.
21. A data transmission device, characterized in that, Includes: A receiving unit, configured to receive a first data packet, where the first data packet includes a first packet segment and the Flowlet identifier of the first packet segment, and the first data packet belongs to a target TCP flow; A determining unit, configured to determine the target TCP flow to which the first data packet belongs, and obtain forwarding information matching the target TCP flow; The forwarded information includes the reference Flowlet identifier of the target TCP flow and the reference forwarding path; wherein, the reference Flowlet identifier is currently the first Flowlet identifier corresponding to the second packet segment, and the second packet segment is the previous packet segment adjacent to the first packet segment in the target TCP flow; the reference forwarding path is the first forwarding path of the second data packet, and the second data packet includes the second packet segment and the first Flowlet identifier; A comparison unit, configured to compare the Flowlet identifier of the first packet segment with the first Flowlet identifier; A forwarding unit, configured to determine whether to forward the first packet segment through the first forwarding path according to the comparison result.
22. The device according to claim 21, wherein, The apparatus maintains a forwarding information table, and the forwarding information table includes the forwarding information of M TCP flows, where M is an integer greater than or equal to 1. Each piece of forwarding information of a TCP flow includes the five-tuple hash value of the corresponding TCP flow; the determining unit is specifically configured to: Calculate the five-tuple hash value of the first data packet according to the five-tuple information of the first data packet; Search for the forwarding information matching the target TCP flow from the forwarding information table according to the five-tuple hash value of the first data packet.
23. The device according to claim 21 or 22, characterized in that, The forwarding unit is specifically configured to: If the Flowlet identifier of the first packet segment is the same as the first Flowlet identifier, forward the first data packet through the first forwarding path; If the Flowlet identifier of the first packet segment is different from the first Flowlet identifier, determine a second forwarding path for the first data packet and forward it through the second forwarding path.
24. The device according to claim 23, characterized in that, The apparatus further includes: An updating unit, if the Flowlet identifier of the first packet segment is different from the first Flowlet identifier and is the second Flowlet identifier, update the reference Flowlet identifier of the target TCP flow to the second Flowlet identifier, and update the reference forwarding path to the second forwarding path.
25. A host, characterized in that, It includes a processor, a memory, and a communication interface. The memory is used to store data processing program codes, and the processor is used to call the data processing program codes to execute the method according to any one of claims 1-8.
26. A switch, characterized in that, It includes a processor, a memory, and a communication interface. The memory is used to store data transmission program codes, and the processor is used to call the data transmission program codes to execute the method according to any one of claims 9-12.
27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a host, the method according to any one of claims 1-8 is implemented.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a switch, the method according to any one of claims 9-12 is implemented.
29. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions are executed by a host, the host is caused to execute the method according to any one of claims 1-8.
30. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a switch, cause the switch to perform the method according to any one of claims 9-12.
Citation Information
Patent Citations
Load balance processing method and device
CN108270687A
Data center load balancing method for asymmetric network
CN110061929A