Scalable E2E network architecture and components supporting low latency and high throughput
By establishing virtual tunnels in the network system and using credit mechanisms to manage multiple data flow paths, the network congestion problem is solved, low-latency and high throughput data transmission is achieved, and the system's scalability and data transmission efficiency are improved.
Patent Information
- Application Number
- CN202210819930.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-30
- Filing Date
- 2022-07-11
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-07-11
AI Technical Summary
Existing network systems are prone to congestion in the flow path, resulting in increased delays and reduced throughput, and traditional flow congestion management techniques fail to scale effectively.
By establishing a virtual tunnel between the source endpoint and the destination endpoint, data transmission is transmitted using multiple data flow paths, and data transmission sequences are updated based on the received credit, using a credit mechanism for congestion prevention and management, including a combination of optimistic flow and scheduled flow to optimize routing and scheduling of data packets.
It realizes low latency and high throughput data transmission, reduces network congestion, improves the scalability of the system and the transmission efficiency of data packets.
Smart Images

Figure CN115695295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to scalable e2e network architecture and components that support low latency and high throughput. Background Art
[0002] Certain network-based systems (e.g., data centers, Internet-based systems, etc.) can communicate by transporting data packets to and from various sources and endpoints. Several layers (e.g., transport layer, network layer, etc.) can be managed to control and improve the transmission and reception of data packets, such as segmentation, error control, logical addressing, routing, path determination, and flow congestion management.
[0003] In some systems, flow congestion can occur in the flow path between networks due to different transmission and processing rates. For example, a smartphone is downloading a file from an off-site server. The server can transmit data at a maximum speed of 100 Mbps, and the smartphone can process data at a maximum speed of 10 Mbps. The server sends data at 50 Mbps, causing congestion in the flow path. To address this, the smartphone can instruct the server (e.g., via the transport layer) to reduce the transmission rate to 10 Mbps.
[0004] Some flow congestion management systems (e.g., those following the architecture outlined above) implement reactive management techniques that create a feedback loop to remove congestion, resulting in feedback delays and potential overshoot. Furthermore, these techniques are not configured to be scalable, but are limited to a single flow path being managed. Summary of the Invention
[0005] On the one hand, the present application relates to a method for managing network services, the method comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; receiving a plurality of credits from the destination endpoint at the source endpoint, the plurality of credits being provided via two or more of the plurality of data flow paths; updating a data transmission sequence at the source endpoint based on the plurality of credits; and providing a plurality of data packets to the destination endpoint based on the data transmission sequence.
[0006] On the other hand, the present application relates to one or more non-transitory computer-readable media having computer-executable instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising multiple data flow paths, each of the multiple data flow paths connecting the source endpoint and the destination endpoint; receiving multiple credits from the destination endpoint at the source endpoint, the multiple credits being provided via two or more of the multiple data flow paths; updating a data transmission sequence at the source endpoint based on the multiple credits; and providing multiple data packets to the destination endpoint based on the data transmission sequence.
[0007] On the other hand, the present application relates to a device for managing network services, wherein a controller includes one or more processors and a memory, wherein the memory stores instructions, which, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel including multiple data flow paths, each of the multiple data flow paths connecting the source endpoint and the destination endpoint; providing multiple credits to the source endpoint via the destination endpoint, the multiple credits being provided via two or more of the multiple data flow paths; updating a data transmission sequence at the source endpoint based on the multiple credits; providing multiple data packets to the destination endpoint based on the data transmission sequence, wherein establishing the virtual tunnel includes establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a high-level diagram of an inter-network system according to some embodiments.
[0009] Figure 2 According to some embodiments, Figure 1 Figure 1 shows a diagram of a virtual tunnel implemented in an inter-network system.
[0010] Figure 3 According to some embodiments, Figure 1 Figure 3 is a diagram of a virtual tunnel with multiple flow paths implemented in an inter-network system.
[0011] Figure 4 According to some embodiments, Figure 1 Figure 2 shows a diagram of a virtual tunnel with two flows implemented in an inter-network system.
[0012] Figure 5 According to some embodiments, Figure 1Figure 2 shows a diagram of a virtual tunnel with two flows implemented in an inter-network system.
[0013] Figure 6 According to some embodiments, Figure 1 Figure 2 shows a diagram of a virtual tunnel with two flows implemented in an inter-network system.
[0014] Figure 7 According to some embodiments, Figure 1 Diagram of a virtual tunnel implemented in an inter-network system to provide credit to a source.
[0015] Figure 8 According to some embodiments, Figure 1 Block diagram of a virtual tunnel implemented in an inter-network system to provide packets to a destination.
[0016] Figure 9 According to some embodiments, Figure 7 The schedule of credit generation implemented in the virtual tunnel.
[0017] Figure 10 According to some embodiments, Figure 1 Block diagram of a system for congestion management implemented in an inter-network system.
[0018] Figure 11 According to some embodiments, Figure 1 Block diagram of a system for congestion management implemented in an inter-network system.
[0019] Figure 12 According to some embodiments, Figure 1 Block diagram of a system for congestion management implemented in an inter-network system.
[0020] Figure 13 According to some embodiments, Figure 1 Block diagram of a system for congestion management implemented in an inter-network system.
[0021] Figure 14 According to some embodiments, Figure 1 Block diagram of flow-based queuing and scheduling implemented in an inter-network system.
[0022] Figure 15 According to some embodiments, Figure 1 Figure 3 is a diagram of a virtual tunnel implemented in an inter-network system for managing unordered packets. DETAILED DESCRIPTION
[0023] Overview
[0024] Referring generally to the accompanying drawings, systems and methods for implementing virtual tunnels within an inter-network system are shown according to some embodiments. The virtual tunnel can surround multiple flow paths and provide a mechanism for managing the routing of data packets between the multiple flow paths, rather than within a single flow path between a source endpoint and a destination endpoint. In addition, the systems and methods disclosed herein may include a routing management system that employs credit-based congestion prevention, allowing virtual tunnels to significantly reduce congestion for the data packets. The credits can be provided by the destination to the source, allowing the source to adjust the scheduling / routing of the data packets to increase throughput and reduce latency. Finally, the systems and methods disclosed herein may also include several flows (e.g., optimistic flows, scheduled flows, etc.) within the virtual tunnel, which provide another layer of mobility for routing the data packets from the source to the destination. In some embodiments, this results in a low-latency and scalable flow management system for maintaining high throughput data packet transmission.
[0025] The embodiments summarized below are illustrative only and are not intended to be limiting in any way. Other aspects, inventive features, and advantages of the devices or processes described herein will become apparent from the detailed description set forth herein, taken in conjunction with the accompanying drawings, in which like reference numerals refer to like elements.
[0026] One embodiment of the present disclosure is a method for managing network traffic. The method includes establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel including multiple data flow paths, each of the multiple data flow paths connecting the source endpoint and the destination endpoint. The method further includes receiving multiple credits from the destination endpoint at the source endpoint, the multiple credits provided via two or more of the multiple data flow paths. The method further includes updating a data transmission sequence at the source endpoint based on the multiple credits. The method further includes providing multiple data packets to the destination endpoint based on the data transmission sequence.
[0027] In some embodiments, providing the plurality of credits to the source endpoint includes generating the plurality of credits at a generation rate, the generation rate being based on a port rate of a destination network interface controller (NIC), the destination NIC being configured to receive the plurality of data packets from the destination endpoint and provide the plurality of data packets to a destination server, determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path, and providing at least one data packet of the plurality of data packets to the destination endpoint via the non-congested data flow path.
[0028] In some embodiments, the method further includes establishing a first data flow and a second data flow within the virtual tunnel, using the first data flow to provide a first subset of the plurality of data packets to the destination endpoint, and using the second data flow to provide a second subset of the plurality of data packets to the destination endpoint, wherein the second data flow provides the second subset based on at least one of the plurality of credits.
[0029] In some embodiments, the method further includes generating a first virtual order queue (VOQ) for the first data flow and a second VOQ for the second data flow at the source endpoint, and scheduling two or more of the multiple data packets into the first VOQ or the second VOQ based on a flow path criterion, the flow path criterion indicating at least one of a priority requirement or a latency requirement of the two or more data packets, wherein the data transmission sequence includes the first VOQ and the second VOQ.
[0030] In some embodiments, the method further includes receiving the first subset and the second subset at the destination endpoint, and reordering the plurality of data packets at the destination endpoint by combining the first subset and the second subset in the same order that provided the first subset and the second subset.
[0031] In some embodiments, the method further includes receiving the first subset and the second subset at the destination endpoint, and reordering the plurality of data packets at the destination endpoint by combining the first subset and the second subset based on cross-flow reordering, wherein the cross-flow reordering is configured to notify the destination endpoint that the first subset is received before the second subset. In some embodiments, the first data flow is an optimistic flow and the second data flow is a scheduled flow.
[0032] In some embodiments, the data transmission sequence includes a transmission order of each of the plurality of data packets configured to be transmitted to the destination endpoint, and a selected data flow of the plurality of data flow paths for each of the plurality of data packets.
[0033] In some embodiments, establishing the virtual tunnel includes establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
[0034] In some embodiments, receiving the plurality of credits at the source endpoint includes receiving the plurality of credits at a first credit rate and, responsive to a determination of one or more congested data flow paths among the plurality of data flow paths, receiving the plurality of credits at a second credit rate.
[0035] In some embodiments, receiving the plurality of credits at the source endpoint includes receiving shaped credits from the plurality of credits, the shaped credits being shaped using one or more transit switches within the virtual tunnel, and receiving an indication of discarded credits from the plurality of credits, the discarded credits being discarded using the one or more transit switches within the virtual tunnel.
[0036] In some embodiments, receiving the indication of the discarded credit comprises receiving a negative acknowledgement (NACK) at the source endpoint, the NACK indicating that a header for the discarded credit was received at the destination endpoint.
[0037] Another embodiment of the present disclosure is one or more non-transitory computer-readable media having computer-executable instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations. The operations include establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel including a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint. The operations further include receiving a plurality of credits from the destination endpoint at the source endpoint, the plurality of credits provided via two or more of the plurality of data flow paths. The operations further include updating a data transmission sequence at the source endpoint based on the plurality of credits. The operations further include providing a plurality of data packets to the destination endpoint based on the data transmission sequence.
[0038] In some embodiments, providing the plurality of credits to the source endpoint includes generating the plurality of credits at a generation rate, the generation rate being based on a port rate of a destination network interface controller (NIC), the destination NIC being configured to receive the plurality of data packets from the destination endpoint and provide the plurality of data packets to a destination server, determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path, and providing at least one data packet of the plurality of data packets to the destination endpoint via the non-congested data flow path.
[0039] In some embodiments, the one or more processors are further configured to establish a first data flow and a second data flow within the virtual tunnel, use the first data flow to provide a first subset of the plurality of data packets to the destination endpoint, and use the second data flow to provide a second subset of the plurality of data packets to the destination endpoint, wherein the second data flow provides the second subset based on at least one of the plurality of credits.
[0040] In some embodiments, the one or more processors are further configured to generate a first virtual order queue (VOQ) for the first data flow and a second VOQ for the second data flow at the source endpoint, and schedule two or more of the plurality of data packets into the first VOQ or the second VOQ based on a flow path criterion, the flow path criterion indicating at least one of a priority requirement or a latency requirement of the two or more data packets. In some embodiments, the data transmission sequence includes the first VOQ and the second VOQ.
[0041] In some embodiments, the one or more processors are further configured to receive the first subset and the second subset at the destination endpoint, and to reorder the plurality of data packets at the destination endpoint by combining the first subset and the second subset in the same order that provided the first subset and the second subset.
[0042] In some embodiments, the one or more processors are further configured to receive the first subset and the second subset at the destination endpoint, and reorder the plurality of data packets at the destination endpoint by combining the first subset and the second subset based on cross-flow reordering, wherein the cross-flow reordering is configured to notify the destination endpoint that the first subset is received before the second subset. In some embodiments, the first data flow is an optimistic flow and the second data flow is a scheduled flow.
[0043] In some embodiments, the data transmission sequence includes a transmission order of each of the plurality of data packets configured to be transmitted to the destination endpoint, and a selected data flow of the plurality of data flow paths for each of the plurality of data packets.
[0044] In some embodiments, establishing the virtual tunnel includes establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
[0045] Another embodiment of the present disclosure is a device for managing network traffic, the device comprising one or more processors and a memory, the memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations include establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising multiple data flow paths, each of the multiple data flow paths connecting the source endpoint and the destination endpoint. The operations include providing multiple credits to the source endpoint via the destination endpoint, the multiple credits being provided via two or more of the multiple data flow paths. The operations include updating a data transmission sequence at the source endpoint based on the multiple credits. The operations include providing multiple data packets to the destination endpoint based on the data transmission sequence. In some embodiments, establishing the virtual tunnel includes establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
[0046] In some embodiments, providing the plurality of credits to the source endpoint includes generating the plurality of credits at a generation rate, the generation rate being based on a port rate of a destination network interface controller (NIC), the destination NIC being configured to receive the plurality of data packets from the destination endpoint and provide the plurality of data packets to a destination server, determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, and allowing at least one data packet of the plurality of data packets to be provided to the destination endpoint via the data flow path.
[0047] In some embodiments, the one or more processors are further configured to establish a first data flow and a second data flow within the virtual tunnel, use the first data flow to provide a first subset of the plurality of data packets to the destination endpoint, and use the second data flow to provide a second subset of the plurality of data packets to the destination endpoint, wherein the second data flow provides the second subset based on at least one of the plurality of credits. In some embodiments, the one or more processors are further configured to generate a first virtual order queue (VOQ) for the first data flow and a second VOQ for the second data flow at the source endpoint, and schedule two or more of the plurality of data packets into the first VOQ or the second VOQ based on a flow path criterion indicating at least one of a priority requirement or a latency requirement of the two or more data packets, wherein the data transmission sequence includes the first VOQ and the second VOQ.
[0048] In some embodiments, the apparatus is provided in an integrated circuit package.
[0049] Inter-network System Overview
[0050] Now refer to Figure 1 , a diagram of an inter-network system 100 is shown, according to some embodiments. System 100 may refer to any type of system that provides for transmission and / or reception of data via data packet communications. System 100 is shown to include a data center 102, a network 104, a network 106, a device 108, and a communication path 110. In some embodiments, system 100 advantageously implements a low-latency and scalable flow management system for data packet transmission while maintaining high throughput (e.g., successful message delivery over a communication channel, etc.).
[0051] Data center 102 can be any location (e.g., a building, a dedicated space within a building, a group of buildings, etc.) for housing computer systems and associated components (e.g., servers, telecommunications, storage systems, etc.). The disclosed systems and methods can be or include processing performed partially or entirely within data center 102. While data center 102 is often referred to as at least one of a source or a destination herein, this is exemplary only and should not be considered limiting. It is contemplated that several data centers (e.g., one data center acting as a source and one data center acting as a destination) or other types of network devices can act as sources and / or destinations (e.g., another device, a terminal, a smartphone, etc.) can be used.
[0052] Network 104 (and similarly, network 106) can be any processing device and computer group capable of sharing resources (e.g., local, provided by a network node, etc.). A set of communication protocols (e.g., TCP / IP, etc.) can be used between the devices to communicate with each other and provide / receive data. The nodes of network 104 can include personal computers, servers, network hardware, or other dedicated or general-purpose hosts. In some embodiments, the nodes are identified by host names and / or network addresses.
[0053] Communication path 110 can be configured to facilitate communication between network 104 and network 106. In some embodiments, communication path 110 includes one or more virtual tunnels. In some embodiments, a virtual tunnel acts as a point-to-point connection between two points (e.g., a source point and a destination point, etc.). As discussed above, a virtual tunnel can be configured to contain several different routing paths for data packets and allow data transmission within the tunnel to be managed as a separate and scalable management technique. The advantages and techniques for implementing virtual tunnels in system 100 are described in more detail below.
[0054] Now refer to Figure 2, a diagram illustrating a system 200 for virtual tunneling according to some embodiments. In some embodiments, system 200 illustrates data being provided from a source-side server (e.g., servers 202, 204) to a destination-side server (e.g., server 216). As mentioned above, the terminal devices (e.g., server 202, server 216, etc.) described herein are meant to be exemplary and should not be considered limiting. Any type of processing device and / or computing device is contemplated. In some embodiments, system 200 includes systems and methods for establishing a virtual tunnel with a top-of-rack (ToR) switch.
[0055] In some embodiments, ToR switching is a network architecture design in which computing equipment (e.g., servers, facilities, other switches, etc.) located within the same or adjacent "racks" is connected to an intra-rack network switch. The intra-rack network switch can, in turn, be connected to an aggregation switch (e.g., via fiber optic cables, etc.). Although the systems and methods disclosed herein generally consider virtual tunnels with ToR switch endpoints, other computer architectures, such as end-of-row (EoR) designs, may be considered. System 200 is shown to include servers 202, 204, NICs 206, 208, a source top-of-rack (S-ToR) switch 210, a tunnel source 211, a high-performance network transport (HPNT) tunnel 110 ("tunnel 110"), a destination top-of-rack (D-ToR) switch 212, a tunnel destination (TDST) 213, a NIC 214, and a server 216.
[0056] In one example, server 202 provides several data packets to network interface controller (NIC) 206. NIC 206 may be provided in a single integrated circuit package. NIC 206 may provide the data packets to S-ToR switch 210. S-ToR switch 210 may include TSRC 211. In some embodiments, TSRC 211 acts as a source endpoint and includes a source address from which data packets are provided and to which credits are provided. TSR1 212 may queue some or all packets originating from some or all attached NIC ports (e.g., ports of NICs 206, 208, etc.) into a tunnel virtual output queue (VOQ) (not shown). In some embodiments, VOQ stands for virtual output queue, virtual in-order queue, or both. For the purposes of this disclosure, VOQ will represent the virtual in-order queue of the disclosed embodiments.
[0057] In some embodiments, virtual output queues may be a technique used in certain network switch architectures (e.g., system 200, etc.) in which, rather than keeping all traffic in a single queue, a separate queue is maintained for each possible output location. This can address common problems, such as "head of line" blocking. In virtual output queues, the physical buffer of each input port can maintain a separate virtual queue for each output. Thus, congestion on an egress port may only block the virtual queue for that particular egress port, while other packets in the same physical buffer destined for different (non-congested) output ports may be in separate virtual queues and therefore still be processed. Using alternative techniques, a blocked packet at a congested egress port may have blocked the entire physical buffer, resulting in head of line blocking.
[0058] Tunnel 110 may act as a virtual tunnel that allows for “latency-sensitive” (LS) flows between source-destination pairs (e.g., S-ToR 210, D-ToR 213, etc.), which may be coupled to and / or include ports connecting the source / destination to respective NICs (e.g., SRC 211, TDST 213, etc.). In some embodiments, tunnel 110 is established between ports of S-ToR 210 and D-ToR 212, which may be links (e.g., ports) connecting D-ToR 212 and NIC 214.
[0059] Tunnel 110 may enclose several flow paths (e.g., 5, 10, etc.), allowing data packet routing to be managed for some or all of the flow paths within tunnel 110, rather than just a single data flow path. In a typical embodiment, TSRC 211 schedules and adds tunnel headers to some or all data packets from flows routed to TDST 213 (the scheduling of packets may be limited by receiving credits from the tunnel destination). Furthermore, TDST 213 may send credit messages and other control messages, such as acknowledgement (ACK) signals and / or negative acknowledgement (NACK) signals, to the source (e.g., TSRC 211). The credit generation rate may be controlled using techniques such as congestion control algorithms (details regarding the credit generation rate and credit process will be described in more detail below). TDST 213 may also reorder packets received from the tunnel and deliver them to the flow destination (e.g., NIC 214, etc.).
[0060] Several data flow paths (e.g., a path from server 202 to TSRC 211, a path from server 204 to TSRC 211, etc.) enter TSRC 211. Here, S-ToR switch 210 and / or TSRC 211 can queue all packets from all attached NIC ports into a tunnel VOQ and schedule packets from the VOQ into tunnel 110. Furthermore, TDST 213 can send credit and control packets (e.g., ACK, NACK, etc.) back to TSRC 211. In some embodiments, in-order delivery is required, and TDST 213 reorders packets received via tunnel 110 and delivers them to NIC 214 in the proper order. NIC 214 can provide the ordered data packets to server 216 to complete the data transfer. In some embodiments, data is provided to S-ToR 210 in-order by the source server (e.g., servers 202, 204, etc.). In some embodiments, NIC 214 receives packets in-order from D-ToR 212. In some embodiments, once a credit is received at a source, this indicates that the source is allowed to send one or more data packets to the destination without causing congestion in the network.
[0061] Now refer to Figure 3 , another implementation of system 200 is presented, according to some embodiments. Figure 3 Tunnel 110 is shown including several spines (e.g., spines 302, 304) and multiple flow paths into and out of TSRC 211 and TDST 213. Spines 302 and 304 can serve as midpoints or connecting nodes between TSRC 211 and TDST 213, allowing several data packets following several data flow paths to connect at one or more nodes during routing to TDST 213. As discussed above, multiple paths can be used to send data packets within tunnel 110.
[0062] Virtual tunnel with multiple streams
[0063] Now refer to Figure 4 , another implementation of system 200 is presented, according to some embodiments. Figure 4 Tunnel 110 is shown as containing two separate flows: an optimistic flow and a scheduled flow. In some embodiments, the optimistic flow is used by high-priority and ultra-low-latency traffic and should primarily be used for the first initial packets (e.g., K packets) of each flow. In some embodiments, the scheduled flow is used for typical low-latency traffic that may / always use credits from the receiver to send packets. In some embodiments, queuing, scheduling, and / or reordering of data packets is performed per flow. Figure 4 Tunnel 110 is shown including areas 402 , 404 and graphs 406 , 408 , and 410 .
[0064] In the first diagram (diagram 406), two flows are shown at TSRC 211. At this stage, S-ToR 210 can create a VOQ for each flow, such that there is an optimistic flow VOQ and a scheduled flow VOQ (this can be seen in diagram 406, left). Diagram 406 also shows a separate output queue for each flow (this can be seen in diagram 406, right). In the second diagram (diagram 408), two flows are shown at area 402. At this stage, the transit switch shows separate output queues. Finally, diagram 410 shows separate reordering contexts for the two flows at TDST 213.
[0065] In some embodiments, several modes for providing data packets via several flows may be considered, such as push mode and pull mode. In push mode, packets can be scheduled without waiting for credit in the optimistic flow. In pull mode, the optimistic flow and / or the scheduling flow only become actively scheduled when credit is received from the tunnel. Therefore, in some embodiments, the scheduling flow is always in pull mode, and the optimistic flow mode changes based on the congestion status reported by the destination (e.g., TDST 213). The optimistic flow may start "fast" (e.g., transmitting at a high rate, etc.) and push traffic with almost no waiting; if congestion occurs, the optimistic flow may return to pull mode. For example, a new tunnel session is started (entering pull mode). TDST 213 signals that the tunnel is congested (entering pull mode). TDST 213 signals that the tunnel congestion has eased (entering push mode).
[0066] Now refer to Figure 5 , another implementation of system 200 is presented, according to some embodiments. Figure 5 A non-limiting use case embodiment can be shown for implementing multiple flows within tunnel 110. For example, the NIC / server (e.g., server 204) tracks the transmitted bytes for each flow and sets the Differentiated Services Code Point (DSCP) bits differently for the first K bytes than for the remaining bytes. S-ToR 210 uses the DSCP field to map the first K bytes of an incoming flow to the optimistic flow and the remaining bytes to the scheduled flow. This process is performed in Figure 5 , where the optimistic flow 502 includes the first K bytes (first 2 bars) from the NIC 208, and the remaining bytes (last 6 bars) from the NIC 208 are provided to the scheduled flow.
[0067] Continuing with the above example, when there is no congestion within tunnel 110, this may result in little to no round trip time (RTT). In some embodiments, when tunnel 110 is uncongested, the first byte of the flow is in an optimistic flow provided in push mode. For short flows (e.g., flows prioritized by completion time, etc.), only the optimistic flow may be used. In some embodiments, more than one optimistic flow and scheduled flows may be implemented, with different VOQs, and thus different scheduling occurs among the flows.
[0068] Now refer to Figure 6 , another implementation of system 200 is presented, according to some embodiments. Figure 6 The reordering process that occurs at D-ToR 214 after providing the data packets to the destination endpoint can be shown. In some embodiments, the ordered packets can be transmitted through tunnel 110, which may require end-to-end (e2e) in-sequence delivery from system 200. Figure 5 The described systems and methods for flow management can occur at the flow level, but can also ensure that data is delivered correctly in order after the data transfer is complete. This can be true even for flows that begin in an optimistic flow and continue in a scheduled flow. Several reordering functions can be implemented, such as intra-flow reordering and cross-flow reordering. In some embodiments, an intra-flow reordering technique is implemented, where packets for each flow are delivered in the same order in which they entered the network. In some embodiments, a cross-flow reordering technique is implemented, where optimistic flow packets are delivered before scheduled flow packets that follow them enter the network.
[0069] Tunnel 110 can provide a reliable communication model for transport. However, link-level transmission between switches can be lossy, and tunnel 110 may not utilize link-level flow control. Link-level flow control mechanisms used in InfiniBand (e.g., Ethernet PFC, credit-based FC, etc.) may not be well-suited for LL traffic. In some embodiments, HOL blocking can introduce unexpected long-tail delays for innocent flows. Deadlock avoidance and recovery mechanisms may increase complexity and latency in system 100 and / or system 200. Furthermore, this may increase the complexity and power consumption of the switch (e.g., requiring more buffers and queues). In some embodiments, when congestion occurs in the network, the HPNT switch may drop packets, as discussed above. End-to-end congestion control mechanisms and load balancing mechanisms ensure that packet drops are rare. When a packet is dropped, the switch can still send the header of the dropped packet to the destination, which then sends a NACK message back to the source to retransmit the packet.
[0070] Still refer to Figure 6 , flows within tunnel 110 (e.g., optimistic flows and scheduled flows) may include sequence numbers to determine packet reordering at the destination and / or for sending ACK and / or NACK messages. In some embodiments, TDST 213 ensures at least one of the following reordering requirements: each flow packet leaves the network in the order in which it entered the network, and a scheduled tunnel packet cannot leave before an optimistic packet that previously entered the network has left.
[0071] In some embodiments, these two conditions ensure that packets for each flow are delivered in order, even for large flows using both flows, optimistic flow packets may not be blocked by scheduled packets, and scheduled packets may not be unnecessarily blocked by optimistic packets. In some embodiments, the source endpoint adds flow sequence numbers to the optimistic packets and the scheduled packets, respectively. In addition, each scheduled packet may carry the last optimistic sequence number (e.g., an anchor sequence number, etc.) used when it entered the network. In some embodiments, the destination endpoint may ensure that flow packets leave the network in the order in which they arrived, and may ensure that each scheduled packet does not leave the network before the optimistic packet and its anchor sequence number have left.
[0072] Now refer to Figure 7 , another implementation of system 200 is presented, according to some embodiments. Figure 7 Credits are shown being provided from the TDST credit generator 702 within the D-ToR 212 to the S-ToR 210. In some embodiments, credits are shaped and / or dropped based on the port rate of the NIC 214 (e.g., and other destination ports connected to the tunnel 110, etc.). The credit queue may be shared among all active tunnels (e.g., flow paths, etc.) for the destination port. Credits may be provided across any and all egress ports and / or flow paths within the tunnel 110. Once the credits arrive at the S-ToR 210, they may be dropped if the selected output queue for the data packet is congested.
[0073] Now refer to Figure 8 , another implementation of system 200 is presented, according to some embodiments. Figure 8 The system includes transit switches 802 and 804, an egress port reordering buffer 806, a TDST credit generator 808, a NIC 810, and a server 812. Figure 9 , according to some embodiments, a diagram 900 showing a timeline of credit generation.
[0074] Now refer to Figure 10 , another embodiment of the system 200 is presented, according to some embodiments. Figure 10 is shown to include a transfer element 1002. In some embodiments, Figure 10 The system is configured to control the buffer side of the S-ToR 210, provide flow-level fairness for large flows, implement configurable bandwidth splitting between short and long flows, provide fairness between ingress port NIC allocations of BW for short flows, and provide fairness between short flows from the NICs where necessary.
[0075] Now refer to Figure 11 , another embodiment of the system 200 is presented, according to some embodiments. Figure 11Shown to include servers 11-2, 11-4, NICs 1106, 1108, and common ToR 1114. In some embodiments, tunnel 110 uses VOQ congestion management for scheduling VOQs containing long flows. In some embodiments, the number of flows in the HPNT VOQ is less than the number of flows in the output queue TOR switch. Now referring to Figure 12 , according to some embodiments, shows another embodiment of the system 200. In some embodiments, when incoming traffic from a NIC port causes congestion, PFC is used to stop traffic from the port.
[0076] Now refer to Figures 11 to 15 According to some embodiments, several other examples and implementations of the system 200 and the implementation of the tunnel 110 therein are disclosed. In network congestion management, Figures 11 to 15 Congestion control update period, credit generation rate update, terminal switch endpoint capabilities, S-ToR congestion management, and support for unordered packets are all exposed and considered.
[0077] The methods and processes disclosed herein may be performed by one or more devices (e.g., servers, processors, controllers, computers, etc.). For example, one or more devices may include a communications interface and processing circuitry including a processor and memory. The processing circuitry may be communicatively connected to the communications interface so that the processing circuitry and its various components can send and receive data via the communications interface. The processor may be implemented as a general-purpose processor, an application-specific integrated circuit (ASIC), one or more field-programmable gate arrays (FPGAs), a set of processing components, or other suitable electronic processing components.
[0078] The communication interface may be or include a wired or wireless communication interface (e.g., a jack, antenna, transmitter, receiver, transceiver, wired terminal, etc.) for communicating data. In various embodiments, communication via the communication interface may be direct (e.g., local wired or wireless communication) or via a communication network (e.g., a WAN, the Internet, a cellular network, etc.). For example, the communication interface may include an Ethernet card and port for sending and receiving data via an Ethernet-based communication link or network. In another example, the communication interface may include a Wi-Fi transceiver for communicating via a wireless communication network. In another example, the communication interface may include a cellular or mobile phone communication transceiver.
[0079] Memory (e.g., memory, memory unit, storage device, etc.) may include one or more devices (e.g., RAM, ROM, flash memory, hard disk storage) for storing data and / or computer code for completing or facilitating the various processes, layers, and modules described in this disclosure. Memory may be or include volatile memory or non-volatile memory. Memory may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in this disclosure. According to an example embodiment, the memory is communicatively connected to the processor via processing circuitry and includes computer code for executing (e.g., by the processing circuitry and / or processor) one or more processes described herein.
[0080] Configuration of the exemplary embodiment
[0081] As utilized herein, the terms "approximately," "about," "substantially," and similar terms are intended to have a broad meaning consistent with common and recognized usage by those of ordinary skill in the art to which the subject matter of this disclosure pertains. Those skilled in the art who examine this disclosure will understand that these terms are intended to allow for a description of certain features described and claimed without limiting the scope of such features to the precise numerical ranges provided. Accordingly, these terms should be interpreted as indicating that insubstantial or inconsequential modifications or variations of the subject matter described and claimed are considered to be within the scope of the present disclosure as set forth in the appended claims.
[0082] It should be noted that the term "exemplary" and variations thereof, as used herein to describe various embodiments, are intended to indicate that such embodiments are possible examples, representations, or illustrations of possible embodiments (and that such terms are not intended to imply that such embodiments are necessarily extraordinary or superlative examples).
[0083] As used herein, the term "coupled" and variations thereof mean that two components are directly or indirectly joined to one another. Such joining may be fixed (e.g., permanent or fixed) or removable (e.g., removable or releasable). Such joining may be achieved using two components that are directly coupled to one another, two components that are coupled to one another using a separate intermediate member and any additional intermediate components that are coupled to one another, or two components that are coupled to one another using an intermediate component that integrally forms a single entity with one of the two components. If "coupled" or variations thereof are modified by additional terms (e.g., directly coupled), the general definition of "coupled" provided above is modified by the plain language meaning of the additional terms (e.g., "directly coupled" means the joining of two components without any separate intermediate components), resulting in a definition that is narrower than the general definition of "coupled" provided above. Such coupling may be mechanical, electrical, or fluidic.
[0084] As used herein, the term "or" is used in its inclusive sense (rather than its exclusive sense), such that when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Unless specifically stated otherwise, conjunction language such as the phrase "at least one of X, Y, and Z" is understood to convey that the elements may be X, Y, Z; X and Y; X and Z; Y and Z; or X, Y, and Z (i.e., any combination of X, Y, and Z). Thus, unless otherwise indicated, such conjunction language is generally not intended to imply that certain embodiments require that at least one of X, at least one of Y, and at least one of Z each be present.
[0085] References herein to element positions (e.g., "top," "bottom," "above," "below") are intended only to describe the orientation of the various elements in the figures. It should be noted that the orientation of the various elements may differ according to other exemplary embodiments, and such variations are intended to be encompassed by the present disclosure.
[0086] The hardware and data processing components used to implement the various processes, operations, illustrative logic, logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device designed to perform the functions described herein, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors and a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuit systems specific to a given function. Memory (e.g., memory, memory unit, storage device) may include one or more devices (e.g., RAM, ROM, flash memory, hard disk storage device) for storing data and / or computer code to accomplish or facilitate the various processes, layers, and modules described in this disclosure. The memory may be or include volatile memory or non-volatile memory and may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described in this disclosure. According to an exemplary embodiment, the memory is communicatively connected to the processor via processing circuitry and includes computer code for executing (e.g., by the processing circuitry or processor) one or more processes described herein.
[0087] The present disclosure contemplates methods, systems, and program products for implementing various operations on any machine-readable medium. The embodiments of the present disclosure may be implemented using an existing computer processor, or by a dedicated computer processor for an appropriate system incorporated for this or another purpose, or by a hard-wired system. Embodiments within the scope of the present disclosure include program products comprising machine-readable media for carrying or having machine-executable instructions or data structures stored thereon. Such machine-readable media may be any available media that can be accessed by a general-purpose or special-purpose computer or other machine with a processor. By way of example, such machine-readable media may include RAM, ROM, EPROM, EEPROM, or other optical disk storage devices, magnetic disk storage devices, or other magnetic storage devices, or any other media that can be used to carry or store desired program code in the form of machine-executable instructions or data structures, and can be accessed by a general-purpose or special-purpose computer or other machine with a processor. The above combinations are also included within the scope of machine-readable media. For example, machine-executable instructions include instructions and data that cause a general-purpose computer, a special-purpose computer, or a dedicated processing machine to perform a specific function or a set of functions.
[0088] Although the figures and description may illustrate a particular order of method steps, the order of such steps may vary from that depicted and described, unless otherwise specified above. In addition, two or more steps may be performed simultaneously, or partially simultaneously, unless otherwise specified above. Such variations may depend, for example, on the software and hardware systems selected and the designer's preferences. All such variations are within the scope of this disclosure. Similarly, software implementations of the described methods may be implemented using standard programming techniques with rule-based logic and other logic to implement the various connection steps, processing steps, comparison steps, and decision steps.
[0089] It is important to note that the construction and arrangement of the various systems (e.g., system 100, system 200, etc.) and methods as shown in the various exemplary embodiments are illustrative only. In addition, any element disclosed in one embodiment may be combined or utilized with any other embodiment disclosed herein. Although only one example of how elements from one embodiment may be combined or utilized in another embodiment is described above, it should be understood that other elements of the various embodiments may be combined or utilized with any other embodiment disclosed herein.
Claims
1. A method for managing network services, the method comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; receiving, at the source endpoint, a plurality of credits from the destination endpoint, the plurality of credits provided via two or more of the plurality of data flow paths; updating, at the source endpoint, a data transmission sequence based on the plurality of credits; providing a plurality of data packets to the destination endpoint based on the updated data transmission sequence; determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path; and At least one data packet of the plurality of data packets is provided to the destination endpoint via the non-congested data flow path.
2. The method according to claim 1, wherein the method further comprises: Establishing a first data flow and a second data flow in the virtual tunnel; using the first data flow to provide a first subset of the plurality of data packets to the destination endpoint; and A second subset of the plurality of data packets is provided to the destination endpoint using the second data flow, wherein the second data flow provides the second subset based on at least one of the plurality of credits.
3. The method according to claim 1, wherein the data transmission sequence comprises: a transmission order for each of the plurality of data packets configured to be transmitted to the destination endpoint; and a selected data flow of the plurality of data flow paths for each of the plurality of data packets.
4. The method of claim 1, wherein establishing the virtual tunnel comprises establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
5. The method of claim 1 , wherein receiving the plurality of credits at the source endpoint comprises: receiving the plurality of credits at a first credit rate; and The plurality of credits are received at a second credit rate in response to a determination of one or more congested data flow paths among the plurality of data flow paths.
6. The method of claim 1 , wherein receiving the plurality of credits at the source endpoint comprises: receiving a shaped credit from the plurality of credits, the shaped credit being shaped using one or more transit switches within the virtual tunnel; and An indication of discarded credits from the plurality of credits is received, the discarded credits being discarded using the one or more transit switches within the virtual tunnel.
7. The method of claim 6, wherein receiving the indication of the discarded credit comprises receiving a negative acknowledgement (NACK) at the source endpoint, the NACK indicating receipt of a header of the discarded credit at the destination endpoint.
8. A method for managing network services, the method comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; receiving, at the source endpoint, a plurality of credits from the destination endpoint, the plurality of credits provided via two or more of the plurality of data flow paths; updating, at the source endpoint, a data transmission sequence based on the plurality of credits; providing a plurality of data packets to the destination endpoint based on the updated data transmission sequence, wherein providing the plurality of credits to the source endpoint comprises: generating the plurality of credits at a generation rate based on a port rate of a destination network interface controller (NIC) configured to receive the plurality of data packets from the destination endpoint and provide the plurality of data packets to a destination server; determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path; and At least one data packet of the plurality of data packets is provided to the destination endpoint via the non-congested data flow path.
9. A method for managing network services, the method comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; receiving, at the source endpoint, a plurality of credits from the destination endpoint, the plurality of credits provided via two or more of the plurality of data flow paths; updating, at the source endpoint, a data transmission sequence based on the plurality of credits; providing a plurality of data packets to the destination endpoint based on the updated data transmission sequence; Establishing a first data flow and a second data flow in the virtual tunnel; using the first data flow to provide a first subset of the plurality of data packets to the destination endpoint without using credits; using the second data flow to provide a second subset of the plurality of data packets to the destination endpoint, wherein the second data stream provides the second subset based on at least one of the plurality of credits, generating, at the source endpoint, a first virtual sequential queue (VOQ) for the first data flow and a second virtual sequential queue (VOQ) for the second data flow; and Two or more of the multiple data packets are scheduled to the first VOQ or the second VOQ based on a flow path criterion, wherein the flow path criterion indicates at least one of a priority requirement or a latency requirement of the two or more data packets, wherein the data transmission sequence includes the first VOQ and the second VOQ.
10. The method according to claim 9, further comprising: receiving the first subset and the second subset at the destination endpoint; The plurality of data packets are reordered at the destination endpoint by combining the first subset and the second subset in the same order that provides the first subset and the second subset.
11. The method according to claim 9, further comprising: receiving the first subset and the second subset at the destination endpoint; reordering the plurality of data packets at the destination endpoint by combining the first subset and the second subset based on a cross-flow reordering, wherein the cross-flow reordering is configured to notify the destination endpoint that the first subset is received before the second subset, The first data flow is an optimistic flow, and the second data flow is a scheduled flow.
12. One or more non-transitory computer-readable media having computer-executable instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; receiving, at the source endpoint, a plurality of credits from the destination endpoint, the plurality of credits provided via two or more of the plurality of data flow paths; updating, at the source endpoint, a data transmission sequence based on the plurality of credits; providing a plurality of data packets to the destination endpoint based on the updated data transmission sequence; determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path; and At least one data packet of the plurality of data packets is provided to the destination endpoint via the non-congested data flow path.
13. The medium of claim 12, wherein providing the plurality of credits to the source endpoint comprises: The plurality of credits is generated at a generation rate that is based on a port rate of a destination network interface controller (NIC) configured to receive the plurality of data packets from the destination endpoint and provide the plurality of data packets to a destination server.
14. The medium of claim 12, wherein the operations further comprise: Establishing a first data flow and a second data flow in the virtual tunnel; using the first data flow to provide a first subset of the plurality of data packets to the destination endpoint; and A second subset of the plurality of data packets is provided to the destination endpoint using the second data flow, wherein the second data flow provides the second subset based on at least one of the plurality of credits.
15. The medium of claim 14, wherein the operations further comprise: generating, at the source endpoint, a first virtual sequential queue (VOQ) for the first data flow and a second virtual sequential queue (VOQ) for the second data flow; and Two or more of the multiple data packets are scheduled to the first VOQ or the second VOQ based on a flow path criterion, wherein the flow path criterion indicates at least one of a priority requirement or a latency requirement of the two or more data packets, wherein the data transmission sequence includes the first VOQ and the second VOQ.
16. The medium of claim 15, further comprising: receiving the first subset and the second subset at the destination endpoint; The plurality of data packets are reordered at the destination endpoint by combining the first subset and the second subset in the same order that provides the first subset and the second subset.
17. The medium of claim 15, further comprising: receiving the first subset and the second subset at the destination endpoint; reordering the plurality of data packets at the destination endpoint by combining the first subset and the second subset based on a cross-flow reordering, wherein the cross-flow reordering is configured to notify the destination endpoint that the first subset is received before the second subset, The first data flow is an optimistic flow, and the second data flow is a scheduled flow.
18. The medium of claim 12, wherein the data transmission sequence comprises: a transmission order for each of the plurality of data packets configured to be transmitted to the destination endpoint; and a selected data flow of the plurality of data flow paths for each of the plurality of data packets.
19. The medium of claim 12, wherein establishing the virtual tunnel comprises establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
20. An apparatus for managing network traffic, the apparatus comprising one or more processors and a memory, the memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: establishing a virtual tunnel between a source endpoint and a destination endpoint, the virtual tunnel comprising a plurality of data flow paths, each of the plurality of data flow paths connecting the source endpoint and the destination endpoint; providing a plurality of credits to the source endpoint via the destination endpoint, the plurality of credits being provided via two or more of the plurality of data flow paths; updating, at the source endpoint, a data transmission sequence based on the plurality of credits; providing a plurality of data packets to the destination endpoint based on the updated data transmission sequence; determining that at least one credit received at the source endpoint indicates a non-congested data flow path among the plurality of data flow paths, wherein the at least one credit is provided to the source endpoint via the non-congested data flow path; and providing at least one data packet of the plurality of data packets to the destination endpoint via the non-congested data flow path, Wherein establishing the virtual tunnel comprises establishing the virtual tunnel within a pre-existing network, the pre-existing network operating under a pre-existing transport layer protocol.
Citation Information
Patent Citations
Network device architecture for consolidating input / output and reducing latency
CN101040489A
Shared-Credit Arbitration Circuit
US20180159789A1