Replay Microarchitecture for Ethernet

JP2025526905A5Pending Publication Date: 2025-12-01TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025508895
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-19
Filing Date
2023-08-17
Publication Date
2025-12-01

AI Technical Summary

Technical Problem

Existing Ethernet-based networks face challenges in achieving high bandwidth, low latency, loss resilience, and distributed control with minimal software overhead, particularly in high-performance computing and AI training data centers.

Method used

A hardware-implementable flow control protocol, such as the Tesla Transport Protocol (TTP), operates over lossy Ethernet networks with minimal CPU involvement, using hardware mechanisms like state machines, hardware replay architectures, and hardware link timers to achieve single-digit microsecond latency.

Benefits of technology

The protocol enables low-latency, lossy Ethernet communication with minimal software intervention, facilitating efficient communication in high-performance computing and AI training systems by maintaining packet order and managing link transitions in hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for communicating over Ethernet-based networks using a transport layer without the assistance of software control mechanisms. In some embodiments, a first node includes a hardware replay architecture configured to replay packets transmitted under a transport layer hardware-only Ethernet protocol. The hardware replay architecture can include a local storage device configured to store the packets transmitted under the transport layer hardware-only Ethernet protocol. The hardware replay architecture can further include a linked list stored in the local storage device and configured to track the order of packets for transmission to another node.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a non-provisional application claiming priority to U.S. Provisional Patent Application No. 63 / 373,016, entitled "TRANSPORT PROTOCOL FOR ETHERNET," filed on August 19, 2022, the technical disclosure of which is incorporated herein by reference in its entirety for all purposes. This application is a non-provisional application claiming priority to U.S. Provisional Patent Application No. 63 / 503,349, entitled "TRANSPORT PROTOCOL FOR ETHERNET," filed on May 19, 2023, the technical disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0002] The present disclosure relates to systems and methods for facilitating communication over a network, and more particularly, to hardware-implementable flow control protocols for communication over Ethernet-based networks. [Background technology]

[0003] The Institute of Electrical and Electronics Engineers (IEEE) has provided various standards for local area networks (LANs), including the IEEE 802 standard, commonly known as Ethernet, collectively known as IEEE 802.3. The IEEE 802.3 Ethernet standard includes specifications for the physical media interface (such as Ethernet cable, optical fiber, or backplane), but not for communication flow control. Protocols such as TCP / IP, RoCE, or InfiniBand can accelerate fabric flow control. While the TCP / IP protocol has latency measured in milliseconds, RoCE or InfiniBand generally have lossless and scaling specifications that can overly constrain the system.

[0004] As high-performance computing (HPC) and artificial intelligence (AI) training data centers become more prevalent, communication network fabrics with high bandwidth, low latency, loss resilience at scale, distributed control, and as little software overhead as possible are desired. Therefore, it may be desirable to develop a network flow control protocol that can operate over lossy Ethernet-based networks with little or no central processing unit (CPU) involvement, while achieving lower latency than existing Ethernet-based networks. Summary of the Invention [Means for solving the problem]

[0005] The systems, methods, and devices of the present disclosure each have several innovative embodiments, no single one of which is solely responsible for all of the desirable attributes disclosed herein. Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below.

[0006] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node including one or more processors configured to implement a transport layer hardware-only Ethernet protocol.

[0007] In some aspects, the technology described herein relates to the first node, wherein the Ethernet protocol is lossy.

[0008] In some aspects, the techniques described herein relate to a first node, wherein the one or more processors are further configured to implement a hardware replay architecture to replay packets transmitted over the first link to the second node, the packets being stored in a local storage device of the first node, and an order of the packets for replay being specified in a linked list.

[0009] In some aspects, the techniques described herein relate to a first node configured to transmit packets to a second node with single-digit microsecond latency.

[0010] In some aspects, the techniques described herein involve one or more processors configured to implement a state machine, with respect to a first node, configured to operate in an open state where a link is open between the first node and a second node, transition from the open state to an intermediate closed state, and transition from the intermediate closed state to the closed state to close the link in response to receiving a close acknowledgement from the second node.

[0011] In some aspects, the technology described herein relates to a first node further including an Ethernet port.

[0012] In some aspects, the techniques described herein relate to a first node, wherein one or more processors are configured to determine to replay a packet on a link between the first node and a second node based on timing and status information associated with the link stored in a first-in, first-out (FIFO) memory, wherein entries in the FIFO memory are accessed according to ticks of hardware link timers associated with the plurality of links.

[0013] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node including one or more processors configured to implement a Layer 2 hardware-only Ethernet protocol.

[0014] In some aspects, the techniques described herein include a hardware-only architecture involving a first node, wherein one or more processors are configured to replay packets transmitted over a first link to a second node.

[0015] In some aspects, the techniques described herein relate to a first node, wherein the one or more processors are further configured to determine to replay the packet over a link associated with the first node based on timing and status information associated with the link stored in a first-in, first-out (FIFO) memory accessed based on ticks of timers associated with the plurality of links.

[0016] In some aspects, the techniques described herein relate to a first node configured to open and close a link with a second node in an Ethernet-based network, the first node including state machine hardware configured to operate in an open state where the link is open between the first node and the second node, transition from the open state to an intermediate closed state, and transition from the intermediate closed state to the closed state to close the link in response to receiving a close acknowledgement from the second node, the first node configured to operate in a lossy network.

[0017] In some aspects, the techniques described herein relate to a first node, wherein the state machine hardware implements a flow control protocol for a transport layer solely in hardware.

[0018] In some aspects, the techniques described herein involve a first node having a latency associated with a flow control protocol of less than 10 microseconds.

[0019] In some aspects, the techniques described herein relate to a first node, wherein the state machine hardware is configured to transition from a closed state to an intermediate open state and transition from the intermediate open state to an open state.

[0020] In some aspects, the techniques described herein relate to a first node, where state machine hardware transitions from an open state to an intermediate closed state in response to sending a request to close the link to a second node or receiving a request to close the link from the second node.

[0021] In some aspects, the techniques described herein relate to a first node, where state machine hardware transitions from an intermediate closed state to a closed state in response to sending an acknowledgment to close the link to a second node.

[0022] In some aspects, the techniques described herein relate to a first node, where the state machine hardware transitions from an intermediate closed state to a closed state without waiting a period of time.

[0023] In some aspects, the techniques described herein relate to a first node that, in an open state, does not retransmit a packet until a negative acknowledgment for the packet is received from the second node or until a predetermined timeout period expires without receiving a negative acknowledgment for the packet.

[0024] In some aspects, the techniques described herein relate to a first node, wherein in an open state, the first node transmits up to N packets without pausing, where N is limited by the size of physical memory allocated to the first node.

[0025] In some aspects, the techniques described herein further include, with respect to the first node, hardware link timers associated with the plurality of links, and a hardware replay architecture configured to replay packets solely in hardware.

[0026] In some aspects, the technology described herein relates to a first node including a hardware replay architecture configured to replay packets transmitted over a first link to a second node using an Ethernet protocol, the hardware replay architecture including: a local storage configured to store a linked list including the packets, the linked list maintaining an order of the packets for transmission to the second node; and logic configured to: determine to replay a first one of the packets in response to at least one of (a) receiving a negative acknowledgment of the first packet from the second node or (b) a timeout associated with the first packet; and retire a second one of the packets in response to receiving an positive acknowledgment of the second packet from the second node, wherein the Ethernet protocol is lossy.

[0027] In some aspects, the techniques described herein relate to a first node, wherein logic circuitry includes a plurality of pipeline stages, and wherein the logic circuitry determines, in a first pipeline stage of the plurality of pipeline stages, to process data associated with a first link rather than a second link between the first node and a second node.

[0028] In some aspects, the techniques described herein relate to a first node, wherein logic circuitry determines to replay a first packet in a second pipeline stage of a plurality of pipeline stages.

[0029] In some aspects, the techniques described herein relate to a first node, wherein the logic circuitry determines, in a second pipeline stage of the plurality of pipeline stages, to replay a third packet of the packets and a first packet of the packets based on an order of the packets maintained by a linked list.

[0030] In some aspects, the techniques described herein relate to a first node, wherein logic circuitry determines to process data associated with the first link rather than the second link based on a link pointer, and wherein the logic circuitry updates the link pointer to point to the second link in a third pipeline stage of the plurality of pipeline stages.

[0031] In some aspects, the techniques described herein relate to a first node, wherein the first node and a second node are in an Ethernet-based network, and the first node communicates with the second node through an Ethernet switch.

[0032] In some aspects, the technology described herein relates to a first node, the first node including a network interface processor (NIP) and a high-bandwidth memory (HBM), wherein the bandwidth of the HBM is at least 1 gigabyte.

[0033] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node including one or more processors configured to implement a transport layer hardware-only Ethernet protocol, the transport layer hardware-only Ethernet protocol being lossy, and the one or more processors including a hardware replay architecture configured to replay packets transmitted under the transport layer hardware-only Ethernet protocol.

[0034] In some aspects, the techniques described herein relate to a first node, wherein the hardware replay architecture includes a local storage device configured to store packets transmitted under an Ethernet protocol in transport layer hardware only.

[0035] In some aspects, the techniques described herein involve a hardware replay architecture for a first node that includes a linked list stored in local storage and configured to track the order of packets for transmission to another node, each element of the linked list corresponding to a respective one of the packets stored in the local storage.

[0036] In some aspects, the techniques described herein relate to a first node, wherein the hardware replay architecture is configured to transmit packets in an order corresponding to the linked list.

[0037] In some aspects, the techniques described herein relate to a first node, wherein the hardware replay architecture is configured to store: a first pointer configured to point to a first element of a linked list, wherein the first pointer indicates not to replay a first packet of packets corresponding to the first element of the linked list; and a second pointer configured to point to a second element of the linked list, wherein the second pointer indicates to replay a second packet of packets corresponding to the second element of the linked list.

[0038] In some aspects, the techniques described herein relate to a first node, wherein the hardware replay architecture replays a second packet and one or more packets following the second packet according to the order of the packets for transmission.

[0039] In some aspects, the techniques described herein relate to a first node, wherein a hardware replay architecture causes a local storage device to discard a first packet and one or more packets preceding a second packet according to an order of packets for transmission.

[0040] In some aspects, the technology described herein relates to a computer-implemented method implemented in a first node for replaying packets transmitted over a first link to a second node using an Ethernet protocol, the computer-implemented method including: storing a linked list including packets, the linked list maintaining an order of the packets for transmission to the second node; determining to replay a first one of the packets in response to at least one of (a) receiving a negative acknowledgment of the first packet from the second node or (b) a timeout associated with the first packet; and retiring a second one of the packets in response to receiving an positive acknowledgment of the second packet from the second node, wherein the Ethernet protocol is lossy.

[0041] In some aspects, the techniques described herein relate to a computer-implemented method, wherein a first node includes a hardware replay architecture including a plurality of pipeline stages, and the hardware replay architecture determines, in a first pipeline stage of the plurality of pipeline stages, to process data associated with the first link rather than the second link.

[0042] In some aspects, the technology described herein relates to a computer-implemented method, wherein a hardware replay architecture determines to replay a first packet at a second pipeline stage of a plurality of pipeline stages.

[0043] In some aspects, the technology described herein relates to a computer-implemented method, wherein a hardware replay architecture determines, in a second pipeline stage of a plurality of pipeline stages, to replay a third one of the packets and a first one of the packets based on an order of the packets maintained by a linked list.

[0044] In some aspects, the technology described herein relates to a computer-implemented method, wherein a first node and a second node are in an Ethernet-based network, and the first node communicates with the second node through an Ethernet switch.

[0045] In some aspects, the technology described herein relates to a computer-implemented method, wherein a first node includes a network interface processor (NIP) and a high-bandwidth memory (HBM), wherein the bandwidth of the HBM is at least 1 gigabyte.

[0046] In some aspects, the techniques described herein relate to a first node for transmitting packets in an Ethernet-based network, the first node including one or more processors including: a first-in, first-out (FIFO) memory configured to store timing and status information associated with a plurality of links, wherein the first node is configured to transmit packets to one or more other nodes over the plurality of links using an Ethernet protocol; a timer configured to tick (count forward) according to a time period, wherein the timer is associated with the plurality of links; and logic circuitry configured to access an entry in the FIFO memory based on each tick on the timer and determine to replay at least one packet associated with the first link based on the timing and status information associated with the first link of the plurality of links, wherein the Ethernet protocol is lossy.

[0047] In some aspects, the techniques described herein relate to a first node, wherein the logic circuitry is configured to access entries of the FIFO memory in a round-robin manner.

[0048] In some aspects, the techniques described herein relate to a first node, wherein the timer is configured to adjust the time period based on a number of active links associated with an entry in the FIFO memory, the active link being included in the plurality of links.

[0049] In some aspects, the techniques described herein relate to a first node, wherein logic circuitry is configured to determine, based on timing and status information associated with a second link of the plurality of links, to retire a packet associated with the second link.

[0050] In some aspects, the techniques described herein relate to a first node, wherein a packet associated with a second link is stored in a local storage device of the first node, and in response to determining, by the logic circuitry, to retire the packet associated with the second link, the local storage device discards the packet associated with the second link.

[0051] In some aspects, the techniques described herein relate to a first node, wherein the logic circuitry is configured to determine to close a second link of the plurality of links based on timing and status information associated with the second link.

[0052] In some aspects, the techniques described herein relate to a first node, where timing and status information associated with a first link of a plurality of links indicates that an acknowledgment of receiving at least one packet associated with the first link has not been received by the first node for a threshold duration for replaying the packet.

[0053] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node including one or more processors configured to implement a transport layer hardware-only Ethernet protocol, the transport layer hardware-only Ethernet protocol being a lossy protocol, and the one or more processors including a hardware link timer configured to determine packets transmitted under the transport layer hardware-only Ethernet protocol for replay.

[0054] In some aspects, the techniques described herein relate to a first node transmitting a first plurality of packets over a first link and a second plurality of packets over a second link according to a transport layer hardware-only Ethernet protocol, wherein the hardware link timer includes a first-in, first-out (FIFO) memory configured to store timing and status information associated with the first link in a first entry of the FIFO memory and timing and status information associated with the second link in a second entry of the FIFO memory.

[0055] In some aspects, the techniques described herein relate to a first node, wherein a hardware link timer includes a timer associated with a plurality of links that ticks according to an arbitrary time period, and wherein the hardware link timer accesses entries in a FIFO memory on round-robin ticks of the timer, the entries including a first entry and a second entry.

[0056] In some aspects, the techniques described herein relate to a first node, wherein a hardware link timer is configured to adjust a time period based on a number of active links associated with an entry in a FIFO memory, the active links including a first link and a second link.

[0057] In some aspects, the techniques described herein relate to a first node, wherein a hardware link timer is configured to determine to replay at least some of a first plurality of packets based on timing and status information associated with a first link stored in a first entry of a FIFO memory, and to determine to retire a second plurality of packets based on timing and status information associated with a second link stored in a second entry of the FIFO memory.

[0058] In some aspects, the techniques described herein relate to a first node, wherein the second plurality of packets is stored in a local storage device of the first node, and in response to a hardware link timer determining to retire the second plurality of packets, the local storage device discards the second plurality of packets.

[0059] In some aspects, the techniques described herein relate to a first node, wherein timing and status information associated with a first link indicates that an acknowledgment of receipt of one of a first plurality of packets has not been received by the first node for a threshold duration for replaying the packet.

[0060] In some aspects, the technology described herein relates to a computer-implemented method implemented in a first node in an Ethernet-based network, the computer-implemented method including: storing timing and status information associated with a plurality of links in a first-in, first-out (FIFO) memory of the first node, the first node being configured to transmit packets to one or more other nodes over the plurality of links using an Ethernet protocol; accessing an entry in the FIFO memory based on each tick of a hardware timer; and determining to replay at least one packet associated with the first link based on the timing and status information associated with a first link of the plurality of links, the Ethernet protocol being lossy.

[0061] In some aspects, the technology described herein relates to a computer-implemented method, in which entries in a FIFO memory are accessed in a round-robin fashion.

[0062] In some aspects, the technology described herein relates to a computer-implemented method further including adjusting a time period of the hardware timer based on a number of active links associated with an entry in the FIFO memory, the active link being included in the plurality of links.

[0063] In some aspects, the techniques described herein relate to a computer-implemented method further including determining to retire a packet associated with a second link of the plurality of links based on timing and status information associated with the second link.

[0064] In some aspects, the technology described herein relates to a computer-implemented method further including replaying at least one packet associated with the first link.

[0065] In some aspects, the technology described herein relates to a computer-implemented method, wherein timing and status information associated with a first link of a plurality of links indicates that acknowledgment of receipt of at least one packet associated with the first link has not been received by the first node for a threshold duration for replaying the packet.

[0066] In some aspects, the technology described herein relates to the above and all of the above-mentioned embodiments. [Brief explanation of the drawings]

[0067] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate examples of the subject matter described herein and not to limit its scope.

[0068] Embodiments of the present disclosure will be described with reference to the accompanying drawings, in which like reference numerals refer to like elements.

[0069] [Figure 1A] 1 is a table illustrating exemplary protocols operating at various layers of the Open Systems Interconnection (OSI) model. [Figure 1B] 1 is a table illustrating exemplary protocols operating at various layers of the Open Systems Interconnection (OSI) model.

[0070] [Figure 2] FIG. 1 illustrates an example state machine for opening and closing links between nodes implementing the Tesla Transport Protocol (TTP), according to an embodiment of the present disclosure.

[0071] [Figure 3A-B] 3 is an exemplary timing diagram illustrating the transmission and reception of packets between two devices implementing TTP, according to an embodiment of the present disclosure.

[0072] [Figure 4]FIG. 2 is an exemplary schematic block diagram of a node implementing a TTP, according to an embodiment of the present disclosure.

[0073] [Figure 5] 3A-3C illustrate exemplary headers of packets transmitted or received according to a TTP, in accordance with an embodiment of the present disclosure.

[0074] [Figure 6] FIG. 1 illustrates an exemplary network and computer environment in which embodiments of the present disclosure may be implemented.

[0075] [Figure 7A] FIG. 2 illustrates opcodes for various types of TTP packets, according to some embodiments of the present disclosure. [Figure 7B] FIG. 2 illustrates opcodes for various types of TTP packets, according to some embodiments of the present disclosure.

[0076] [Figure 8] FIG. 1 illustrates an exemplary physical storage device for storing packets for replaying packets transmitted and / or received under a lossy protocol such as TTP, in accordance with some embodiments of the present disclosure.

[0077] [Figure 9] FIG. 1 illustrates an example data structure (e.g., a linked list) for tracking and maintaining a transmission order for transmitting and replaying packets, in accordance with some embodiments of the present disclosure.

[0078] [Figure 10] FIG. 1 is an example block diagram of at least a portion of a hardware replay architecture for replaying packets transmitted over multiple links, in accordance with some embodiments of the present disclosure.

[0079] [Figure 11]FIG. 10 is an example block diagram of a hardware link timer that implements a timeout checking mechanism for replaying packets without software assistance, in accordance with some embodiments of the present disclosure.

[0080] [Figure 12] FIG. 1 illustrates an example routine for replaying packets transmitted from a node, consistent with some embodiments of the present disclosure.

[0081] [Figure 13] FIG. 10 illustrates an example routine for determining whether to replay one or more links associated with a node. DETAILED DESCRIPTION OF THE INVENTION

[0082] The following detailed description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be implemented in a multitude of different ways, as defined and encompassed by, for example, the claims. This description refers to the drawings, in which like reference numbers and / or terminology may indicate identical or functionally similar elements. It will be understood that the elements depicted in the drawings are not necessarily drawn to scale. It will also be understood that certain embodiments may include more elements than shown in the drawings and / or a subset of the elements depicted in the drawings. Furthermore, some embodiments may incorporate any suitable combination of features from two or more drawings. Headings are provided for convenience only and do not affect the scope or meaning of the claims.

[0083] Generally described, one or more aspects of the present disclosure correspond to systems and methods that use hardware mechanisms (e.g., without software assistance) to control network traffic flow. More specifically, some embodiments of the present disclosure disclose flow control protocols that are compatible with Ethernet standards and can be implemented via hardware circuitry to achieve low latency, such as single-digit microsecond latencies. In some embodiments, single-digit microsecond latencies are achieved at least in part through the use of hardware-controlled state machines to streamline the opening and closing of communication links between nodes in a network. Furthermore, the disclosed flow control protocols (e.g., Tesla Transport Protocol (TTP)) may limit the number of packets transmitted / retransmitted over an established link and / or the duration of a wait period before transitioning to the next state of the hardware-controlled state machine, thereby contributing to achieving low communication latency. Advantageously, the flow control protocols disclosed herein enable pure hardware implementations up to Layer 4 (the transport layer) of the Open Systems Interconnection (OSI) model.

[0084] Some aspects of the present disclosure relate to flow control designed to run solely in hardware. Such flow control can be implemented without software flow control or central processing unit (CPU) / kernel involvement. This enables IEEE 802.3 Ethernet functionality where latency is only physically limited or prime. For example, single-digit microsecond latency can be achieved.

[0085] TeslaTransport Protocol over Ethernet (TTP) is a hardware-only Ethernet flow control protocol that can be implemented up to the transport layer in the OSI model. Layer 2 (L2) Ethernet flow control can be implemented solely in hardware. Layer 3 and / or Layer 4 Ethernet flow control can also be implemented solely in hardware. Link control, timers, congestion, and replay functions can be implemented in hardware. TTP can be implemented in network interface processors and network interface cards. TTP allows for full I / O batching. TTP is a lossy protocol. Lossy protocols can recover lost data. For example, lossy protocols can recover lost or corrupted packets by replaying (e.g., retransmitting) them until receipt is acknowledged.

[0086] The L2 header, state machine, and opcodes in this disclosure can define this hardware-only protocol (eg, TTP) that can recover from lost packets in an N-to-N link set.

[0087] Furthermore, some embodiments of the present disclosure disclose a hardware replay architecture (e.g., microarchitecture) that can replay transmitted and / or received packets under lossy protocols such as TTP. As described above, TTP (or TTPoE) is a hardware-only Ethernet flow control protocol. TTP can facilitate the implementation of ultra-low latency (e.g., single-digit microsecond) fabrics for HPC and / or AI training systems. To implement lossy Ethernet flow control protocols without the assistance of software control mechanisms, some aspects of the present disclosure describe a hardware replay architecture that can buffer, hold, acknowledge, and / or replay packets so that lost or corrupted packets can be replayed and recovered until receipt is acknowledged.

[0088] To replay packets transmitted and / or received according to a lossy Ethernet protocol such as TTP using hardware-only resources, some embodiments of the disclosed hardware replay architecture utilize physical storage and data structures to store packets transmitted and / or received on different links, particularly to maintain the order of transmitted packets when replay occurs. In some embodiments, the physical storage may be any type of local storage or cache (e.g., a low-level cache) that stores, buffers, or holds packets associated with one or more links. The physical storage may be size-limited, such as having a size in units of megabytes (MB) or kilobytes (KB). In some embodiments, the data structure may include one or more linked lists, each of which may record and / or track the order of transmitted packets for a link established between a first communication node and a second communication node. Advantageously, implementing a replay mechanism for a lossy protocol using a hardware replay architecture that employs size-limited physical storage and linked lists that keep track of packet order on various links allows communication nodes to operate in a TTP-compliant manner under limited hardware resources (e.g., when virtual processing or storage resources are unavailable).

[0089] Furthermore, some embodiments of the present disclosure relate to a hardware link timer that implements timeout checking without the assistance of a software control mechanism. Rather than employing multiple timers to track timeouts for each link, some aspects of the present disclosure describe a hardware link timer that employs a single timer that can track timeouts across multiple links through cooperation with a first-in, first-out (FIFO) memory. More specifically, entries in the FIFO memory may store link status and / or timer information, and the hardware link timer may access the entries in the FIFO memory in a round-robin manner to determine whether packets associated with the link can be discarded or need to be saved. If the hardware link timer determines that packets associated with the link can be discarded, under constrained hardware resources, more space is available to store packets associated with another link. If the hardware link timer determines that one or more packets associated with the link should be saved, the saved packets associated with the link may enable the communication node hosting the hardware link timer to replay the saved packets.

[0090] Ethernet is an established standard technology for wired communications. In recent years, Ethernet has also been used in various vehicle applications in the automotive industry. Typically, the latency associated with Ethernet communications ranges from hundreds of microseconds to over a few milliseconds. In addition to physical limitations (e.g., the speed at which signals travel over the communications medium), the complexity of the associated protocols for controlling data flow over Ethernet typically presents another bottleneck in terms of latency. For example, software-controlled management may generally be desirable to follow Transport Control Protocol (TCP) or User Datagram Protocol (UDP). Software-controlled or software-assisted network flow control management tends to increase the latency associated with communications.

[0091] However, such limitations on latency may make Ethernet technology less suitable for applications such as high-performance computing (HPC) and artificial intelligence (AI) training data centers, where latency within one microsecond may be desirable to improve system performance and efficiency. Protocols such as remote direct memory access (RDMA) over converged Ethernet (RoCE) or InfiniBand over Ethernet (IBoE) can help reduce latency, but they may involve significant system design complexity or cost. For example, RoCE or InfiniBand have lossless network and scaling specifications that may be difficult to implement. Implementing RoCE or InfiniBand may also result in significant software control overhead or involve a centralized token control mechanism with limited bandwidth. Furthermore, systems implementing RoCE or InfiniBand may be prone to sleep (e.g., frequently paused).

[0092] To address at least some of the above problems, some embodiments of the present disclosure disclose a flow control protocol (e.g., Tesla Transport Protocol (TTP)) operable on an Ethernet-based network or a peer-to-peer (P2P) network. The flow control protocol may be fully implementable via hardware without the assistance of software control mechanisms to reduce communication latency to within single-digit microseconds. The flow control protocol may be implemented without the intervention of software resources, such as a general-purpose processor or central processing unit that executes computer-readable instructions or an operating system. Furthermore, due to several mechanisms built into the flow control protocol (e.g., one or more of a limit on the number of packets that cannot be sent before pausing, a limit on the number of links that can be established simultaneously, a hardware-controlled state machine, or a proposed header format for packets sent or received according to the TTP), no virtual resources (e.g., a virtualized processor or memory) are required to implement the flow control protocol.

[0093] In some embodiments, the state machine expedites transitions between various states to open and close communication links between nodes. The state machine may be maintained and implemented by hardware without the involvement of software, firmware, drivers, or other types of programmable instructions. Thus, transitions between various states of the state machine may be accelerated compared to implementations of other protocols that utilize software support, such as the Transmission Control Protocol (TCP) applicable to Ethernet-based networks.

[0094] In some embodiments, the headers of packets transmitted and received according to the TTP (e.g., the TTP header) support operation at Layers 2 through 4 of the Open Systems Interconnection (OSI) model. The headers may include fields that are recognizable by existing Ethernet-based network devices or infrastructure. Thus, compatibility of the TTP with existing Ethernet standards may be maintained. Advantageously, this allows for economical use of existing infrastructure and / or supply chains, provides many system design options, and can achieve system-level reuse or redundancy.

[0095] As described above, a node may implement or operate under TTP (e.g., communicate with another node using TTP) using purely hardware resources without the assistance of software-controlled mechanisms. To operate under TPP using purely hardware resources, a node may employ a hardware replay architecture to replay packets that may be lost in transmission. In some embodiments, the hardware replay architecture may include local storage, such as one or more caches, for storing packets transmitted and / or received over one or more links, each of which may be opened or closed according to the TTP. In contrast to protocols such as TCP or UDP, where virtual resources with nearly unlimited processing power and storage capacity are typically available through software-controlled network flow control management, the size of a cache (e.g., a low-level cache) employed by a hardware replay architecture in a node operating under TTP may be limited in size. For example, the size of the cache may be in units of megabytes (MB) or kilobytes (KB), such as 256 KB. In order to communicate with each other over one or more links established according to a lossy communication protocol such as TTP with limited local storage, packets associated with one or more links should be appropriately managed (e.g., saved or discarded) so that some packets are saved for replay while other packets are discarded to avoid cache overflow.

[0096] In some examples, a first node transmitting N packets to a second node using a link established under the TTP may utilize a cache to store the N packets, where N is any positive integer that may be limited by the size of the cache. The first node may continuously transmit some or all of the N packets to the second node as long as constraints from the TTP and / or network conditions permit. To accommodate replay packets, the cache may continue to store already transmitted packets until an acknowledgment that the packets were received is received from the second node. Once an acknowledgment that the packets were received is received, the cache may discard the packets to make room for storing packets to be transmitted over the link between the first node and the second node or other nodes, or over other links. In contrast, if a negative acknowledgment of the packets is received (e.g., the second node notifies the first node that the packets were not received) or if a timeout occurs without receiving an acknowledgment or negative acknowledgment of receiving the packets from the second node, the first node may replay the packets (e.g., retransmit the packets to the second node). In connection with replaying a packet, the first node may discard other packets for which an acknowledgment of receipt has been received.

[0097] In some examples, the order in which packets are transmitted and replayed may be the same. For example, a first node may transmit N packets in a particular order (e.g., packet 1, packet 2 through packet N). If the fifth packet is replayed (e.g., in response to the first node receiving a negative acknowledgment for the fifth packet from the second node, in response to a timeout occurring without receiving any acknowledgment, or an acknowledgment that the fifth packet was received), and an acknowledgment for packets 1 through 4 is received, the cache may discard packets 1 through 4 rather than the fifth packet, so that the node may replay the fifth packet. Additionally and / or optionally, when replaying the fifth packet, the first node may replay packets transmitted after the fifth packet (assuming N>5) in the same order as they were previously transmitted.

[0098] In some examples, the hardware replay architecture of the first node may utilize a linked list in conjunction with a cache to maintain ordering between the initial transmission of some or all of the N packets and any subsequent replays. The linked list may include N elements, each element including a respective one of the N packets and a reference to the next element corresponding to the next packet. When transmitting and / or replaying the N packets, the hardware replay architecture may further utilize one or more pointers to one or more elements in the linked list to determine whether a packet should be retained for replay or can be discarded (e.g., to conserve storage resources). For example, if N is 9 (e.g., 9 packets sent from a first node to a second node), in the linked list, the first element may include the first packet and the first reference, the first reference may point to the second element, the second element may include the second packet and the second reference, the second reference may point to the third element, the eighth element may include the eighth packet and the eighth reference, the eighth reference may point to the ninth element, and the ninth element may include the ninth packet. The hardware replay architecture may maintain and update three pointers that point to the three elements. Suppose a node sends packets 1 through 9 and receives an acknowledgment from the second node that receives packets 1 through 7 but does not receive packets 8 and 9. The first pointer may point to the first element of the linked list, the second pointer may point to the eighth element of the linked list, and the third pointer may point to the ninth element of the linked list. Thus, with the hardware replay architecture, the cache may discard packets and replay packets based on the three pointers. More specifically, the cache may replay the packet pointed to by the second pointer (e.g., the eighth packet) through the packet pointed to by the third pointer (e.g., the ninth packet) and discard the remaining packets (e.g., the packet pointed to by the first pointer before the packet pointed to by the second pointer).Additionally, and optionally, some or all of the hardware replay architecture may operate in a pipelined manner to increase the throughput of the node. Advantageously, the use of caches and linked lists to implement the replay functionality allows a first node to communicate with a second node using TTP under limited hardware resources without the assistance of software control mechanisms.

[0099] As described above, a node operating under the TTP protocol may include a hardware link timer to implement a timeout checking mechanism for replaying packets without software assistance. In contrast to other Ethernet protocols (e.g., TCP or UDP), which typically employ software to track timeouts across multiple links using multiple timers (e.g., one timer for each link), a hardware link timer enables a node to determine which packets transmitted over which links to replay, and if and when replay is desired under limited hardware resources (e.g., when large resource pools of virtual and / or physical address space and computer resources are not available). In some embodiments, the hardware link timer may periodically perform timing checks on established links (e.g., active links) associated with the node. The hardware link timer may include a first-in, first-out (FIFO) memory that stores timing and status information associated with each active link and may check the timing and status associated with each active link in a round-robin manner. A hardware link timer may utilize a single programmable timer to schedule points in time for multiple active links and / or packets, and retrieve timing and status information associated with each of the multiple active links and / or packets, which may be used to determine whether to replay packets associated with the link or discard the packets through further information retrieval.

[0100] In some examples, the FIFO memory can store timing information associated with one or more links established between a first node and other nodes. For example, the first node may include a hardware link timer that uses the FIFO memory to store timing information associated with M links established between the first node and one or more other nodes, where M is a positive integer greater than 1. Instead of using M timers, each timer tracking timing information for a corresponding link, the hardware link timer may utilize a single timer (e.g., a timer that ticks once per programmable time period) to track and / or update timing information for each of the M links through accessing the FIFO memory in a round-robin (e.g., circular) manner. Specifically, the hardware link timer may access entries in the FIFO memory one at a time, with each accessed entry in the FIFO memory corresponding to one of the M links. In some embodiments, the time period of each tick may vary and may be on the order of hundreds of microseconds to single-digit microseconds. For example, the time period of a tick may be up to 100 microseconds or even down to 1 microsecond. Additionally, the hardware link timer may adjust the time period of a tick based on the number of links represented by the entries in the FIFO memory (e.g., M). For example, as M increases (e.g., more links represented by entries in the FIFO memory), the time period of a tick may decrease, and as M decreases (e.g., fewer links represented by entries in the FIFO memory), the time period of a tick may increase. Thus, the time interval at which link status and / or timing information is checked may remain constant while the time period of a tick varies inversely with the number of links represented by entries in the FIFO memory.

[0101] In some examples, timing and / or status information associated with one of the M links may indicate a time period during which the link has not received an acknowledgment for a transmitted received packet. Assuming a first node has transmitted N packets to a second node over the link, an entry in the FIFO memory may store timing and / or status information that, when accessed through a round-robin manner under a particular time period of ticks, indicates that no acknowledgment of receipt of any of the N packets has been received for a predetermined duration. Upon accessing the entry in the FIFO memory, a hardware link timer may utilize the timing and / or status information stored in the entry to retrieve the N packets that may be stored in a local storage device (e.g., a low-level cache) of the first node for replaying the N packets. Alternatively, timing and / or status information associated with one of the M links may be stored in an entry in the FIFO memory to indicate that the link may be closed (e.g., all packets transmitted by the first node have been received by the second node). Upon accessing the FIFO memory entry, the timing and / or status information stored in the FIFO memory entry indicates that the link may be closed, so the hardware link timer may utilize the timing and / or status information stored in the entry to search for packets that may still be stored in the first node's local storage and discard the packets. Advantageously, by utilizing a single timer for multiple links and / or packets that ticks under an adjustable period, and a FIFO memory that stores timing and / or status information for multiple links, the first node may replay packets at the appropriate times to achieve low latency and free up hardware resources occupied by inactive links (e.g., closed links) for use by active links to operate under limited computer and storage resources.

[0102] While various aspects are described according to example embodiments and feature combinations, those skilled in the art will understand that the examples and feature combinations are exemplary in nature and should not be construed as limiting. More specifically, aspects of the present application may be applicable to various types of networks and communication protocols in various contexts. Furthermore, while particular architectures of circuit block diagrams or state machines for controlling network flow are described, such example circuit block diagrams or state machines or architectures should not be construed as limiting. Accordingly, those skilled in the relevant art will understand that aspects of the present application are not necessarily limited in their application to any particular type of network, network infrastructure, or example interactions between nodes of a network.

[0103] Tesla Transport Protocol (TTP) 1A-1B are tables illustrating the OSI model (having seven layers) along with example protocols associated with each layer. Figure 1A illustrates example protocols including the TCP and UDP protocols operating on Layer 4 (e.g., the transport layer) of the OSI model. Figure 1B illustrates example protocols including the Tesla Transport Protocol (TTP) operating on Layer 4 of the OSI model.

[0104] 1A, in addition to TCP or UDP operating on Layer 4, other exemplary protocols or applications operating with TCP or UDP may include Hypertext Transfer Protocol (HTTP), Teletype Network Protocol (Telnet), File Transfer Protocol (FTP), which operate on Layer 7; Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Moving Picture Experts Group (MPEG), which operate on Layer 6; Network File System (NFS) and Structured Query Language (SQL), which operate on Layer 5; Internet Protocol version 4 (IPv4) / Internet Protocol version 6 (IPv6), which operate on Layer 3; etc. When TCP or UDP operates on Layer 4, the implementation of Layer 4 typically includes software as shown in FIG. 1A.

[0105] In addition to TTP operating at Layer 4 as shown in FIG. 1B, other exemplary protocols or applications operating with TTP may include Pytorch operating at Layer 7, FFMPEG, High Efficiency Video Coding (HEVC), YUV operating at Layer 6, RDMA operating at Layer 5, IPv4 / IPv6 operating at Layer 3, etc. In contrast to FIG. 1A, when TTP operates at Layer 4, implementations of Layers 1-4 of the OSI model can be performed solely in hardware without software involvement as shown in FIG. 1B. Advantageously, a pure hardware implementation of Layers 1-4 of the OSI model based on TTP as shown in FIG. 1B can reduce the latency of communications over Ethernet-based networks compared to the implementation as shown in FIG. 1A.

[0106] FIG. 2 illustrates an example state machine 200 for opening and closing links between nodes implementing TTP, according to an embodiment of the present disclosure. The state machine 200 may be implemented by a network interface processor or a network interface card. There may be one state machine 200 for each Ethernet link between nodes on each node that communicates over the Ethernet link. For example, if a network interface processor may communicate with five network interface cards over five TTP links, the network interface processor may include five instances of the state machine 200, one instance per link. In this example, each of the five network interface cards may have one instance of the state machine 200 for communicating with the network interface processor. In some embodiments, nodes that communicate with each other using the state machine 200 may form a peer-to-peer network.

[0107] 2, state machine 200 includes a closed state 202, an open receive state 204, an open transmit state 206, an open state 208, a closed receive state 210, and a closed transmit state 212. State machine 200 may begin in closed state 202, which may indicate that a communication link is not currently open between a first node maintaining state machine 200 and a second node with which a communication link is to be established. Furthermore, separate copies of state machine 200 may be maintained, updated, and transitioned through by nodes operating in accordance with the Tesla Transport Protocol (TTP) disclosed in this disclosure. Furthermore, if a node operating in accordance with the TTP communicates with multiple nodes simultaneously or overlapping in time, the node may maintain multiple independent state machines 200 for each link.

[0108] State machine 200 may then transition differently depending on whether the first node sends or receives a request to establish a communication link to the second node. If the first node sends a request to open a communication link to the second node, state machine 200 may transition from closed state 202 to open send state 206. On the other hand, if the first node receives a request to open a communication link from the second node, state machine 200 may transition from closed state 202 to open receive state 204.

[0109] While in the open-send state 206, the state machine 200 may remain in the open-send state 206 or transition back to the closed state 202 or to the open state 208 depending on various criteria. If the first node receives an open negative response (e.g., a message denying the request to open the link) from the second node, the state machine 200 may transition from the open-send state 206 back to the closed state 202. On the other hand, if the first node receives an open positive response (e.g., a message accepting the request to open the link) from the second node, the state machine 200 may transition from the open-send state 206 to the open state 208. Alternatively, if the first node does not receive an open negative response or an open positive response from the second node within a certain period of time, the first node may time out, after which the first node can retransmit a request to open a communication link to the second node and remain in the open-send state 206.

[0110] As described above, if, while in the closed state 202, the first node receives a request from a second node to open a communication link, the state machine 200 may transition from the closed state 202 to the open-receive state 204. In the open-receive state 204, the state machine 200 may transition variously depending on whether the first node accepts or rejects the request to open the link from the second node. For example, the first node may choose to send an open negative response (e.g., rejecting the request to open the link) to the second node. In such a situation, the state machine 200 may transition back to the closed state 202, and the first node may send or receive further requests to open the link from the second node or other nodes. Alternatively, in the open-receive state 204, the first node may send an open positive response to the second node and then transition to the open state 208.

[0111] While in the open state 208, the first node and the second node may send and receive packets to each other over the established communication link. This link may be a wired Ethernet link. The first node may remain in the open state 208 until some condition occurs. In some embodiments, while in the open state 208, the state machine 200 may transition from the open state 208 to the closed receive state 210 in response to receiving a request to close the communication link, which may cause the first node and the second node to send and receive packets. Alternatively, the state machine 200 may transition from the open state 208 to the closed transmit state 212 in response to the first node sending a request to the second node to close the communication link. In addition to a request to close the communication link, the state machine 200 may transition from the open state 208 to the closed receive state 210 or the closed transmit state 212 if the communication link has been idle for more than a threshold amount of time.

[0112] While in the close receive state 210, the state machine 200 may transition back to the close state 202 if the first node sends a close acknowledgement (e.g., a message acknowledging or accepting the request to close the link) to the second node. Otherwise, the state machine 200 may remain in the close receive state 210 if the first node sends a close negative acknowledgement (e.g., a message denying or not acknowledging the request to close the link) to the second node.

[0113] While in the send close state 212, state machine 200 may transition back to the close state 202 if the first node receives a close acknowledgement (e.g., a message acknowledging or accepting the request to close the link) from the second node. Otherwise, state machine 200 may remain in the send close state 212 if the first node receives a negative close acknowledgement (e.g., a message denying or not acknowledging the request to close the link) sent from the second node. In the send close state 212, if the first node does not receive a reply from the second node within a timeout threshold, the first node may resend the request to close the communication link to the second node.

[0114] In some embodiments, state machine 200 may be maintained and implemented by hardware without the involvement of software, firmware, drivers, or other types of programmable instructions. Thus, transitions between the various states of state machine 200 may be accelerated compared to implementations of other protocols that include software support, such as Transmission Control Protocol (TCP) applicable to Ethernet-based networks.

[0115] In some embodiments, instead of continuing to transmit packets waiting to be sent and stored in a transmit queue, the first node may immediately stop transmitting packets in the transmit queue while transmitting a close acknowledgment to the second node in response to receiving a request to close the link from the second node while in the close receive state 210. Advantageously, refraining from continuing to transmit packets for an indefinite amount of time after receiving a request to close the link allows the first node to transition from the open state 208 back to the closed state 202 with less transition period and less time uncertainty.

[0116] Additionally, the number of packets that can be transmitted consecutively by the first node or the second node while in the open state 208 may be limited. For example, while in the open state 208, the first node may transmit only N packets consecutively before stopping packet transmission, where N may be a positive integer from 1 to over 1000. The number N may be limited by physical memory. In some embodiments, N may be limited or constrained by the size of physical memory (e.g., dynamic random access memory, etc.) available to the first node. Specifically, N may be proportional to the size of physical memory associated with the first node or the second node. For example, if 1 gigabyte (GB) of physical memory is allocated to the first node, N may be up to 1 million. In some embodiments, N may be in the range of tens of thousands or hundreds of thousands. The amount of physical memory for exchanging packets during the open state 208 may be tracked. Advantageously, limiting the number of packets that can be transmitted consecutively by the first node or the second node may reduce computational and storage resources for implementing the state machine 200. In contrast to protocols (e.g., TCP) that typically presume unlimited software and hardware resource availability through virtualization (e.g., virtualized memory or processing resources), limiting the number of packets sent allows TTP to operate under constrained computational and storage resources.

[0117] In some embodiments, after receiving or transmitting a close acknowledgment to the other, the first node or the second node does not further wait to close the link. For example, while in the close-send state 212, the first node may immediately transition to the close state 202 in response to receiving a close acknowledgment sent from the second node. Instead of waiting another predetermined or random period of time to monitor whether the second node has additional packets to send, the first node may transition from the close-send state 212 back to the close state 202 in a short amount of time. Advantageously, this increases the accuracy and reduces the latency associated with transitioning between states in the state machine 200, thereby allowing TTP to readily communicate at lower latencies than protocols such as TCP.

[0118] 3A-3B illustrate exemplary timing diagrams showing the transmission and reception of packets between two devices implementing TTP, according to an embodiment of the present disclosure. FIG. 3A illustrates a scenario in which none of the packets transmitted from device A to device B are lost, while FIG. 3B illustrates another scenario in which some of the packets transmitted from device A to device B are lost. FIG. 3A-3B can be understood in conjunction with state machine 200. Device A and device B are two exemplary nodes communicating via TTP.

[0119] 3A, while in the closed state 202, device A may send a TTP_OPEN with packet ID=0 to device B. After sending the TTP_OPEN to device B at (1), the state machine maintained by device A may transition from the closed state 202 to an open send state 206. Furthermore, after receiving the TTP_OPEN from device A at (1), the state machine maintained by device B may transition from the closed state 202 to an open receive state 204.

[0120] Then, at (2), after receiving the TTP_OPEN_ACK from device B, the state machine maintained by device A may transition from the open-send state 206 to the open state 208. Further, at (2), after sending the TTP_OPEN_ACK to device A, the state machine maintained by device B may transition from the open-receive state 204 to the open state 208.

[0121] At (3), while in the open state 208, device A may send four packets (e.g., TTP_PAYLOAD ID=1-4) consecutively or consecutively to device B before receiving any response from device B. In some embodiments, the number of packets that device A may send to device B before receiving any response from device B is limited. In response to the packets received from device A, device B may send four packets (e.g., TTP_ACK ID=1-4) at (4) acknowledging receipt of the four packets sent by device A.

[0122] At (5), device A sends a TTP_CLOSE (with packet ID=5) to device B. After sending the TTP_CLOSE, the state machine maintained by device A may transition from the open state 208 to the close send state 212. In response to receiving the TTP_CLOSE from device A, the state machine maintained by device B may transition from the open state 208 to the close receive state 210.

[0123] Thereafter, at (6), device B may send a TTP_CLOSE_ACK (with packet ID=5) to device A. After sending the TTP_CLOSE_ACK to device A, the state machine maintained by device B may transition from the close receive state 210 back to the close state 202. After receiving the TTP_CLOSE_ACK from device B, the state machine maintained by device A may transition from the close send state 212 back to the close state 202. In this manner, the link / connection between device A and device B may be close.

[0124] FIG. 3B illustrates a “lossy” flow control feature associated with a flow control protocol (e.g., TTP) disclosed in the present disclosure, where lossy may indicate that lost or corrupted packets are retransmitted after receiving a negative acknowledgment.

[0125] 3B, while in the closed state 202, device A may send a TTP_OPEN with packet ID=0 to device B. After sending the TTP_OPEN to device B at (1), the state machine maintained by device A may transition from the closed state 202 to an open send state 206. Furthermore, after receiving the TTP_OPEN from device A at (1), the state machine maintained by device B may transition from the closed state 202 to an open receive state 204.

[0126] Then, at (2), after receiving the TTP_OPEN_ACK from device B, the state machine maintained by device A may transition from the open-send state 206 to the open state 208. Further, at (2), after sending the TTP_OPEN_ACK to device A, the state machine maintained by device B may transition from the open-receive state 204 to the open state 208.

[0127] At (3), while in the open state 208, device A may send four packets (e.g., TTP_PAYLOAD ID=1-4) consecutively or consecutively to device B before receiving any response from device B. However, due to some network conditions, device B may not receive some of the packets (e.g., TTP_PAYLOAD ID=3). Thus, at (4), device B may send three packets (e.g., TTP_ACK ID=1-2 and TTP_NACK ID=3) that acknowledge receipt of the two packets (ID=1-2) sent by device A but notify that the packet with TTP_PAYLOAD ID=3 was not received.

[0128] After receiving a packet (e.g., TTP_NACK ID=3) from device B, device A retransmits two packets (e.g., TTP_PAYLOAD ID=3-4) to device B at (5). Notably, the retransmission of two packets after receiving a packet (e.g., TTP_NACK ID=3) reflects the "lossy" feature of TTP. In some embodiments, device A may retransmit some of the packets after a timeout occurs (e.g., if a local counter exceeds a certain value). Advantageously, the "lossy" feature allows TTP to control or scale network flow without limitations due to the existence of a peer-to-peer link between device A and device B, and allows TTP to achieve link-specific recovery in large systems where some traffic loss is expected.

[0129] In (6), after receiving the two packets (e.g., TTP_PAYLOAD ID=3-4), device B may send two packets (e.g., TTP_ACK ID=3-4) to device A to acknowledge receipt of the retransmitted packets (e.g., TTP_PAYLOAD ID=3-4).

[0130] At (7), device A may send a packet (e.g., TTP_CLOSE ID=5) to device B in an attempt to close the link between device A and device B. Further, at (7), the state machine maintained by device A may transition from the open state 208 to the closed transmit state 212, and the state machine maintained by device B may transition from the open state 208 to the closed receive state 210.

[0131] At (8), device B may send a packet (e.g., TTP_CLOSE_ACK ID=5) to device A to acknowledge and agree to close the link. The state machine maintained by device B may transition from the close receive state 210 back to the close state 202. In response to receiving the packet (e.g., TTP_CLOSE_ACK ID=5) from device B, the state machine maintained by device A may transition from the close send state 212 back to the close state 202.

[0132] In some embodiments, device A and / or device B may not transition to the open state 208 or send or receive data packets until the process of negotiating the link is complete. For example, device A may not send data packets to or accept data packets from device B until device A receives a TTP_OPEN_ACK from device B. In these embodiments, there may be no need to impose a timeout period when closing a link between device A and device B, especially when a TTP_OPEN is sent from device A or device B immediately after the previous link between device A and device B was closed.

[0133] FIG. 4 illustrates an exemplary block diagram of a node 400 implementing TTP, according to an embodiment of the present disclosure. As shown in FIG. 4, the node 400 may include a transmit (TX) path and a receive (RX) path. As shown in FIG. 4, the node 400 includes, at its front end, a physical coding sublayer (PCS)+physical medium attachment (PMA) block 402 that handles communications over Layer 1 (e.g., the physical layer) of the OSI model. In some embodiments, the PCS+PMA block 402 operates based on a reference clock 404 having a frequency of 156.25 MHz. In other embodiments, the PCS+PMA block 402 may operate under a different clock frequency. The PCS+PMA block 402 may be compatible with Ethernet or the IEEE 802.3 standard. In operation for processing data on the RX path, the PCS+PMA block 402 receives RX serdes[3:0] as an input and rearranges the RX serdes[3:0] into an output (e.g., RX frame 408) that is processed by the TTP medium access control (MAC) block 410. In operation for processing data on the TX path, the PCS+PMA block 402 receives TX frame 412 from the TTP MAC block 410 as an input and rearranges the data format to output TX serdes[3:0].

[0134] On the RX path, the TTP MAC block 410 receives an RX frame 408 as an input and outputs RDMA receive data 416 to a system-on-chip (SoC) 420. On the TX path, the TTP MAC block 410 receives RDMA transmit data 418 from the SoC 420 and outputs a TX frame 412 to the PCS+PMA block 402. As shown in FIG. 4, the TTP MAC block 410 may process operations on layers 2 through 4 of the OSI model. The TTP MAC block 410 may include a TTP finite state machine (FSM) 422. The TTP FSM 422 may maintain and update the state machine 200 as shown in FIG. 2. As described above, for each communication link that the node 400 has established with one or more other nodes, the TTP FSM 422 may maintain and update a corresponding state machine (e.g., the state machine 200) to control the flow associated with each communication link.

[0135] In some embodiments, PCS+PMA block 402 and TTP MAC block 410 may be implemented in hardware, such as in the form of an application specific integrated circuit (ASIC) or field programmable gate array (FPGA). Thus, PCS+PMA block 402 and TTP MAC block 410 may operate without the assistance or involvement of software / firmware / drivers. Advantageously, PCS+PMA block 402 and TTP MAC block 410 may handle communications up to Layers 1-4 of the OSI model without software assistance to reduce latency associated with communications up to Layers 1-4.

[0136] FIG. 5 illustrates an exemplary header 500 for a packet transmitted or received in accordance with TTP. As shown in FIG. 5, the exemplary header 500 has 64 bytes. The first 16 bytes include headers for Ethernet Layer 2 (e.g., data link layer) and virtual local area network (VLAN) operation. The second 16 bytes include ETHTYPE followed by an optional Layer 3 Internet Protocol (IP) header. To support Layer 2 operation under TTP, ETHTYPE can be set to a specific value (e.g., 0x9AC6). When ETHTYPE is set to a specific value, the header 500 may signal to a network device processing the header 500 that the header 500 is formatted in accordance with TTP. The third 16 bytes include optional fields for Layer 3 (IP) operation and Layer 4 operation under UDP. The third 16 bytes and the fourth 16 bytes end with fields for Layer 4 operation under TTP. TTP may be referred to as TTP over Ethernet (TTPoE). TTP is labeled as TTPoE in Figure 5.

[0137] Advantageously, the exemplary header 500 enables the TTP to support operation over Ethernet-based networks from at least Layers 2 through 4 of the OSI model. Specifically, existing Ethernet switches and hardware may support the operations associated with the TTP.

[0138] FIG. 6 illustrates an exemplary network and computer environment 600 in which embodiments of the present disclosure may be implemented. The exemplary network and computer environment 600 may be utilized in a high-performance computing or artificial intelligence training data center. As an example, the network and computer environment 600 may be used in neural network training to generate data for use by an autonomous driving system for a vehicle (e.g., an automobile). As shown in FIG. 6, the exemplary network and computer environment 600 includes an Ethernet switch 608, hosts 602A-602E, Peripheral Component Interconnect Express (PCIe) hosts 604A-604N, and computing tiles 606A-606N. While FIG. 6 shows five hosts 602A-602E, any suitable number of hosts greater than or less than five may be implemented. Furthermore, the number of PCIe hosts and the number of computing tiles may be any suitable positive integer.

[0139] Each of the hosts 602A-602E includes a network interface card (NIC), a central processing unit (CPU), and dynamic random access memory (DRAM). While shown as a CPU, in some embodiments the CPU may be embodied as any type of single-core, single-threaded, multi-core, or multi-threaded processor, microprocessor, digital signal processor (DSP), microcontroller, or other processor or processing / control circuitry. Although shown as DRAM, in some embodiments the DRAM may alternatively or additionally be embodied as any type of volatile or non-volatile memory or data storage device, such as static random access memory (SRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The DRAM may store various data and program code used during operation of the hosts 602A-602E, including operating systems, application programs, libraries, drivers, etc.

[0140] In some embodiments, the NIC may implement TTP for communicating with the Ethernet switch 608. Each NIC may communicate with the Ethernet switch 608 using TTP as a flow control protocol to manage links established between each NIC and a network interface processor (NIP) through the Ethernet switch 608. In some embodiments, the NIC may include the PCS+PMA block 402 and the TTP MAC block 410 of FIG. 4. In some embodiments, the NIC may implement TTP without software / firmware assistance.

[0141] As shown in FIG. 6, each of the PCIe hosts 604A-604N may include a network interface processor (NIP) and high-bandwidth memory (HBM). In some embodiments, the bandwidth supported by the HBM may be 32 gigabytes (GB) per compute. Each of the PCIe hosts 604A-604N may communicate with each of the computing tiles 606A-606N. Each of the computing tiles 606A-606N may include storage, input / output, and computational resources. The computing tile 606A may include a system-on-wafer having an array of processors for high-performance computing. In a particular application, each of the computing tiles 606A-606N may perform 9 peta-floating-point operations per second (PFLOPS), store data 11 gigabytes (GB) in size using static random access memory (SRAM), or facilitate input / output operations at a bandwidth of 36 terabytes (TB) per second.

[0142] In some embodiments, each of the NICs in the hosts 602A-602E may open and close a communication link with each of the NIPs in the PCIe hosts 604A-604N. Specifically, by implementing the state machine 200 of FIG. 2, one NIC and one NIP may open and close a communication link with each other. To open and close a communication link, the NIC and the NIP may use packets including the opcodes of FIGS. 7A-7B to perform the desired operation. For example, to open a link with a NIP, the NIC may send a packet including the opcode TTP_OPEN (shown in FIG. 7A) to the NIP requesting to open a communication link. After receiving the packet with the opcode TTP_OPEN, the NIP may transition from the closed state 202 to the open-receive state 204 of FIG. 2. After sending a packet with the opcode TTP_OPEN_ACK (shown in FIG. 7A), the NIP may transition from the open-receive state 204 to the open state 208, as shown in FIG. 2. In some embodiments, once a communication link is established (e.g., when the NIC and NIP are both in the open state 208), the NIC and NIP may send or receive packets from each other using the header 500 of Figure 5. In other words, each packet sent or received between the NIC and the NIP may include the header 500 of Figure 5.

[0143] As shown in FIG. 6 , communication and data exchange between each of the hosts 602A-602E, each of the PCIe hosts 604A-604N, each of the computing tiles 606A-606N, or the Ethernet switch 608 may be based on TTP. The low latency achieved through TTP (compared to TCP) using the techniques described above may enable high bandwidth and high-speed communication between the various elements of FIG. 6 . In some embodiments, at least a portion of the NIP or at least a portion of the NIC shown in FIG. 6 may be implemented similarly or identically to the node 400 of FIG. 4 . While not fully shown in FIG. 6 , in some embodiments, each of the NICs and NIPs may include a port 610 that may receive and transmit packets. In some embodiments, the port 610 is an Ethernet port.

[0144] Figures 7A-7B illustrate opcodes for various types of TTP packets according to an embodiment of the present disclosure. The TTP packets illustrated in Figures 7A and 7B are utilized in Figures 2, 3A, and 3B to open and close links between nodes in the network. TTP packets can be exchanged between nodes in the network and computer environment of Figure 6. The TTP packets illustrated in Figures 7A and 7B can be best understood in conjunction with Figures 2, 3A, and 3B.

[0145] Packet replay hardware architecture Referring again to FIG. 4 , which illustrates an exemplary block diagram of a node 400 that transmits and / or receives packets using TTP, a replay hardware architecture will now be described. As described above, the node 400 may include blocks such as a physical coding sublayer (PCS)+physical medium attachment (PMA) block 402 and a TTP medium access control (MAC) block 410 that includes a TTP FSM 422 for processing communications from Layer 1 to Layer 4 of the OSI model without software assistance, thereby reducing the latency associated with communications from Layer 1 to Layer 4. Furthermore, the TTP medium access control (MAC) block 410 of the node 400 may include a hardware replay architecture that includes at least a TTP (peer link) tag block 436, an RX data path 432, an RX storage 432-1 (e.g., on die SRAM), a TX data path 434, and a TX storage 434-1 (e.g., on die SRAM). The hardware replay architecture can replay packets lost during transmission under a lossy protocol such as TTP. Optionally, the TTP medium access control (MAC) block 410 of the node 400 may further include a TTP MAC RDMA address encoding block 438 that may receive and encode RDMA transmit data 418 from a system on chip (SoC) 420 .

[0146] In some embodiments, the hardware replay architecture of node 400 for replaying packets may include at least the following circuitry: TTP tag block 436, RX data path 432, RX storage 432-1, TX storage 434-1, and TX data path 434. As described above, the hardware replay architecture may utilize physical storage and data structures to store packets transmitted and / or received on various links, particularly to maintain the order of transmitted packets when replay occurs. In some embodiments, the physical storage utilized by the hardware replay architecture may be any suitable type of local storage or cache (e.g., a low-level cache) that may store, buffer, and / or hold packets associated with one or more links. The physical storage may be limited in size, such as having a size in units of megabytes (MB) or kilobytes (KB). In some examples, the physical storage may be deployed as part of TX data path 434, or more specifically, as part of TX storage 434-1. Physical storage may also be deployed as part of RX data path 432, and more specifically, as part of RX storage 432-1. For example, physical storage may be RX storage 432-1 and TX storage 434-1, and the RX storage 432-1 and TX storage 434-1 utilized by the hardware replay architecture associated with each of RX data path 432 and TX data path 434 may be 256 KB in size. In other examples, physical storage may be deployed within and as part of TTP tag block 436 (e.g., as local storage deployed within TTP tag block 436).

[0147] Note that any other suitable size of physical storage may be employed by the hardware replay architecture within the TTP medium access control (MAC) block 410 of node 400. In some embodiments, the data structure utilized by the hardware replay architecture (e.g., within the TTP tag block 436) may include one or more linked lists, each of which may record and / or track the order of packets transmitted for a corresponding link established between a first communication node and a second communication node. In some embodiments, the TTP tag block 436 may utilize linked lists in conjunction with physical storage (e.g., the RX storage 432-1 and the TX storage 434-1) to maintain and manage stored packets and replay packets transmitted over multiple links.

[0148] 8 and 9 illustrate example physical storage and data structures (e.g., TX linked list 952) utilized by a node (e.g., node 400 or device A of FIG. 3B) within an Ethernet-based network implementing TTP for replaying or retransmitting packets, according to some embodiments of the present disclosure. Figures 8 and 9 can be understood in conjunction with reference to FIG. 3B, which illustrates that in response to receiving a negative acknowledgment packet (e.g., TTP_NACK ID=3) indicating that a packet (TTP_PAYLOAD ID=3) was not received, device A replays two packets (e.g., TTP_PAYLOAD ID=3-4).

[0149] 8, device A of FIG. 3B may store packet 1 (e.g., packet TTP_PAYLOAD ID=1 of FIG. 3B), packet 2 (e.g., packet TTP_PAYLOAD ID=2 of FIG. 3B), packet 3 (e.g., packet TTP_PAYLOAD ID=3 of FIG. 3B), packet 4 (e.g., packet TTP_PAYLOAD ID=4 of FIG. 3B), and packet 5 (e.g., packet TTP_CLOSE ID=5 of FIG. 3B) in a physical storage device, such as packet physical cache 802, for transmission and / or replay. As described above, packet physical cache 802 may be TX storage device 434-1 and / or physical storage device deployed within TTP tag block 436. In some embodiments, packet physical cache 802 may have two storage spaces: packet physical tag 804 and packet physical data 806. For each packet (e.g., packets 1 to 5), the packet physical tag 804 may include a physical address pointer that points to a physical address in the packet physical data that stores the packet. For example, a physical address pointer 808 associated with packet 4 stored in an entry in the packet physical tag 804 may point to an entry in the packet physical data 806 in which packet 4 (e.g., packet TTP_PAYLOAD ID=4 in FIG. 3B) is stored. As illustrated in FIG. 8, device A may transmit packets 1, 2, 3, 4, and 5 in an order 820 (e.g., transmitting packet 1 first and packet 5 last). However, device A may not store packets 1 to 5 in the packet physical data 806 based on the order 820. Specifically, device A transmits packet 3 before packets 4 and 5, but the address 810 in the packet physical data 806 that stores packet 3 may follow addresses 812 and 814 in the packet physical data 806 that store packets 4 and 5, respectively.

[0150] FIG. 9 illustrates a TX linked list 952 that may be utilized by node 400 and / or device A of FIG. 3B to maintain the order of packet transmission between previous transmissions and replays. TX linked list 952 may be part of TTP tag block 436 of node 400. As discussed above in the description of FIG. 8, device A of FIG. 3B may store packets 1 through 5 at various addresses in packet physical data 806 that do not reflect the order 820 in which packets 1 through 5 are transmitted. Nevertheless, device A may utilize TX linked list 952 to keep track of and maintain the desired order in which packets 1 through 5 are transmitted. As shown in FIG. 9, TX linked list 952 includes five elements 960, 962, 964, 968, and 970, each of which corresponds to or is associated with one of packets 1 through 5. FIG. 9 illustrates that TX linked list 952 tracks and maintains the order 820 in which packets 1 through 5 are transmitted. For example, in TX linked list 952, element 964 corresponding to packet 3 comes before and points to element 968 corresponding to packet 4, which comes before and points to element 970 corresponding to packet 5. Thus, by utilizing TX linked list 952, device A may maintain the order of packet transmissions during previous transmissions and replays, and replay may be triggered in response to receiving a TTP_NACK ID=3 packet indicating that a packet having TTP_PAYLOAD ID=3 was not received by device B of FIG. 3B. Replay may be triggered in response to a timeout or negative acknowledgment, according to any suitable principles and advantages disclosed herein.

[0151] As shown in FIG. 9, device A of FIG. 3B may further use one or more pointers 972, 974, and 976 stored in memory to determine which packets to replay. As shown in (3) and (4) of FIG. 3B, device A transmits four packets (e.g., TTP_PAYLOAD IDs=1-4) and receives three packets (e.g., TTP_ACK IDs=1-2 and TTP_NACK ID=3) that acknowledge receipt of the two packets transmitted by device A (IDs=1-2), but indicate that the packet with TTP_PAYLOAD ID=3 was not received. In response, device A may set pointer 972 to point to element 964 corresponding to packet 3 to indicate that device A will replay packets starting with packet 3. Device A may further set pointer 974 to point to element 968 corresponding to packet 4 to indicate that device A will replay packet 4 in addition to packet 5. To indicate that device A may replay packets 3 and 4 before sending packet 5, the device may further set pointer 976 to point to element 970 corresponding to packet 5. Furthermore, device A may set elements 960 and 962 of TX linked list 952 to null to indicate that device A may remove packets 1 and 2 from packet physical data 806 and packet physical tag 804 addresses (not shown in FIG. 8 ) to free up more storage space for device A to store transmitted or received packets.

[0152] Thereafter, based on TX linked list 952, pointer 972, and pointer 974, device A may replay packets 3 and 4, as shown at (5) in FIG. 3B. Next, device A may receive an acknowledgment of receipt of packets 3 and 4, as shown at (6) in FIG. 3B. In response, based on TX linked list 952, device A may transmit packet 5 (e.g., packet TTP_CLOSE ID=5 in FIG. 3B) corresponding to element 970 of TX linked list 952 to complete the transmission and replay of packet 1 through packet 5. Additionally and / or optionally, device A may free the storage occupied by packets 1 through 5 after all packets corresponding to elements of TX linked list 952 have been transmitted and replayed. In some embodiments, device A may indicate that an address in packet physical tag 804 and an address in packet physical data 806 are released and freed for use in conjunction with other linked lists corresponding to other packets by setting free list entry 832 and free list entry 834, respectively, to particular values.

[0153] 10 illustrates an example block diagram of the TTP tag block 436 of FIG. 4 , which is part of a hardware replay architecture for replaying packets transmitted over multiple links, in accordance with some embodiments of the present disclosure. As shown in FIG. 10 , the TTP tag block 436 may include a memory that stores a TX linked list 1020 and logic circuits 1012, 1014, 1016, and 1018, which operate in pipeline stages 1002, 1004, 1006, and 1008, respectively. The logic circuits 1012, 1014, 1016, and 1018 may be implemented by any suitable physical circuitry. In some examples, some or all of the logic circuits 1012, 1014, 1016, and 1018 may be implemented by dedicated circuitry, such as in the form of an application-specific integrated circuit (ASIC). In some examples, some or all of logic circuits 1012, 1014, 1016, and 1018 may be implemented with programmable logic gates or general-purpose processing circuitry, such as in the form of a field programmable gate array (FPGA) or a digital signal processor (DSP). In operation, TX linked list 1020 may function similarly to TX linked list 952 of FIG. 9. In some embodiments, TX linked list 1020 tracks the order of N packets, including packet 1022, packet 1024, and packet 1026, and node 400 may transmit the N packets tracked by TX linked list 1020 over a particular link. TTP tag block 436 further includes pointer 1032, pointer 1034, and pointer 1036, which point to packet 1022, packet 1024, and packet 1026, respectively. TTP tag block 436 may store pointer 1032, pointer 1034, and pointer 1036 in any suitable storage element (not shown in FIG. 10). In particular applications, the N packets including packet 1022, packet 1024, and packet 1026 in TX linked list 1020 may be stored in a physical storage device such as TX storage device 434-1 in TX data path 434 of node 400.In such an application, TX linked list 1020 may include pointers to packets 1022, 1024, and 1026. In other applications, N packets including packet 1022, packet 1024, and packet 1026 may be part of TX linked list 1020 stored in physical storage within TTP tag block 436.

[0154] In some embodiments, node 400 may store in TX storage 434-1 (or other physical storage of node 400) N packets (including packet 1022, packet 1024, and packet 1026) transmitted to the second node using the link established under the TTP, where N is any positive integer that may be limited by the size of TX storage 434-1. Node 400 may continue to transmit some or all of the N packets to the second node as constraints from the TTP and / or network conditions permit. To accommodate replay of the N packets, including packet 1022, packet 1024, and packet 1026, TX storage 434-1 may continue to store one or more previously transmitted packets (e.g., packet 1022) until an acknowledgment that one or more packets have been received is received from the second node. Packets may be stored until receipt of a previously transmitted packet is acknowledged. When an acknowledgment of receipt of a packet is received, TX storage 434-1 may discard the packet to make space for storing packets to be transmitted over the link between node 400 and the second node and / or one or more other nodes, or over other links. In contrast, if a negative acknowledgment of the packet is received (e.g., the second node notifies node 400 that the packet was not received), or if a timeout occurs without receiving an acknowledgment or negative acknowledgment of receipt of the packet from the second node, node 400 may replay (e.g., retransmit) the packet still stored in TX storage 434-1. In conjunction with replaying a packet, node 400 may discard other packets for which an acknowledgment of receipt was received. In some embodiments, TX linked list 1020 may cooperate with TX storage 434-1 to maintain ordering between previous transmissions and any subsequent replays of some or all of the N packets, including packets 1022, 1024, and 1026.As shown in FIG. 10, the TX linked list 1020 includes N elements, each corresponding to one of the N packets, or each of the N packets and a reference to the next element corresponding to the next packet.

[0155] When transmitting and / or replaying N packets, to determine whether the packets should be kept for replay or discarded by TX storage 434-1 to save storage resources, TTP tag block 436 may further utilize pointer 1032, pointer 1034, and pointer 1036, which respectively point to three elements in TX linked list 1020. Taking the example where N is 9 (e.g., 9 packets transmitted from node 400 to a second node), in TX linked list 1020, the first element corresponds to the first packet (e.g., packet 1022) and the first reference, the first reference points to the second element, the second element corresponds to the second packet and the second reference, the second reference points to the third element, the eighth element corresponds to the eighth packet (e.g., packet 1024) and the eighth reference, the eighth reference points to the ninth element, and the ninth element corresponds to the ninth packet (e.g., packet 1026). The TTP tag block 436 may maintain and update three pointers 1032, 1034, and 1036, which point to the first element (e.g., packet 1022), the eighth element (e.g., packet 1024), and the ninth element (e.g., packet 1026), respectively.

[0156] Further, assuming node 400 receives an acknowledgment from a second node that transmits packets 1 through 9 and receives packets 1 through 7, but not packets 8 and 9, then pointer 1032 points to the first element (e.g., packet 1022) of TX linked list 1020, then pointer 1034 points to the eighth element (e.g., packet 1024) of TX linked list 1020, and then pointer 1036 points to the ninth element (e.g., packet 1026) of TX linked list 1020. Thus, TTP tag block 436 allows TX storage device 434-1 to discard some or all of the N packets, including packets 1022, 1024, and 1026, and to replay some or all of the N packets based on pointers 1032, 1034, and 1036. More specifically, TX storage 434-1 may replay packet 1024, which is pointed to by pointer 1034, through packet 1026, which is pointed to by pointer 1036 (in this case, only packets 1024 and 1026 are replayed). TX storage 434-1 may further discard the remaining packets (e.g., it may discard packet 1022, which is pointed to by pointer 1032, and other packets previously transmitted before packet 1024, in this case, seven packets including packet 1022).

[0157] 10, some or all of TTP tag block 436 (e.g., logic circuits 1012, 1014, 1016, and 1018) may operate in a pipelined manner to increase the throughput of node 400. Logic circuits 1012, 1014, 1016, and 1018 may operate in conjunction with TX linked list 1020 to determine whether a packet should be replayed or discarded / retired from TX storage 434-1 or other physical storage of node 400 that stores the packet. As shown in FIG. 10, logic circuits 1012, 1014, 1016, and 1018 may operate in their respective pipeline stages according to the clock that TTP tag block 436 operates on. Specifically, logic circuit 1012 operates in the initial pipeline stage 1002 (labeled "Q0"), logic circuit 1014 operates in the first pipeline stage 1004 (labeled "Q1"), logic circuit 1016 operates in the second pipeline stage 1006 (labeled "Q2"), and logic circuit 1018 operates in the third pipeline stage 1008 (labeled "Q3").

[0158] During operation, logic circuitry 1012 may select one of the data streams for processing in the TTP link tag pipeline. As shown in initial pipeline stage 1002, logic circuitry 1012 may select one of the transmit stream ("TX QUEUE"), receive stream ("RX QUEUE"), or acknowledgment stream ("ACK QUEUE") for processing in the TTP link tag pipeline based on a control signal (e.g., "PICK"). In the TTP link tag pipeline, logic circuitry determines whether to replay one or more packets of the selected data stream or to retire one or more packets of the selected data stream. The TTP link tag pipeline may also determine to reject acknowledgments of packets sent after another packet that the TTP tag pipeline has determined to replay.

[0159] Assuming that logic circuit 1012 selects a transmit stream to prepare for packet replay, in the first pipeline stage 1004, logic circuit 1014 determines which link to evaluate for replay. This may include reading a tag associated with the link. As shown in FIG. 10 , logic circuit 1014 may potentially select one of two links (e.g., “MOOSE” and “CAT”) to replay, where each link may be established between the same endpoint or different endpoints. For example, both links “MOOSE” and “CAT” may be established between node 400 and a second node, or link “MOOSE” may be established between node 400 and a second node, while link “CAT” may be established between node 400 and a third node. Logic circuit 1014 may select a link (e.g., “CAT”) to replay based on a link pointer that points to the selected link.

[0160] Next, in second pipeline stage 1006, logic circuit 1016 may determine which packets transmitted over link “CAT” to replay or retire. In some embodiments, logic circuit 1016 may determine to replay some of the packets transmitted over link “CAT,” while retiring other packets based on whether an acknowledgment or negative acknowledgement of receipt is received. For example, logic circuit 1016 may determine to replay packet 1024 if a negative acknowledgement of packet 1024 is received or no acknowledgment of packet 1024 is received for a period of time that triggers a timeout. In contrast, logic circuit 1016 may determine to retire packet 1022 in response to receiving an acknowledgment of packet 1022. Additionally and / or optionally, logic circuit 1016 may further determine to replay and / or retire other packets transmitted over link “CAT” based on TX link list 1020. For example, based on the order of packets transmitted over link “CAT” specified by TX link list 1020, indicating that packet 1026 was transmitted after packet 1024, logic circuit 1016 may determine to replay packet 1026 in addition to replaying packet 1024 in response to receiving a negative acknowledgment for packet 1024. Assuming that an acknowledgment for a packet transmitted between packet 1022 and packet 1024 has been received, logic circuit 1016 may cause TX storage device 434-1 to further retire packets transmitted between packet 1022 and packet 1024 to conserve available storage space within TX storage device 434-1. In second pipeline stage 1006, acknowledgments for packets can be rejected in connection with determining to replay previously transmitted packets. Retiring a packet can include allowing other data to be written to memory in place of the packet and / or removing the packet from memory.

[0161] Thereafter, in a third pipeline stage 1008, logic circuitry 1018 may update the link pointer pointing to link “CAT” to point to another link (e.g., link “MOOSE”). Thus, in subsequent pipeline operations, logic circuits 1012, 1014, 1016, and 1018 may operate to determine whether to replay packets associated with link “MOOSE” based on another TX linked list (not shown in FIG. 10) that contains, references, or corresponds to packets transmitted over link “MOOSE.” Advantageously, using TX storage 434-1 and TX linked list 1020 to implement the replay functionality enables node 400 to communicate with a second node using TTP under limited hardware resources without the assistance of software control mechanisms.

[0162] Hardware link timer FIG. 11 shows an example block diagram of a hardware link timer 1100 that implements a timeout checking mechanism for replaying packets without software assistance. In some embodiments, the hardware link timer 1100 may be part of the node 400 of FIG. 4. Some or all of the hardware link timer 1100 may be deployed within the TTP tag block 436 of FIG. 4. As noted above, in contrast to other Ethernet protocols (e.g., TCP or UDP) where software is typically employed to track timeouts across multiple links using multiple timers (e.g., one timer per link), the hardware link timer 1100 enables the node 400 to determine which packets sent over which links to replay, and if replay is desired, when replaying occurs under limited hardware resources (e.g., when a large resource pool of virtual and / or physical address space and computer resources is not available). In some embodiments, the hardware link timer 1100 may periodically perform timing checks on an established link (e.g., an active link) utilized by the node 400 to communicate with one or more other nodes according to the TTP.

[0163] 11, hardware link timer 1100 may include first-in, first-out (FIFO) memory 1104, timer 1102, and logic circuits 1120, 1112, 1114, 1116, and 1118, which may be part of TTP tag block 436 for replaying packets. FIFO memory 1104 may store timing and status information associated with each active link. Hardware link timer 1100 may check the timing and status associated with each active link stored in FIFO memory 1104 in a round-robin manner. More specifically, the hardware link timer 1100 may start checking the timing and status information associated with the first link stored in the first entry of the FIFO memory 1104, moving toward the timing and status information associated with the Nth link stored in the Nth entry of the FIFO memory 1104, and then check the timing and status information associated with the first link stored in the first entry of the FIFO memory 1104 again. The hardware link timer 1100 may utilize the timer 1102 to schedule points in time for reading timing and status information associated with multiple active links and / or packets. The reading of the timing and status information may be used to determine whether to replay a packet associated with a link or to retire and / or discard the packet through further information retrieval. Note that the node 400 of FIG. 4 may include multiple hardware link timers similar to those shown in FIG. 11, and each hardware link timer may determine whether timeouts exist associated with multiple links.

[0164] In some embodiments, FIFO memory 1104 can store timing information associated with one or more links established between node 400 and other nodes. For example, node 400 may include a hardware link timer 1100 that uses FIFO memory 1104 to store timing information associated with M links established between node 400 and one or more other nodes, where M is a positive integer greater than 1. Instead of using M timers, each timer tracking timing information for a corresponding link, hardware link timer 1100 may utilize timer 1102 (e.g., a hardware clock that ticks once per programmable time period) to track and / or update timing information for each of the M links through accessing FIFO memory 1104 in a round-robin (e.g., circular) manner. Specifically, hardware link timer 1100 may access entries in FIFO memory 1104 in a round-robin manner, once per tick of timer 1102, with each accessed entry in FIFO memory 1104 corresponding to one of the M links.

[0165] In some embodiments, the time period of each timer 1102 tick may vary and may be on the order of hundreds of microseconds to single-digit microseconds. For example, the time period of the timer 1102 tick may be up to 100 microseconds, or even down to 1 microsecond. Additionally, the hardware link timer 1100 may adjust the time period of the timer 1102 tick based on the number of links (e.g., M) represented by the entries in the FIFO memory 1104. For example, as M (e.g., more links represented by the entries in the FIFO memory 1104) increases, the time period of the timer 1102 tick may decrease, and as M (e.g., fewer links represented by the entries in the FIFO memory 1104) decreases, the time period of the timer 1102 tick may increase. In this manner, the time interval at which link status and / or timing information is checked may remain constant while the time period of the timer 1102 tick varies inversely with the number of links represented by the entries in the FIFO memory 1104.

[0166] In some embodiments, timing and / or status information associated with one of the M links may indicate a time period during which the link has not received an acknowledgment for a received packet that was sent. Assuming node 400 has transmitted N packets over the link to a second node, one entry in FIFO memory 1104 may store timing and / or status information that, when accessed through a round-robin manner under a particular time period of timer 1102 ticks, indicates that no acknowledgment of receipt of any of the N packets has been received for a predetermined duration (e.g., 20 microseconds, 50 microseconds, 100 microseconds, 200 microseconds, 300 microseconds, 400 microseconds, 500 microseconds, and / or any duration therebetween). Upon accessing an entry in FIFO memory 1104, hardware link timer 1100 may utilize logic circuits 1120, 1112, 1114, 1116, and 1118 to check the timing and / or status information stored in the entry and retrieve N packets that may be stored in local storage (e.g., TX storage 434-1 or other local storage) of node 400 for replaying the N packets.

[0167] Alternatively, timing and / or status information associated with one of the M links may be stored in one entry of the FIFO memory 1104 to indicate that the link may be closed (e.g., all packets transmitted by the first node have been received by the second node). Upon accessing the entry of the FIFO memory 1104, the hardware link timer 1100, utilizing logic circuits 1120, 1112, 1114, 1116, and 1118, checks the timing and / or status information stored in the entry, searches for a packet that may still be stored in local storage of the node 400 (e.g., TX storage 434-1), and may discard the packet because the timing and / or status information stored in the entry of the FIFO memory 1104 indicates that the link may be closed. Advantageously, by utilizing a single timer (e.g., timer 1102) that ticks under adjustable periods for multiple links and / or packets, and a FIFO memory 1104 that stores timing and / or status information for multiple links, node 400 may replay packets at the appropriate times to achieve low latency and free up hardware resources occupied by inactive links (e.g., closed links) for use by active links to operate under limited computational and storage resources.

[0168] As shown in Figure 11, logic circuits 1120, 1112, 1114, 1116, and 1118 may operate at different pipeline stages, similar to logic circuits 1012, 1014, 1016, and 1018 shown in Figure 10. As shown in Figure 11, logic circuits 1120, 1112, 1114, 1116, and 1118 may operate in conjunction with timer 1102 and FIFO memory 1104 to determine when packets transmitted over one or more links need to be replayed or can be retired / discarded from local storage, such as TX storage 434-1, or whether one or more links can be closed. As shown in Figure 11, logic circuits 1120, 1112, 1114, 1116, and 1118 may operate at their respective pipeline stages according to a clock operated by hardware link timer 1100. Specifically, logic circuits 1120 and 1112 may operate in the initial pipeline stage (labeled "Q0"), logic circuit 1114 may operate in the first pipeline stage (labeled "Q1"), logic circuit 1116 may operate in the second pipeline stage (labeled "Q2"), and logic circuit 1118 may operate in the third pipeline stage (labeled "Q3").

[0169] In operation, in the initial pipeline stage Q0, logic circuit 1120 may select timing and status information to be used for logic circuit 1112's timing and status information search (e.g., TIMER link search). As shown in FIG. 11, the timing and status information may come from an entry from FIFO memory 1104 (e.g., the oldest entry that entered FIFO memory 1104 before all other entries) or from other sources (e.g., alternative preferred link search information). As shown in FIG. 11, in the initial pipeline stage Q0, the timing and status information associated with “Link A” in FIFO memory 1104 is selected by logic circuit 1112 based on a control signal (e.g., “Pick”) that selects “TIMER link search” rather than “TX traffic” or “RX traffic.” “TX traffic” may correspond to packets transmitted over a link established by node 400 (e.g., “Link B”), while “RX traffic” may correspond to packets received over another link established by node 400 (e.g., “Link D”).

[0170] In the first pipeline stage Q1, the logic circuit 1114 determines which link is being queried based on timing and status information received from the initial pipeline stage Q0. As shown in FIG. 11 , the logic circuit 1114 determines that “Link A” is being queried, for a later determination of whether “Link A” needs to be replayed or can be closed. Next, in the second pipeline stage Q2, the logic circuit 1116 determines whether “Link A” can be closed based on timing and status information associated with “Link A” accessed from the FIFO memory 1104. If the timing and status information associated with “Link A” indicates that “Link A” can be closed, the logic circuit 1116 may trigger packets associated with “Link A” to be retired / discarded from local storage (e.g., TX storage 434-1). If the timing and status information associated with "Link A" indicates that "Link A" is still active / open, the operation of the hardware link timer 1100 proceeds to a third pipeline stage Q3, where the logic circuit 1118 determines whether to replay packets sent over "Link A" or how to update the timing and status information associated with "Link A."

[0171] In the third pipeline stage Q3, the logic circuit 1118 may determine to replay at least some packets associated with “Link A” based on status and timing information associated with “Link A” accessed from the FIFO memory 1104. For example, the status and timing information associated with “Link A” may include a “TIMER BIT” that, when set (e.g., to a logical one), may indicate that an acknowledgment of receiving at least one of the packets associated with “Link A” has not been received by the node 400 for a threshold duration for replaying the packet. In some embodiments, the threshold duration may be adjustable and may be 20 microseconds, 50 microseconds, 100 microseconds, 200 microseconds, 300 microseconds, 400 microseconds, 500 microseconds, and / or any suitable duration therebetween. The threshold duration may range from 20 microseconds to 500 microseconds. In some embodiments, the “TIMER BIT” associated with “Link A” (and / or other links) may be set based on the number of times “Link A” has been queried from the FIFO memory 1104 and the time duration of the timer 1102.

[0172] If the “TIMER BIT” is asserted, logic circuit 1118 may replay packets associated with “Link A.” The assertion of the “TIMER BIT” may indicate that a timeout associated with one or more packets has occurred (e.g., a threshold duration has been reached without receiving an acknowledgment or negative acknowledgment). Furthermore, logic circuit 1118 may update timing and status information associated with “Link A” stored in FIFO memory 1104 in response to the replay of “Link A.” For example, logic circuit 1118 may clear the “TIMER BIT” (e.g., set the “TIMER BIT” from a logic 1 to a logic 0). On the other hand, if the status and timing information associated with “Link A” indicates not to replay one or more packets on “Link A” (e.g., the “TIMER BIT” is not asserted, which corresponds to a logic 0 in FIG. 11 ), logic circuit 1118 may not replay “Link A.” In such a situation, if the timing and status information associated with "Link A" indicates that "Link A" should be replayed the next time it is queried, the logic circuit 1118 may further set the "TIMER BIT" to a logic one.

[0173] Exemplary Methods for Replay and Link Timing 12, an exemplary packet replay procedure 1200 is described for replaying packets transmitted from a node, such as node 400 or device A of FIG. 3B. Packet replay procedure 1200 may be implemented, for example, by TTP tag block 436 of FIG. 4 or other components of node 400. Procedure 1200 begins at block 1202, where TTP tag block 436 may store a linked list containing packets transmitted from node 400 over a first link to a second node using an Ethernet protocol. For example, the linked list may be TX linked list 1020, which contains or references packets 1022, 1024, and 1026 to maintain the order of packets 1022, 1024, and 1026 for transmission to the second node.

[0174] At block 1204, the TTP tag block 436 may determine to replay a first one of the packets in response to at least one of (a) receiving a negative acknowledgment of the first packet from the second node, or (b) a timeout associated with the first packet. For example, the TTP tag block 436 may determine to replay packet 1024 in response to (a) receiving a negative acknowledgment of packet 1024 from the second node, or (b) a timeout associated with packet 1024 indicating that an acknowledgment of packet 1024 was not received for a threshold time period.

[0175] In block 1206, the TTP tag block 436 may retire a second one of the packets in response to receiving an acknowledgment of the second packet from the second node. For example, the TTP tag block 436 may retire packet 1022 in response to receiving an acknowledgment of packet 1022 from the second node.

[0176] FIG. 13 illustrates an exemplary link timeout procedure 1300 for determining whether to replay one or more links associated with a node, such as node 400 or device A of FIG. 3B. Link timeout procedure 1300 may be implemented, for example, by hardware link timer 1100 or node 400 of FIG. 11. Procedure 1300 begins at block 1302, where hardware link timer 1100 or node 400 stores timing and status information associated with multiple links in a FIFO memory, and node 400 transmits packets over the multiple links to one or more other nodes using an Ethernet protocol. For example, hardware link timer 1100 may store timing and status information associated with multiple links in FIFO memory 1104.

[0177] In block 1304, the hardware link timer 1100 or the node 400 may access an entry in the FIFO memory based on each tick of a hardware timer deployed within the hardware link timer 1100 or the node 400. For example, the hardware link timer 1100 may access an entry in the FIFO memory 1104 based on each tick of the timer 1102.

[0178] In block 1306, the hardware link timer 1100 or the node 400 may determine to replay at least one packet associated with a first link of the plurality of links based on timing and status information associated with the first link. For example, the hardware link timer 1100 may determine to replay at least one packet associated with or transmitted over “Link A” based on timing and status information associated with “Link A.”

[0179] ·Conclusion The foregoing disclosure is not intended to limit the disclosure to the precise form or particular field of use disclosed. Accordingly, various alternative embodiments and / or modifications to the disclosure, whether expressly described or implied herein, are contemplated as possible in light of the present disclosure. While embodiments of the present disclosure have been described in this manner, those skilled in the art will recognize that changes can be made in form and detail without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

[0180] It should be understood that not necessarily all objectives or advantages will be achieved in accordance with any particular example described herein. Thus, for example, those skilled in the art will recognize that some examples may be manipulated to achieve or optimize one advantage or advantages as taught herein without necessarily achieving other objectives or advantages that may be taught or suggested herein.

[0181] All of the processes described herein may be embodied in software code modules executed by a computer system including multiple computers or processors, and may be fully automated via the software code modules. The code modules may be stored on any type of non-transitory computer-readable medium or other computer storage device. Some or all of the methods may be embodied in dedicated computer hardware.

[0182] Many other variations beyond those described herein will be apparent from this disclosure. For example, depending on the example, acts, events, or functions of any of the algorithms described herein may be performed in a different order, added, merged, or omitted entirely (e.g., not all acts or events described may be necessary to implement an algorithm). Furthermore, in some embodiments, acts or events may be performed simultaneously rather than sequentially, for example, through multithreading, interrupt processing, or multiple processors or processor cores, or on other parallel architectures. Furthermore, various tasks or processes may be performed by different machines and / or computer systems that may function together.

[0183] The various illustrative logic blocks and modules stored in connection with the examples disclosed herein may be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. The processor may be a microprocessor, but in alternative examples, the processor may be a controller, microcontroller, or state machine, combinations thereof, etc. A processor may include electrical circuitry that processes computer-executable instructions. In some embodiments, a processor includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, a microprocessor in conjunction with a DSP core, or any other such configuration. While described herein primarily with reference to digital technology, a processor may also include primarily analog components. The computing environment can include any type of computing system, including, but not limited to, a computing system based on a microprocessor, mainframe computer, digital signal processor, portable computing device, device controller, or computing engine within an appliance, to name a few.

[0184] Elements of a method, process, routine, or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor device, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium. An exemplary storage medium may be coupled to the processor device such that the processor device can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor device. The processor device and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor device and the storage medium may reside as discrete components in a user terminal.

[0185] The processes described herein or illustrated in the figures of this disclosure may be initiated in response to an event, such as a predetermined or dynamically determined schedule, on demand when initiated by a user or system administrator, or in response to some other event. When such processes begin, a set of executable program instructions stored on one or more non-transitory computer-readable media (e.g., hard drives, flash memory, removable media, etc.) may be loaded into memory (e.g., RAM) of a server or other computing device. The executable instructions may then be executed by a hardware-based computer processor of the computing device. In some embodiments, such processes, or portions thereof, may be implemented across multiple computing devices and / or multiple processors, either serially or in parallel.

[0186] In particular, conditional language such as "can," "could," "might," or "may" is typically understood within the context in which it is generally used to convey that some examples include certain features, elements, and / or steps, while other examples do not, unless otherwise specified. Thus, such conditional language is generally not intended to imply that the features, elements, and / or steps are somehow methodical to the examples, or that the examples necessarily include logic for determining whether those features, elements, and / or steps should be included in or performed in any particular example, with or without user input or prompting.

[0187] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is typically understood in the context in which it is commonly used to indicate that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless otherwise specified. Thus, such disjunctive language generally does not imply, and should not be intended to imply, that some instances require at least one of X, at least one of Y, or at least one of Z, respectively, to be present.

[0188] Any process descriptions, elements, or blocks in the flow diagrams described herein and / or shown in the accompanying drawings should be understood as potentially representing modules, segments, or portions of code, including executable instructions for implementing specific logical functions or elements in the process. As will be appreciated by those skilled in the art, alternative examples are included within the scope of the examples described herein, in which elements or functions may be omitted, performed, or described in a different order than that shown or described, including substantially simultaneously or in reverse order, depending on the functionality involved.

[0189] It is emphasized that many variations and modifications can be made to the above-described examples, and that the elements thereof are to be understood as being among the other acceptable examples, and all such modifications and variations are intended to be included herein within the scope of this disclosure.

[0190] Any process descriptions, elements, or blocks in the flow diagrams described herein and / or shown in the accompanying drawings should be understood as potentially representing modules, segments, or portions of code, including executable instructions for implementing specific logical functions or elements in the process. As will be appreciated by those skilled in the art, depending on the functionality involved, alternative implementations are included within the scope of the examples described herein, in which elements or functions may be omitted, performed substantially simultaneously, or in the order shown or described, including in reverse order.

[0191] Unless otherwise specified, articles such as "a" or "an" should generally be construed to include one or more listed items. Thus, phrases such as "a device configured to" are intended to include one or more listed devices. Such one or more listed devices may also be collectively configured to perform the stated enumeration. For example, "a processor configured to perform enumerations A, B, and C" may include a first processor configured to perform enumeration A working in conjunction with a second processor configured to perform enumerations B and C.

Claims

1. a first node, a hardware replay architecture configured to replay packets transmitted over a first link to a second node using an Ethernet protocol; The hardware replay architecture includes: a local storage device configured to store a linked list containing the packets, the linked list maintaining an order of the packets for transmission to the second node; 1. A logic circuit including a plurality of pipeline stages, the logic circuit comprising: determining, in a first pipeline stage of the plurality of pipeline stages, to process data associated with the first link but not a second link between the first node and the second node; determining to replay the first one of the packets in response to at least one of: (a) receiving a negative acknowledgment of the first packet from the second node; or (b) a timeout associated with the first packet; logic configured to retire the second one of the packets in response to receiving an acknowledgment of the second packet from the second node; A first node, wherein the Ethernet protocol is lossy.

2. 2. The first node of claim 1, wherein the logic circuitry determines to replay the first packet at a second pipeline stage of the plurality of pipeline stages.

3. 3. The first node of claim 2, wherein the logic circuitry determines to replay a third one of the packets and the first one of the packets in the second pipeline stage of the plurality of pipeline stages based on the order of the packets maintained by the linked list.

4. 2. The first node of claim 1, wherein the logic circuitry determines to process data associated with the first link rather than the second link based on a link pointer, and wherein the logic circuitry updates the link pointer to point to the second link at a third pipeline stage of the plurality of pipeline stages.

5. The first node of claim 1 , wherein the first node and the second node are in an Ethernet-based network, and the first node communicates with the second node through an Ethernet switch.

6. 6. The first node of claim 5, wherein the first node comprises a network interface processor (NIP) and a high bandwidth memory (HBM), the bandwidth of the HBM being at least 1 gigabyte.

7. A first node for Ethernet-based communications, comprising: one or more processors configured to implement a transport layer hardware-only Ethernet protocol; the transport layer hardware-only Ethernet protocol is lossy; the one or more processors comprise a hardware replay architecture including a plurality of pipeline stages, the hardware replay architecture configured to replay packets transmitted under the transport layer hardware-only Ethernet protocol; The first node, wherein replaying the packet includes determining, at a pipeline stage among the plurality of pipeline stages, to process data associated with a first link rather than a second link between the first node and a second node.

8. the hardware replay architecture comprises:

8. The first node of claim 7, comprising a local storage device configured to store the packets transmitted under the transport layer hardware-only Ethernet protocol.

9. the hardware replay architecture comprises:

9. The first node of claim 8, further comprising a linked list stored in the local storage device and configured to track the order of the packets for transmission to another node, each element of the linked list corresponding to each of the packets stored in the local storage device.

10. The first node of claim 9 , wherein the hardware replay architecture is configured to transmit packets in an order corresponding to the linked list.

11. the hardware replay architecture comprises: a first pointer configured to point to a first element of the linked list, the first pointer indicating that a first packet of the packets corresponding to the first element of the linked list is not to be replayed; a second pointer configured to point to a second element of the linked list, the second pointer indicating to replay a second packet of the packets corresponding to the second element of the linked list.

12. 12. The first node of claim 11, wherein the hardware replay architecture replays the second packet and one or more packets following the second packet according to an order of the packets for transmission.

13. 13. The first node of claim 12, wherein the hardware replay architecture causes the local storage device to discard the first packet and one or more packets preceding the second packet according to an order of the packets for transmission.

14. 1. A computer-implemented method implemented at a first node for replaying packets transmitted over a first link to a second node using an Ethernet protocol, the first node comprising a hardware replay architecture including a plurality of pipeline stages; storing a linked list containing the packets, the linked list maintaining an order of the packets for transmission to the second node; determining, in a first pipeline stage of the plurality of pipeline stages, to process data associated with a first link but not a second link between the first node and a second node; determining to replay a first one of the packets in response to at least one of: (a) receiving a negative acknowledgment of the first packet from the second node; or (b) a timeout associated with the first packet; retiring the second one of the packets in response to receiving an acknowledgment of the second packet from the second node; A computer-implemented method, wherein the Ethernet protocol is lossy.

15. 15. The computer-implemented method of claim 14, wherein the hardware replay architecture determines to replay the first packet at a second pipeline stage of the plurality of pipeline stages.

16. 16. The computer-implemented method of claim 15, wherein the hardware replay architecture determines to replay a third one of the packets and the first one of the packets in the second pipeline stage of the plurality of pipeline stages based on the order of the packets maintained by the linked list.

17. 15. The computer-implemented method of claim 14, wherein the first node and the second node are in an Ethernet-based network, and the first node communicates with the second node through an Ethernet switch.

18. 20. The computer-implemented method of claim 17, wherein the first node comprises a network interface processor (NIP) and a high-bandwidth memory (HBM), the bandwidth of the HBM being at least 1 gigabyte.