Link timer for Ethernet

By adopting hardware playback architecture and state machine flow control protocols in Ethernet networks, the problems of high latency and system complexity in the prior art are solved, and low latency and efficient communication are achieved, suitable for high-performance computing and artificial intelligence training.

CN120035982APending Publication Date: 2025-05-23TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070910.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-19
Filing Date
2023-08-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing Ethernet networks have problems such as high latency, insufficient bandwidth, system complexity and cost in high performance computing and artificial intelligence training data centers, and it is difficult to meet the needs of low latency, large bandwidth and distributed control.

Method used

A hardware-based flow control protocol is adopted to realize low-latency Ethernet communication using the hardware playback architecture and state machine. By implementing hardware playback and link control under the hardware Ethernet protocol in the transport layer, the dependence on the central processing unit is reduced.

Benefits of technology

It realizes communication with low latency and low software overhead in lossy Ethernet networks, meets the needs of high-performance computing and artificial intelligence training, and improves system performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035982A_ABST
    Figure CN120035982A_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for communicating in an Ethernet-based network using a transport layer without the aid of a software control mechanism. In some embodiments, a first node includes a hardware link timer configured to determine packets to be sent under a transport layer hardware only Ethernet protocol for playback. The hardware link timer may include a first-in first-out (FIFO) memory configured to store timing and state information associated with one or more links established by the first node. The hardware link timers may also include timers associated with one or more links that are ticked according to the time period.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a non-provisional patent application of U.S. Provisional Patent Application No. 63 / 373,016, filed on August 19, 2022, entitled “TRANSPORT PROTOCOL FOR ETHERNET”, and claims priority, the technical disclosure of which is hereby incorporated by reference in its entirety for all purposes. This application is a non-provisional patent application of U.S. Provisional Patent Application No. 63 / 503,349, filed on May 19, 2023, entitled “TRANSPORT PROTOCOL FOR ETHERNET”, and claims priority, the technical disclosure of which is hereby incorporated by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure relates to systems and methods for facilitating network communications. More specifically, embodiments of the present disclosure relate to a hardware-implementable flow control protocol for communicating over an Ethernet-based network. Background Art

[0004] The Institute of Electrical and Electronics Engineers (IEEE) has provided various standards for local area networks (LANs), collectively referred to as IEEE802, including the IEEE 802.3 standard, commonly known as Ethernet. The IEEE 802.3 Ethernet standard has specifications for the physical media interface (Ethernet cable, optical fiber, backplane, etc.), but no specifications for the flow control of the communication. Protocols such as TCP / IP, RoCE, or InfiniBand can accelerate the fabric flow control. The TCP / IP protocol typically has millisecond latency, while RoCE or InfiniBand have lossless and scaling specifications, which may overly constrain the system.

[0005] As high-performance computing (HPC) and artificial intelligence (AI) training data centers become more common, communication network structures with high bandwidth, low latency, large-scale lossy resilience, distributed control, and as little software overhead as possible are needed. Therefore, it may be necessary to develop network flow control protocols that can run over lossy Ethernet-based networks with little or no central processing unit (CPU) involvement while achieving lower latency than existing Ethernet-based networks. Summary of the invention

[0006] The systems, methods, and devices of the present disclosure each have several innovative embodiments, no single one of which is solely responsible for all of the desirable attributes disclosed herein.The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below.

[0007] In some aspects, the techniques described herein relate to a first node for Ethernet-based communications, the first node comprising: one or more processors configured to implement a transport layer hardware-only Ethernet protocol.

[0008] In some aspects, techniques described herein relate to a first node where an Ethernet protocol is lossy.

[0009] In some aspects, the techniques described herein relate to a first node, wherein one or more processors are further configured to implement a hardware replay architecture to replay packets sent via a first link to a second node, wherein the packets are stored in a local storage device of the first node, and wherein an order of the packets for replay is specified in a linked list.

[0010]

[0013] In some aspects, techniques described herein relate to a first node configured to send packets to a second node with single-digit microsecond latency.

[0011] In some aspects, the technology described herein relates to a first node, wherein one or more processors are configured to implement a state machine, the state machine being configured to: operate in an open state in which a link between the first node and a second node is open; transition from the open state to an intermediate closed state; and in response to receiving a close confirmation from the second node, transition from the intermediate closed state to the closed state to close the link.

[0012] In some aspects, the techniques described herein relate to a first node that also includes an Ethernet port.

[0013] In some aspects, the technology described herein relates to a first node in which one or more processors are configured to determine to replay packets on a link between the first node and a second node based on timing and status information associated with the link stored in a first-in-first-out (FIFO) memory, wherein entries of the FIFO memory are accessed based on ticks of a hardware link timer associated with multiple links.

[0014] In some aspects, the techniques described herein relate to a first node for Ethernet-based communications, the first node comprising: one or more processors configured to implement a layer 2 hardware-only Ethernet protocol.

[0015] In some aspects, the techniques described herein relate to a first node in which one or more processors include a hardware-only architecture configured to replay packets sent over a first link to a second node.

[0016] In some aspects, the techniques described herein relate to a first node, wherein one or more processors are further configured to determine to replay a packet on a link associated with the first node based on timing and status information associated with the link stored in a first-in-first-out (FIFO) memory, the FIFO memory being accessed based on ticks of timers associated with multiple links.

[0017] In some aspects, the technology described herein relates to a first node, wherein the first node is configured to open and close a link with a second node in an Ethernet-based network, the first node comprising: state machine hardware, the state machine hardware configured to: operate in an open state in which the link between the first node and the second node is open; transition from the open state to an intermediate closed state; and in response to receiving a close confirmation from the second node, transition from the intermediate closed state to the closed state to close the link, wherein the first node is configured to operate in a lossy network.

[0018] In some aspects, the techniques described herein relate to a first node in which state machine hardware implements a flow control protocol for a transport layer in hardware only.

[0019] In some aspects, techniques described herein relate to a first node wherein a delay associated with a flow control protocol is less than 10 microseconds.

[0020] In some aspects, techniques described herein relate to a first node in which state machine hardware is configured to: transition from an off state to an intermediate on state; and transition from the intermediate on state to the on state.

[0021] In some aspects, techniques described herein relate to a first node in which state machine hardware transitions from an open state to an intermediate closed state in response to sending or receiving a request to close a link to or from a second node.

[0022] In some aspects, techniques described herein relate to a first node in which state machine hardware transitions from an intermediate shutdown state to a shutdown state in response to sending a confirmation to a second node to shut down a link.

[0023] In some aspects, techniques described herein relate to a first node in which state machine hardware transitions from an intermediate shutdown state to a shutdown state without waiting for a period of time.

[0024] In some aspects, the techniques described herein relate to a first node, wherein in an open state, the first node does not retransmit a packet until a negative acknowledgment of the packet is received from a second node or a predetermined timeout period expires without receiving a negative acknowledgment of the packet.

[0025] In some aspects, techniques described herein relate to a first node, wherein in an open state, the first node sends a maximum of N packets without pausing, and wherein N is limited by a size of physical memory allocated to the first node.

[0026] In some aspects, the techniques described herein relate to a first node, further comprising: a hardware link timer associated with a plurality of links; and a hardware replay architecture configured to replay packets in hardware only.

[0027] In some aspects, the technology described herein relates to a first node, the first node comprising: a hardware replay architecture configured to replay packets sent to a second node over a first link using an Ethernet protocol, wherein the hardware replay architecture comprises: a local storage device configured to store a linked list comprising packets, wherein the linked list maintains an order of packets for sending to the second node; and a logic circuit system configured to determine to replay a first packet of packets in response to at least one of: (a) receiving a negative acknowledgment of the first packet from the second node, or (b) a timeout associated with the first packet; and to withdraw the second packet in response to receiving an acknowledgment of the second packet of packets from the second node, wherein the Ethernet protocol is lossy.

[0028] In some aspects, the techniques described herein relate to a first node wherein logic circuitry includes a plurality of pipeline stages, and wherein the logic circuitry determines to process data associated with a first link at a first pipeline stage of the plurality of pipeline stages instead of a second link between the first node and a second node.

[0029] In some aspects, techniques described herein relate to a first node in which logic circuitry determines to replay a first packet at a second pipeline stage of a plurality of pipeline stages.

[0030] In some aspects, techniques described herein relate to a first node in which logic circuitry determines to replay a third packet of packets and a first packet of packets at a second pipeline stage of a plurality of pipeline stages based on an order of packets maintained by a linked list.

[0031] In some aspects, the techniques described herein relate to a first node in which logic circuitry determines, based on a link pointer, to process data associated with a first link instead of a second link, and in which the logic circuitry updates the link pointer at a third pipeline stage of a plurality of pipeline stages to point to the second link.

[0032] In some aspects, the techniques described herein relate to a first node, wherein the first node and a second node are located in an Ethernet-based network, and wherein the first node communicates with the second node through an Ethernet switch.

[0033] In some aspects, techniques described herein relate to a first node, wherein the first node includes a network interface processor (NIP) and a high bandwidth memory (HBM), and wherein a bandwidth of the HBM is at least one gigabyte.

[0034] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node comprising: one or more processors configured to implement a transport layer hardware-only Ethernet protocol, wherein the transport layer hardware-only Ethernet protocol is lossy, and wherein the one or more processors include a hardware replay architecture, the hardware replay architecture configured to replay packets sent under the transport layer hardware-only Ethernet protocol.

[0035] In some aspects, the techniques described herein relate to a first node in which a hardware replay architecture includes a local storage device configured to store packets sent under a transport layer hardware-only Ethernet protocol.

[0036] In some aspects, the techniques described herein relate to a first node, wherein a hardware replay architecture comprises: a linked list stored in a local storage device and configured to track an order of packets for transmission to another node, wherein each element of the linked list corresponds to each packet in the packets stored in the local storage device.

[0037] In some aspects, techniques described herein relate to a first node in which a hardware replay architecture is configured to send packets in an order corresponding to a linked list.

[0038] In some aspects, the technology described herein relates to a first node in which a hardware replay architecture is configured to store: a first pointer configured to point to a first element of a linked list, wherein the first pointer indicates that a first group of groups corresponding to the first element of the linked list is not to be replayed; and a second pointer configured to point to a second element of the linked list, wherein the second pointer indicates that a second group of groups corresponding to the second element of the linked list is to be replayed.

[0039] In some aspects, techniques described herein relate to a first node in which a hardware replay architecture replays a second packet and one or more packets subsequent to the second packet according to an order of packets for transmission.

[0040] In some aspects, techniques described herein relate to a first node in which a hardware replay architecture causes a local storage device to discard a first packet and one or more packets preceding a second packet according to an order of packets for transmission.

[0041] In some aspects, the technology described herein relates to a computer-implemented method implemented at a first node for replaying packets sent to a second node over a first link using an Ethernet protocol, the computer-implemented method comprising: storing a linked list comprising packets, wherein the linked list maintains an order of packets for sending to the second node; determining to replay a first packet of the packets in response to at least one of: (a) receiving a negative acknowledgment of the first packet from the second node, or (b) a timeout associated with the first packet; and withdrawing the second packet in response to receiving an acknowledgment of the second packet of the packets from the second node, wherein the Ethernet protocol is lossy.

[0042] In some aspects, the technology described herein relates to a computer-implemented method in which a first node includes a hardware replay architecture that includes multiple pipeline stages, and in which the hardware replay architecture determines that data associated with a first link but not a second link is to be processed at a first pipeline stage of the multiple pipeline stages.

[0043] In some aspects, the techniques described herein relate to a computer-implemented method in which a hardware replay architecture determines to replay a first packet at a second pipeline stage of a plurality of pipeline stages.

[0044] In some aspects, the techniques described herein relate to a computer-implemented method in which a hardware replay architecture determines a third packet of packets and a first packet of packets to replay at a second pipeline stage of a plurality of pipeline stages based on an order of the packets maintained by a linked list.

[0045] In some aspects, the technology described herein relates to a computer-implemented method in which a first node and a second node are located in an Ethernet-based network, and in which the first node communicates with the second node through an Ethernet switch.

[0046] In some aspects, the techniques described herein relate to a computer-implemented method wherein a first node includes a network interface processor (NIP) and a high bandwidth memory (HBM), and wherein a bandwidth of the HBM is at least one gigabyte.

[0047] In some aspects, the technology described herein relates to a first node for sending packets in an Ethernet-based network, the first node comprising: one or more processors, the one or more processors comprising: a first-in-first-out (FIFO) memory configured to store timing and status information associated with multiple links, wherein the first node is configured to send packets to one or more other nodes through the multiple links using an Ethernet protocol; a timer configured to tick according to a time period, wherein the timer is associated with the multiple links; and a logic circuit system configured to: access entries of the FIFO memory based on corresponding ticks on the timer; and determine to replay at least one packet associated with a first link of the multiple links based on the timing and status information associated with the first link, wherein the Ethernet protocol is lossy.

[0048] In some aspects, techniques described herein relate to a first node in which logic circuitry is configured to access entries of a FIFO memory in a polling manner.

[0049] In some aspects, techniques described herein relate to a first node, wherein a timer is configured to adjust a time period based on a number of active links associated with an entry of a FIFO memory, wherein the active link is included in a plurality of links.

[0050] In some aspects, the techniques described herein relate to a first node in which logic circuitry is configured to determine to withdraw packets associated with a second link of a plurality of links based on timing and state information associated with the second link.

[0051] In some aspects, the technology described herein relates to a first node, wherein packets associated with a second link are stored in a local storage device of the first node, and wherein in response to determining that the packets associated with the second link are to be withdrawn, the logic circuit system causes the local storage device to discard the packets associated with the second link.

[0052] In some aspects, the techniques described herein relate to a first node in which logic circuitry is configured to determine to shut down a second link of a plurality of links based on timing and status information associated with the second link.

[0053] In some aspects, the techniques described herein relate to a first node, wherein timing and status information associated with a first link of a plurality of links indicates that an acknowledgment of receiving at least one packet associated with the first link has not been received by the first node within a threshold duration for replaying packets.

[0054] In some aspects, the technology described herein relates to a first node for Ethernet-based communications, the first node comprising: one or more processors configured to implement a transport layer hardware-only Ethernet protocol, wherein the transport layer hardware-only Ethernet protocol is lossy, and wherein the one or more processors include a hardware link timer, the hardware link timer configured to determine packets sent for replay under the transport layer hardware-only Ethernet protocol.

[0055] In certain aspects, the technology described herein relates to a first node, wherein the first node sends a first plurality of packets over a first link and sends a second plurality of packets over a second link according to a transport layer hardware-only Ethernet protocol, and wherein the hardware link timer includes: a first-in-first-out (FIFO) memory configured to store timing and status information associated with the first link in a first entry of the FIFO memory, and to store timing and status information associated with the second link in a second entry of the FIFO memory.

[0056] In some aspects, the technology described herein relates to a first node, wherein a hardware link timer includes a timer associated with a plurality of links that ticks according to a time period, wherein the hardware link timer accesses entries of a FIFO memory in a polling manner of the timer, wherein the entries include a first entry and a second entry.

[0057] In some aspects, techniques described herein relate to a first node, wherein a hardware link timer is configured to adjust a time period based on a number of active links associated with an entry of a FIFO memory, and wherein the active links include a first link and a second link.

[0058] In some aspects, the technology described herein relates to a first node, wherein a hardware link timer is configured to: determine to replay at least some of a first plurality of packets based on timing and status information associated with a first link stored in a first entry of a FIFO memory; and determine to withdraw a second plurality of packets based on timing and status information associated with a second link stored in a second entry of the FIFO memory.

[0059] In some aspects, the techniques described herein relate to a first node, wherein a second plurality of packets are stored in a local storage of the first node, and wherein in response to determining that the second plurality of packets are to be withdrawn, a hardware link timer causes the local storage to discard the second plurality of packets.

[0060] In some aspects, the techniques described herein relate to a first node, wherein timing and state information associated with a first link indicates that an acknowledgement of receiving a packet of a first plurality of packets has not been received by the first node within a threshold duration for replaying packets.

[0061] In some aspects, the technology described herein relates to a computer-implemented method implemented at a first node in an Ethernet-based network, the computer-implemented method comprising: storing timing and status information associated with multiple links in a first-in-first-out (FIFO) memory of the first node, wherein the first node is configured to send packets to one or more other nodes over the multiple links using an Ethernet protocol; accessing entries of the FIFO memory based on corresponding ticks of a hardware timer; and determining to replay at least one packet associated with a first link of the multiple links based on the timing and status information associated with the first link, wherein the Ethernet protocol is lossy.

[0062] In some aspects, the techniques described herein relate to a computer-implemented method in which entries of a FIFO memory are accessed in a polling manner.

[0063] In some aspects, the techniques described herein relate to a computer-implemented method, further comprising adjusting a time period of a hardware timer based on a number of active links associated with an entry of a FIFO memory, wherein the active link is included in a plurality of links.

[0064] In some aspects, the techniques described herein relate to a computer-implemented method, further comprising determining to withdraw packets associated with a second link of a plurality of links based on timing and state information associated with the second link.

[0065] In some aspects, the techniques described herein relate to a computer-implemented method further comprising causing at least one packet associated with a first link to be replayed.

[0066] In some aspects, the technology described herein relates to a computer-implemented method in which timing and status information associated with a first link of a plurality of links indicates that an acknowledgment of receiving at least one packet associated with the first link has not been received by a first node within a threshold duration for replaying packets.

[0067] In some aspects, the techniques described herein relate to all of the embodiments described and discussed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Throughout the drawings, reference numerals are repeatedly used to indicate correspondence between reference elements. The drawings are provided to illustrate examples of the subject matter described herein and not to limit the scope thereof.

[0069] Embodiments of the present disclosure are described with reference to the accompanying drawings, wherein like reference numerals represent like elements, and in which:

[0070] Figure 1A-1B is a table showing example protocols operating at different layers of the Open Systems Interconnection (OSI) model.

[0071] Figure 2 Depicted is an example state machine for opening and closing links between nodes implementing the Tesla Transport Protocol (TTP) according to an embodiment of the present disclosure.

[0072] Figure 3A-3B is an example timing diagram depicting the transmission and reception of packets between two devices implementing TTP according to an embodiment of the present disclosure.

[0073] Figure 4 An exemplary schematic block diagram of a node implementing TTP according to an embodiment of the present disclosure is illustrated.

[0074] Figure 5 Depicted are example headers of packets sent or received according to a TTP according to an embodiment of the present disclosure.

[0075] Figure 6 An example network and computing environment is illustrated in which embodiments of the present disclosure may be implemented.

[0076] Figure 7A-7B Operation codes for different types of TTP packets are shown according to some embodiments of the present disclosure.

[0077] Figure 8 An example physical storage device for storing packets for replaying packets sent and / or received under a lossy protocol, such as TTP, is illustrated according to some embodiments of the present disclosure.

[0078] Fig. 9 Depicted are example data structures (eg, linked lists) for tracking and maintaining a transmission order for sending and replaying packets according to some embodiments of the present disclosure.

[0079] Fig.10 An example block diagram of at least a portion of a hardware replay architecture for replaying packets sent over multiple links according to some embodiments of the present disclosure is illustrated.

[0080] Fig.11 An example block diagram of a hardware link timer implementing a timeout checking mechanism for replaying packets without software assistance is illustrated in accordance with some embodiments of the present disclosure.

[0081] Fig.12An illustrative routine for replaying packets sent from a node is illustrated in accordance with some embodiments of the present disclosure.

[0082] Fig.13 An example routine for determining whether to replay one or more links associated with a node is depicted. DETAILED DESCRIPTION

[0083] The following detailed description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be embodied in a variety of different ways, for example, as defined and covered by the claims. In this specification, reference is made to the accompanying drawings in which the same reference numerals and / or terms may represent the same or functionally similar elements. It should be understood that the elements shown in the drawings are not necessarily drawn to scale. In addition, it should be understood that some embodiments may include more elements and / or a subset of the elements shown in the drawings than shown in the drawings. In addition, some embodiments may combine any suitable combination of features from two or more drawings. The headings are provided for convenience only and do not affect the scope or meaning of the claims.

[0084] In general, one or more aspects of the present disclosure correspond to systems and methods for controlling network traffic using hardware mechanisms (e.g., without software assistance). More specifically, some embodiments of the present disclosure disclose a flow control protocol compatible with Ethernet standards that can be implemented by hardware circuits to achieve low latency, such as delays in the single digit microsecond range. In some embodiments, the single digit microsecond delay is at least partially achieved by utilizing a hardware-controlled state machine to simplify the opening and closing of communication links between network nodes. In addition, the disclosed flow control protocol (e.g., Tesla Transport Protocol (TTP)) can limit the number of packets sent / retransmitted on an established link and / or the duration of a waiting period before transitioning to the next state of a hardware-controlled state machine. This helps to achieve low-latency communication. Advantageously, the flow control protocol disclosed herein enables a pure hardware implementation of up to Layer 4 (transport layer) of the Open Systems Interconnection (OSI) model.

[0085] Some aspects of the present disclosure relate to flow control designed to run only on hardware. Such flow control can be implemented without software flow control or central processing unit (CPU) / kernel involvement. This can allow IEEE 802.3 Ethernet capabilities with latency limited only or substantially limited by physics. For example, single-digit microsecond latency can be achieved.

[0086] Tesla Ethernet Transport Protocol (TTP) is a hardware-only Ethernet flow control protocol that can implement the transport layer in the OSI model. Layer 2 (L2) Ethernet flow control is implemented only in hardware. Layer 3 and / or Layer 4 Ethernet flow control can also be implemented only in hardware. Link control, timers, congestion, and replay functions can be implemented in hardware. TTP can be implemented in network interface processors and network interface cards. TTP can enable full I / O batching configuration. TTP is a lossy protocol. In a lossy protocol, lost data can be recovered. For example, in a lossy protocol, any lost or damaged packets can be replayed (e.g., retransmitted) and recovered until reception is confirmed.

[0087] The L2 header, state machine, and opcodes in the present disclosure may define such a hardware-only protocol (eg, TTP) that may recover from lost packets in an N-to-N set of links.

[0088] In addition, some embodiments of the present disclosure disclose a hardware replay architecture (e.g., microarchitecture) that is capable of replaying packets sent and / or received under a lossy protocol (such as TTP). As described above, TTP (or TTPoE) is a hardware-only Ethernet flow control protocol. TTP can facilitate the implementation of extremely low latency (e.g., (multiple) single-digit microseconds) structures for HPC and / or AI training systems. In order to implement a lossy Ethernet flow control protocol without the assistance of a software control mechanism, some aspects of the present disclosure describe a hardware replay architecture that can buffer, hold, confirm and / or replay packets so that any lost or damaged packets can be replayed and recovered until reception is confirmed.

[0089] In order to replay packets sent and / or received according to a lossy Ethernet protocol (such as TTP) using only hardware resources, some embodiments of the disclosed hardware replay architecture utilize physical storage devices and data structures to store packets sent and / or received in different links and maintain the order of the sent packets, especially when replay occurs. In some embodiments, the physical storage device can be any type of local storage device or cache (e.g., a low-level cache) that stores, buffers, or saves packets associated with one or more links. The size of the physical storage device can be limited, such as a size of megabytes (MB) or kilobytes (KB). In some embodiments, the data structure can include one or more linked lists, each of which can record and / or track the order of packets sent for a link established between a first communication node and a second communication node. Advantageously, implementing a replay mechanism for a lossy protocol using a hardware replay architecture (which employs a physical storage device of limited size and linked lists that track the order of packets for various links) allows communication nodes to operate in accordance with TTP under limited hardware resources (e.g., when virtual processing or storage resources are not available).

[0090] In addition, some embodiments of the present disclosure relate to a hardware link timer that implements timeout checking without the help of a software-controlled mechanism. Some aspects of the present disclosure describe a hardware link timer that uses a single timer that can track timeouts on multiple links by coordinating with a first-in, first-out (FIFO) memory, rather than using multiple timers to track timeouts on each link. More specifically, entries of the FIFO memory can store the state and / or timer information of the link, and the hardware link timer can access the entries of the FIFO memory in a polling manner to determine whether a packet associated with the link can be discarded or needs to be retained. If the hardware link timer determines that a packet associated with the link can be discarded, then under limited hardware resources, more space can be used to store packets associated with another link. If the hardware link timer determines that one or more packets associated with the link should be retained, the (multiple) reserved packets associated with the link can enable the communication node hosting the hardware link timer to replay the (multiple) reserved packets.

[0091] Ethernet is an established standard technology for wired communications. In recent years, Ethernet has also been used in the automotive industry for a variety of vehicle applications. Typically, the latency associated with Ethernet communications ranges from hundreds of microseconds to more than a few milliseconds. In addition to physical limitations (e.g., the speed of signal transmission on the communication medium), the complexity of the associated protocols that control Ethernet data flow often becomes another bottleneck for latency. For example, in order to comply with the Transmission Control Protocol (TCP) or the User Datagram Protocol (UDP), software-controlled management may often be required. Software-controlled or software-assisted network traffic control management tends to increase the latency associated with the communication.

[0092] However, this limitation on latency may make Ethernet technology less suitable for applications such as high-performance computing (HPC) and artificial intelligence (AI) training data centers, where latency in the single microsecond range may be required to improve system performance and efficiency. Although protocols such as Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) or InfiniBand over Ethernet (IBoE) can help reduce latency, they may introduce greater system design complexity or cost. For example, RoCE or InfiniBand have lossless networking and scaling specifications (which may be difficult to implement). Implementing RoCE or InfiniBand may also result in significant software control overhead, or involve centralized token control mechanisms with limited bandwidth. In addition, systems implementing RoCE or InfiniBand may experience frequent pauses (e.g., frequent pauses).

[0093] To address at least part of the above problems, some embodiments of the present disclosure disclose a flow control protocol (e.g., Tesla Transport Protocol (TTP)) that can operate on an Ethernet-based network or a peer-to-peer (P2P) network. The flow control protocol can be implemented entirely in hardware without the help of a software control mechanism, thereby controlling communication delays to within single-digit microseconds. The flow control protocol can be implemented without involving software resources such as a general-purpose processor or central processing unit that executes computer-readable instructions or an operating system. In addition, since some mechanisms are built into the flow control protocol (e.g., a limit on the number of packets that can be sent before pausing, a limit on the number of links that can be established simultaneously, a hardware-controlled state machine, or one or more of the proposed header formats of packets sent or received according to the TTP), virtualized resources (e.g., a virtualized processor or memory) are not required to implement the flow control protocol.

[0094] In some embodiments, the state machine accelerates the transition between different states to open and close the communication link between the nodes. The state machine can be maintained and implemented by hardware without involving software, firmware, drivers or other types of programmable instructions. Therefore, compared with the implementation of other protocols supported by software (such as the Transmission Control Protocol (TCP) applicable to Ethernet-based networks), the transition between different states of the state machine can be accelerated.

[0095] In some embodiments, a header for packets sent and received in accordance with the TTP (e.g., a TTP header) supports operations from Layer 2 to Layer 4 of the Open Systems Interconnection (OSI) model. The header may include fields that are recognizable by existing Ethernet-based network devices or infrastructure. Thus, compatibility of the TTP with existing Ethernet standards may be preserved. Advantageously, this may allow economical use of existing infrastructure and / or supply chains, resulting in more system design options, and enabling system-level reuse or redundancy.

[0096] As described above, a node can be implemented or operated under a TTP (e.g., using a TTP to communicate with another node), using only hardware resources without the help of a software-controlled mechanism. In order to operate under a TPP of pure hardware resources, a node can use a hardware replay architecture to replay packets that may be lost in transmission. In some embodiments, the hardware replay architecture may include a local storage device, such as one or more caches for storing packets sent and / or received on one or more links, each of which can be opened or closed according to the TTP. Compared to protocols such as TCP or UDP (where virtualized resources with almost unlimited processing power and storage capacity can usually be obtained through software-controlled network traffic control management), the size of the cache (e.g., low-level cache) used by the hardware replay architecture in a node operating under a TTP can be limited. For example, the size of the cache can be megabytes (MB) or kilobytes (KB), such as 256KB. In order to communicate with each other through one or more links established according to a lossy communication protocol (such as TTP) with limited local storage, packets associated with the one or more links should be adequately managed (e.g., retained or discarded) so that some packets are retained for replay and other packets are discarded to avoid cache overflow.

[0097] In some examples, a first node that sends N packets to a second node using a link established under a TTP may utilize a cache to store the N packets, where N is any positive integer that may be subject to cache size constraints. As long as the constraints of the TTP and / or network conditions allow, the first node may continuously send some or all of the N packets to the second node. In order to accommodate the replayed packets, the cache may continue to store the packets that have been sent until a confirmation of the receipt of the packets is received from the second node. When the confirmation of the receipt of the packets is received, the cache may discard the packets to make room to store packets to be sent on the link or other links between the first node and the second node or other nodes. On the contrary, if a negative confirmation of the packet is received (e.g., the second node notifies the first node that the packet has not been received) or a timeout occurs without receiving a confirmation or negative confirmation of the receipt of the packet from the second node, the first node may replay the packet (e.g., retransmit the packet to the second node). In association with the replay of the packet, the first node may discard other packets for which a confirmation of the receipt has been received.

[0098] In some examples, the order in which packets are sent and replayed can be the same. For example, a first node can send N packets (e.g., a first packet, a second packet to an Nth packet) in a particular order. If the fifth packet is replayed (e.g., in response to the first node receiving a negative acknowledgment of the fifth packet from the second node, in response to a timeout occurring without receiving an acknowledgment of the fifth packet or an acknowledgment of its receipt), and acknowledgments for the first to fourth packets have been received, the cache can discard the first to fourth packets, but not the fifth packet, so that the node can replay the fifth packet. Additionally and / or optionally, when replaying the fifth packet, the first node can replay packets sent after the fifth packet in the same order as previously sent (assuming N>5).

[0099] In some examples, the hardware replay architecture of the first node can utilize a linked list coordinated with a cache to maintain the order between the first transmission of some or all of the N packets and any subsequent replays. The linked list may include N elements, each of which includes each of the N packets and a reference to the next element corresponding to the next packet. When sending and / or replaying N packets, the hardware replay architecture can also utilize one or more pointers pointing to one or more elements in the linked list to determine whether the packet is to be retained for replay or can be discarded (e.g., to save storage resources). Taking N as 9 (e.g., 9 packets sent from the first node to the second node) as an example, in the linked list, the first element may include the first packet and the first reference, wherein the first reference points to the second element; the second element may include the second packet and the second reference, wherein the second reference points to the third element; the eighth element may include the eighth packet and the eighth reference, wherein the eighth reference points to the ninth element; the ninth element may include the ninth packet. The hardware replay architecture can maintain and update three pointers to three elements. Assuming that the node has sent the first to ninth packets, and has received confirmation of the reception of the first to seventh packets from the second node but has not received confirmation of the reception of the eighth and ninth packets, the first pointer may point to the first element of the linked list, the second pointer may point to the eighth element of the linked list, and the third pointer may point to the ninth element of the linked list. Therefore, the hardware replay architecture may cause the cache to discard the packet and replay the packet based on three pointers. More specifically, the cache may replay the packet pointed to by the second pointer (e.g., the eighth packet) through the packet pointed to by the third pointer (e.g., the ninth packet), and discard the remaining packets (e.g., the packet pointed to by the first pointer before the packet pointed to by the second pointer). Additionally and optionally, some or all of the hardware replay architectures may operate in a pipelined manner to improve the throughput of the node. Using a cache and a linked list to implement the replay function may enable the first node to communicate with the second node using TTP under limited hardware resources without the help of a software-controlled mechanism.

[0100] As described above, a node operating under the TTP protocol may include a hardware link timer to implement a timeout check mechanism to replay packets without software assistance. Compared to other Ethernet protocols (e.g., TCP or UDP) where software is typically used to track timeouts on multiple links using multiple timers (e.g., a link's timer), a hardware link timer may allow a node to determine which (or which) packets sent through which (or which) links are to be replayed, and if replay is required, when to replay under limited hardware resources (e.g., when a large resource pool of virtual and / or physical address space and computing resources is not available). In some embodiments, a hardware link timer may periodically perform timing checks on established links (e.g., active links) associated with a node. A hardware link timer may include a first-in, first-out (FIFO) memory that may store timing and status information associated with each active link and check the timing and status associated with each active link in a polling manner. A hardware link timer may utilize a single programmable timer to schedule the time points of multiple active links and / or packets to read out the timing and status information associated with each of the multiple active links and / or packets. The read timing and status information can be used to determine, through further information lookup, whether to replay packets associated with the link or to discard them.

[0101] In some examples, the FIFO memory may store timing information associated with one or more links established between the first node and (multiple) other nodes. For example, the first node may include a hardware link timer that uses the FIFO memory to store timing information associated with M links established between the first node and one or more other nodes, where M is a positive integer greater than 1. The hardware link timer may utilize a single timer (e.g., a timer that ticks once within a programmable time period) to track and / or update the timing information of each link in the M links by accessing the FIFO memory in a polling (e.g., cyclic) manner, rather than using M timers (where each timer tracks the timing information of the corresponding link). Specifically, when a single timer ticks once, the hardware link timer may access the entries of the FIFO memory one at a time, where each access entry of the FIFO memory corresponds to one link in the M links. In some embodiments, the time period of each tick may vary and may be between hundreds of microseconds and single-digit microseconds. For example, the time period of a tick may be as high as 100 microseconds or as low as 1 microsecond. In addition, the hardware link timer can adjust the time period of the tick based on the number of links (e.g., M) represented by the entries of the FIFO memory. For example, when M increases (e.g., more links are represented by the entries of the FIFO memory), the time period of the tick can be shortened; and when M decreases (e.g., fewer links are represented by the entries of the FIFO memory), the time period of the tick can be increased. Therefore, if the time period of the tick changes disproportionately with the number of links represented by the entries of the FIFO memory, the time interval in which the status and / or timing information of the links is checked can remain unchanged.

[0102] In some examples, timing and / or status information associated with one of the M links may indicate how long the link has not received an acknowledgment of receipt of a transmitted packet. Assuming that a first node has sent N packets to a second node via a link, an entry of the FIFO memory may store timing and / or status information that, when accessed by polling within a specific time period of a tick, indicates that an acknowledgment of receipt of any of the N packets has not been received within a predetermined duration. When accessing an entry of the FIFO memory, the hardware link timer may utilize the timing and / or status information stored in the entry to look up the N packets that may be stored in a local storage device (e.g., a low-level cache) of the first node to replay the N packets. Alternatively, timing and / or status information associated with one of the M links may be stored in an entry of the FIFO memory to indicate that the link may be closed (e.g., all packets sent by the first node have been received by the second node). When accessing the entries of the FIFO memory, the hardware link timer can use the timing and / or state information stored in the entries to find packets that can still be stored in the local storage device of the first node, and discard these packets because the timing and / or state information stored in the entries of the FIFO memory indicates that the link can be shut down. Advantageously, by utilizing a single timer for multiple links and / or packets that ticks within an adjustable period and a FIFO memory that stores timing and / or state information for multiple links, the first node can replay packets at appropriate timing to achieve low latency and free up hardware resources occupied by inactive links (e.g., shut down links) for active links to operate with limited computing and storage resources.

[0103] Although various aspects will be described according to illustrative embodiments and feature combinations, it will be appreciated by those skilled in the relevant art that these examples and feature combinations are illustrative in nature and should not be construed as limitations. More specifically, aspects of the present application may be applicable to various types of networks and communication protocols under different contexts. In addition, although the specific architecture of a circuit block diagram or a state machine for controlling network flow will be described, such illustrative circuit block diagrams, state machines or architectures should not be construed as limitations. Therefore, it will be appreciated by those skilled in the relevant art that aspects of the present application are not necessarily limited to being applied to the illustrative interactions between any particular type of network, network infrastructure or network nodes.

[0104] Tesla Transport Protocol (TTP)

[0105] Figure 1A-1B is a table showing the OSI model (having seven layers) and example protocols associated with each layer. Figure 1A Example protocols of TCP and UDP protocols operating at layer 4 (eg, transport layer) of the OSI model are shown. Figure 1BAn example protocol of the Tesla Transport Protocol (TTP) operating at layer 4 of the OSI model is shown.

[0106] like Figure 1A As shown, in addition to TCP or UDP operating at layer 4, other example protocols or applications operating with TCP or UDP may include: Hypertext Transfer Protocol (HTTP), Telnet, File Transfer Protocol (FTP) operating at layer 7; Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Motion Picture Experts Group operating at layer 6; Network File System (NFS) and Structured Query Language (SQL) operating at layer 5; Internet Protocol Version 4 (IPv4) / Internet Protocol Version 6 (IPv6) operating at layer 3; etc. By operating TCP or UDP at layer 4, the implementation of layer 4 generally involves the following: Figure 1A Software shown.

[0107] like Figure 1B As shown, in addition to TTP operating at layer 4, other example protocols or applications operating with TTP may include: Pytorch operating at layer 7; FFMPEG, High Efficiency Video Coding (HEVC), YUV operating at layer 6; RDMA operating at layer 5; IPv4 / IPv6 operating at layer 3; and so on. Figure 1A In contrast, with a TTP operating at layer 4, the implementation of layers 1 to 4 of the OSI model can be performed solely in hardware without involving Figure 1B Advantageously, with Figure 1A Compared to the implementation shown, Figure 1B As shown, communication delays on Ethernet-based networks can be reduced through pure hardware implementation of Layers 1 to 4 of the OSI model based on TTP.

[0108] Figure 2 An example state machine 200 for opening and closing links between nodes implementing TTP according to an embodiment of the present disclosure is depicted. State machine 200 can be implemented by a network interface processor or a network interface card. There can be one state machine 200 for each Ethernet link between nodes on each node that communicates via an Ethernet link. For example, if a network interface processor can communicate with five network interface cards via five TTP links, the network interface processor can include five instances of state machine 200, one instance for each link. In this example, each of the five network interface cards can have an instance of state machine 200 for communicating with the network interface processor. In some embodiments, nodes that communicate with each other using state machine 200 can form a peer-to-peer network.

[0109] like Figure 2As shown, the state machine 200 includes a closed state 202, an open receive state 204, an open send state 206, an open state 208, a closed receive state 210, and a closed send state 212. The state machine 200 can start from the closed state 202, which can indicate that there is currently no communication link open between the first node maintaining the state machine 200 and the second node to establish a communication link with it. In addition, individual copies of the state machine 200 can be maintained, updated, and transformed by nodes operating based on the Tesla Transfer Protocol (TTP) disclosed in the present disclosure. In addition, if a node operating based on the TTP communicates with multiple nodes simultaneously or overlaps in time, the node can retain multiple independent state machines 200 for each link.

[0110] Then, the state machine 200 may transition differently depending on whether the first node sends or receives a request to establish a communication link to the second node. If the first node sends a request to open a communication link to the second node, the state machine 200 may transition from the closed state 202 to the open send state 206. On the other hand, if the first node receives a request to open a communication link from the second node, the state machine 200 may transition from the closed state 202 to the open receive state 204.

[0111] While in the open send state 206, the state machine 200 may remain in the open send state 206, or transition back to the closed state 202 or advance to the open state 208 according to various criteria. If the first node receives an open-nack (e.g., a message denying the request to open a link) from the second node, the state machine 200 may transition from the open send state 206 back to the closed state 202. On the other hand, if the first node receives an open-ack (a message accepting the request to open a link) from the second node, the state machine 200 may transition from the open send state 206 to the open state 208. Alternatively, if the first node does not receive an open-nack or an open-ack from the second node within a certain period of time, the first node may time out, and the first node may then retransmit the request to open a communication link to the second node and remain in the open send state 206.

[0112] As described above, when in the closed state 202, if the first node receives a request to open a communication link from the second node, the state machine 200 can be transformed from the closed state 202 to the open receiving state 204. In the open receiving state 204, the state machine 200 can be transformed differently according to whether the first node accepts or rejects the request to open the link from the second node. For example, the first node can choose to send an open-nack (e.g., reject the request to open the link) to the second node. In this case, the state machine 200 can be transformed back to the closed state 202, wherein the first node can also send or receive a request to open the link from the second node or other nodes. Alternatively, in the open receiving state 204, the first node can send an open-ack to the second node and then transition to the open state 208.

[0113] When in the open state 208, the first node and the second node can send and receive packets to each other through the established communication link. The link can be a wired Ethernet link. The first node can remain in the open state 208 until a certain situation occurs. In some embodiments, in response to receiving a request to close the communication link, the state machine 200 can transition from the open state 208 to the closed receive state 210, which allows the first node and the second node to send and receive packets when in the open state 208. Alternatively, in response to the first node sending a request to close the communication link to the second node, the state machine 200 can transition from the open state 208 to the closed send state 212. In addition to the request to close the communication link, if the communication link has been idle for more than a threshold amount of time, the state machine 200 can transition from the open state 208 to the closed receive state 210 or the closed send state 212.

[0114] While in the CloseReceive state 210, if the first node sends a close-ack (e.g., a message confirming or accepting the request to close the link) to the second node, the state machine 200 may transition back to the Close state 202. Otherwise, if the first node sends a close-nack (e.g., a message rejecting or not confirming the request to close the link) to the second node, the state machine 200 may remain in the CloseReceive state 210.

[0115] While in the close send state 212, if the first node receives a close-ack (e.g., a message confirming or accepting the request to close the link) from the second node, the state machine 200 may transition back to the close state 202. Otherwise, if the first node receives a close-nack (e.g., a message rejecting or not confirming the request to close the link) sent from the second node, the state machine 200 may remain in the close send state 212. In the close send state 212, if the first node does not receive a reply from the second node within a timeout threshold, the first node may resend the request to close the communication link to the second node.

[0116] In some embodiments, the state machine 200 can be maintained and implemented by hardware without involving software, firmware, drivers, or other types of programmable instructions. Thus, transitions between different states of the state machine 200 can be accelerated compared to implementations of other protocols that involve software support, such as the Transmission Control Protocol (TCP) applicable to Ethernet-based networks.

[0117] In some embodiments, the first node can immediately stop sending packets in the transmission queue and, when in the closed receive state 210, send a close-ack to the second node in response to a request to close the link from the second node, rather than letting the sent packets wait in the transmission queue for transmission and storage. Advantageously, avoiding continuing to send packets for an indefinite amount of time after receiving a request to close the link can enable the first node to transition back from the open state 208 to the closed state 202 with fewer transition periods and less time uncertainty.

[0118] In addition, the number of packets that the first node or the second node can continuously send during the open state 208 may be limited. For example, when in the open state 208, the first node can only successively send N packets before stopping packet transmission, where N can be a positive integer from 1 to over 1000. The number N can be limited by the physical memory. In some embodiments, N can be limited or constrained by the size of the physical memory available to the first node (e.g., dynamic random access memory, etc.). Specifically, N can be proportional to the size of the physical memory associated with the first node or the second node. For example, if 1 gigabyte (GB) of physical memory is allocated to the first node, N can be up to one million. In some embodiments, N can be within the tens of thousands or hundreds of thousands. During the open state 208, the amount of physical memory used for packet exchange can be tracked. Advantageously, limiting the number of packets that the first node or the second node can continuously send can reduce the computational and storage resources required to implement the state machine 200. Contrary to protocols (e.g., TCP) that typically assume unlimited software and hardware resources available through virtualization (e.g., virtualized memory or processing resources), limiting the number of sent packets allows the TTP to operate with more constrained computational and storage resources.

[0119] In some embodiments, the first node or the second node no longer waits to close the link after receiving or sending a close-ack to the other node. For example, when in the close send state 212, in response to receiving a close-ack sent from the second node, the first node can immediately transition to the close state 202. The first node can transition back to the close state 202 from the close send state 212 in a shorter time, rather than waiting for another predetermined or random time period to monitor whether the second node has additional packets to send. Advantageously, this improves accuracy and shortens the delay associated with transitions between states of the state machine 200, thereby allowing TTP to facilitate communication with lower delays than protocols such as TCP.

[0120] Figure 3A-3B An example timing diagram describing the transmission and reception of packets between two devices implementing a TTP according to an embodiment of the present disclosure is illustrated. Figure 3A The diagram shows the situation where no packets are lost from device A to device B. Figure 3B Another situation is illustrated where some of the transmitted packets from device A to device B are lost. Figure 3A-3B This can be understood in conjunction with state machine 200. Device A and device B are two example nodes communicating via TTP.

[0121] like Figure 3A As shown, device A can send a TTP_OPEN with packet ID=0 to device B while in the closed state 202. After sending the TTP_OPEN to device B at (1), the state machine maintained by device A can transition from the closed state 202 to the open send state 206. In addition, after receiving the TTP_OPEN from device A at (1), the state machine maintained by device B can transition from the closed state 202 to the open receive state 204.

[0122] Then, after receiving the TTP_OPEN_ACK from device B at (2), the state machine maintained by device A may transition from the open send state 206 to the open state 208. Furthermore, after sending the TTP_OPEN_ACK to device A at (2), the state machine maintained by device B may transition from the open receive state 204 to the open state 208.

[0123] At (3), while in the open state 208, device A may continuously or successively send four packets (e.g., TTP_PAYLOAD ID=1 to 4) to device B before receiving any response from device B. In some embodiments, the number of packets that device A may send to device B before receiving any response from device B is limited. In response to the packets received from device A, at (4), device B may send four packets (e.g., TTP_ACK ID=1 to 4) to acknowledge receipt of the four packets sent by device A.

[0124] At (5), device A sends a TTP_CLOSE (Packet ID=5) to device B. After sending the TTP_CLOSE, the state machine maintained by device A may transition from the Open state 208 to the Close Send state 212. In response to receiving the TTP_CLOSE from device A, the state machine maintained by device B may transition from the Open state 208 to the Close Receive state 210.

[0125] Thereafter, at (6), device B may send a TTP_CLOSE_ACK (packet ID=5) to device A. After sending the TTP_CLOSE_ACK to device A, the state machine maintained by device B may transition from the close receive state 210 back to the close state 202. After receiving the TTP_CLOSE_ACK from device B, the state machine maintained by device A may transition from the close send state 212 back to the close state 202. Thus, the link / connection between device A and device B may be close.

[0126] Figure 3B A "lossy" flow control feature associated with a flow control protocol (eg, TTP) disclosed in the present disclosure is illustrated, where lossy may indicate retransmission of lost or damaged packets after receiving a negative acknowledgement.

[0127] like Figure 3B As shown, device A can send a TTP_OPEN with packet ID=0 to device B while in the closed state 202. After sending the TTP_OPEN to device B at (1), the state machine maintained by device A can transition from the closed state 202 to the open send state 206. In addition, after receiving the TTP_OPEN from device A at (1), the state machine maintained by device B can transition from the closed state 202 to the open receive state 204.

[0128] Then, after receiving the TTP_OPEN_ACK from device B at (2), the state machine maintained by device A may transition from the open send state 206 to the open state 208. Furthermore, after sending the TTP_OPEN_ACK to device A at (2), the state machine maintained by device B may transition from the open receive state 204 to the open state 208.

[0129] At (3), when in the open state 208, device A may continuously or successively send four packets (e.g., TTP_PAYLOAD ID=1 to 4) to device B before receiving any response from device B. However, due to certain network conditions, device B may not be able to receive some packets (e.g., TTP_PAYLOAD ID=3). Therefore, at (4), device B may send three packets (e.g., TTP_ACK ID=1 to 2, TTP_NACK ID=3) to confirm the receipt of the two packets sent by device A (ID=1 to 2), but to notify that the packet of TTP_PAYLOAD ID=3 was not received.

[0130] After receiving a packet from device B (e.g., TTP_NACK ID=3), at (5), device A retransmits two packets to device B (e.g., TTP_PAYLOAD ID=3 to 4). It is worth noting that retransmitting two packets after receiving a packet (e.g., TTP_NACK ID=3) may reflect the "lossy" feature of TTP. In some embodiments, device A may retransmit some packets after a timeout occurs (e.g., when a local counter exceeds a certain value). Advantageously, since there is a peer link between device A and device B, the "lossy" feature enables TTP to control or scale network traffic without restriction, and enables TTP to implement link-specific recovery in large systems where some traffic is expected to be lost.

[0131] At (6), after receiving two packets (e.g., TTP_PAYLOAD ID=3 to 4), device B can send two packets (e.g., TTP_ACK ID=3 to 4) to device A to acknowledge receipt of the retransmitted packets (e.g., TTP_PAYLOAD ID=3 to 4).

[0132] At (7), device A may send a packet (e.g., TTP_CLOSE ID=5) to device B to attempt to close the link between device A and device B. In addition, at (7), the state machine maintained by device A may transition from the open state 208 to the closed send state 212, and the state machine maintained by device B may transition from the open state 208 to the closed receive state 210.

[0133] At (8), Device B may send a packet (e.g., TTP_CLOSE_ACK ID = 5) to Device A to confirm and agree to close the link. The state machine maintained by Device B may transition back from the closed receive state 210 to the closed state 202. In response to receiving a packet (e.g., TTP_CLOSE_ACK ID = 5) from Device B, the state machine maintained by Device A may transition back from the closed transmit state 212 to the closed state 202.

[0134] In some embodiments, Device A and / or Device B may not transition to the open state 208 or may not send or receive packets until the process of negotiating the link is complete. For example, Device A may not send a packet to Device B or receive a packet from Device B until Device A receives a TTP_OPEN_ACK from Device B. In these embodiments, when the link between Device A and Device B is closed, especially when TTP_OPEN is immediately sent from Device A or Device B after a previous link closure between Device A and B, it may not be necessary to impose a timeout.

[0135] Figure 4 FIG. illustrates an example block diagram of a node 400 implementing TTP according to an embodiment of the present disclosure. As Figure 4 shown, the node 400 may include a transmit (TX) path and a receive (RX) path. As Figure 4 shown, at the front end of the node 400 includes a physical coding sublayer (PCS) + physical medium attachment (PMA) block 402, which processes communications on the first layer (e.g., physical layer) of the OSI model. In some embodiments, the PCS + PMA block 402 operates based on a reference clock 404 with a frequency of 156.25 MHz. In other embodiments, the PCS + PMA block 402 may operate at different clock frequencies. The PCS + PMA block 402 may be compatible with Ethernet or the IEEE802.3 standard. In the operation of processing data on the RX path, the PCS + PMA block 402 receives the RX sequence [3:0] as an input and rearranges the RX sequence [3:0] into an output (e.g., RX frame 408) for processing by the TTP media access control (MAC) block 410. In the operation for processing data on the TX path, the PCS + PMA block 402 receives the TX frame 412 as an input from the TTP MAC block 410 and rearranges the data format to output the TX sequence [3:0].

[0136] On the RX path, the TTP MAC block 410 receives RX frames 408 as input and outputs RDMA receive data 416 to a system on chip (SoC) 420. On the TX path, the TTP MAC block 410 receives RDMA transmit data 418 from the SoC 420 and outputs TX frames 412 to the PCS+PMA block 402. Figure 4 As shown, the TTP MAC block 410 can handle operations on layers 2 to 4 of the OSI model. The TTP MAC block 410 can include a TTP finite state machine (FSM) 422. The TTP FSM 422 can maintain and update the following information: Figure 2 State machine 200 is shown. As described above, for each communication link established between node 400 and one or more other nodes, TTP FSM 422 may maintain and update a corresponding state machine (eg, state machine 200) to control traffic associated with the corresponding communication link.

[0137] In some embodiments, the PCS+PMA block 402 and the TTP MAC block 410 may be implemented in hardware, such as in the form of an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Thus, the PCS+PMA block 402 and the TTP MAC block 410 may operate without the assistance or involvement of software / firmware / drivers. Advantageously, the PCS+PMA block 402 and the TTP MAC block 410 may handle communications from layer 1 to layer 4 of the OSI model without software assistance to reduce delays associated with communications in layer 1 to layer 4.

[0138] Figure 5 An example header 500 of a packet sent or received according to a TTP is depicted. Figure 5 As shown, the example header 500 has 64 bytes. The first 16 bytes include a header for Ethernet Layer 2 (e.g., data link layer) and virtual local area network (VLAN) operations. The second 16 bytes include ETHTYPE, followed by an optional Layer 3 Internet Protocol (IP) header. To support Layer 2 operations based on TTP, ETHTYPE can be set to a specific value (e.g., 0x9AC6). When ETHTYPE is set to a specific value, header 500 can signal a network device that processes header 500 that header 500 is formatted based on TTP. The third 16 bytes include optional fields for Layer 3 (IP) operations and Layer 4 operations under UDP. At the end of the third 16 bytes and the fourth 16 bytes is a field for Layer 4 operations under TTP. TTP may refer to TTP over Ethernet (TTPoE). TTP in Figure 5 It is marked as TTPoE.

[0139] Advantageously, the example header 500 allows the TTP to support operations from at least Layer 2 to Layer 4 of the OSI model on Ethernet-based networks. In particular, existing Ethernet switches and hardware can support the operations associated with the TTP.

[0140] Figure 6 An example network and computing environment 600 is illustrated in which embodiments of the present disclosure may be implemented. The example network and computing environment 600 may be used in a high performance computing or artificial intelligence training data center. As an example, the network and computing environment 600 may be used for neural network training to generate data for use by an autonomous driving system of a vehicle (e.g., a car). Figure 6 As shown, the example network and computing environment 600 includes an Ethernet switch 608, hosts 602A to 602E, Peripheral Component Interconnect Express (PCIe) hosts 604A to 604N, and computing blocks 606A to 606N. Figure 6 There are five hosts 602A to 602E in FIG. 602A, but any suitable number of hosts greater than or less than five may be implemented. In addition, the number of PCIe hosts and the number of computing blocks may be any suitable positive integers.

[0141] Each of the hosts 602A to 602E includes a network interface card (NIC), a central processing unit (CPU), and a dynamic random access memory (DRAM). Although illustrated as a CPU, in some embodiments, the CPU may be embodied as any type of single-core, single-threaded, multi-core or multi-threaded processor, microprocessor, digital signal processor (DSP), microcontroller, or other processor or processing / control circuit. Although illustrated as DRAM, in some embodiments, the DRAM may alternatively or additionally be embodied as any type of volatile or non-volatile memory or data storage device, such as static random access memory (SRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM). The DRAM may store various data and program codes used during the operation of the hosts 602A to 602E, including operating systems, applications, libraries, drivers, etc.

[0142] In some embodiments, the NIC may implement TTP for communicating with the Ethernet switch 608. Each NIC may communicate with the Ethernet switch 608 using TTP as a flow control protocol to manage the link established between each NIC and a network interface processor (NIP) via the Ethernet switch 608. In some embodiments, the NIC may include Figure 4 The PCS+PMA block 402 and the TTP MAC block 410. In some embodiments, the NIC may implement TTP without the assistance of software / firmware.

[0143] like Figure 6 As shown, each of the PCIe hosts 604A to 604N may include a network interface processor (NIP) and a high bandwidth memory (HBM). In some embodiments, the bandwidth supported by the HBM may be 32 gigabytes (GB) per calculation. Each of the PCIe hosts 604A to 604N may communicate with each of the computing blocks 606A to 606N. Each of the computing blocks 606A to 606N may include storage, input / output, and computing resources. The computing block 606A may include a system on a chip having a processor array for high-performance computing. In some applications, each of the computing blocks 606A to 606N may perform 9 peta floating point operations per second (PFLOPS), use a static random access memory (SRAM) to store data of 11 gigabytes (GB), or facilitate input / output operations with a bandwidth of 36 terabytes (TB) per second.

[0144] In some embodiments, each NIC in the hosts 602A to 602E can open and close the communication link with each NIP in the PCIe hosts 604A to 604N. Specifically, a NIC and a NIP can be implemented by Figure 2 The state machine 200 of NIC and NIP can be used to open and close the communication link between each other. Figure 7A-7B For example, to open a link with the NIP, the NIC may send a packet including the operation code TTP_OPEN (such as Fig. 7A After receiving a packet with the operation code TTP_OPEN, the NIP can Figure 2 The closed state 202 is changed to the open receiving state 204. When sending a packet with the operation code TTP_OPEN_ACK (such as Fig. 7A After that, NIP can Figure 2 The open receive state 204 shown transitions to the open state 208. In some embodiments, after the communication link is established (e.g., when both the NIC and the NIP are in the open state 208), the NIC and the NIP may use Figure 5 In other words, each packet sent or received between the NIC and the NIP may include Figure 5 The header is 500.

[0145] like Figure 6As shown, communication and data exchange between each of hosts 602A to 602E, each of PCIe hosts 604A to 604N, each of compute blocks 606A to 606N, or Ethernet switch 608 can be based on TTP. By using the above techniques, a shorter latency (compared to TCP) is achieved through TTP, enabling Figure 6 high-bandwidth and high-speed communication between various components. In some embodiments, Figure 6 at least a portion of the NIP or at least a portion of the NIC shown can be implemented similarly or identically to Figure 4 node 400. Although Figure 6 not shown in

[0146] Figure 7A-7B In some embodiments, each of the NIC and NIP can include a port 610 through which packets can be received and sent. In some embodiments, port 610 is an Ethernet port. Fig. 7A and Figure 7B The TTP packets shown in Figure 2 and Figure 3A and Figure 3B are used to close and open links between nodes for a network. TTP packets can be exchanged between nodes in a Figure 6 network and computing environment. Fig. 7A and Figure 7B The TTP packets shown can be combined with Figure 2 and Figure 3A and Figure 3B for better understanding.

[0147] Packet replay hardware architecture

[0148] Referring again to Figure 4, which illustrates an example block diagram of a node 400 that sends and / or receives packets using TTP, a replay hardware architecture will be described. As described above, the node 400 may include blocks such as a physical coding sublayer (PCS) + physical medium attachment (PMA) block 402 and a TTP media access control (MAC) block 410, which includes a TTP FSM 422 for handling communications from layer 1 to layer 4 of the OSI model without software assistance to reduce delays associated with communications in layer 1 to layer 4. In addition, the TTP media access control (MAC) block 410 of the node 400 may include a hardware replay architecture that includes at least a TTP (peer link) tag block 436, an RX data path 432, an RX storage device 432-1 (e.g., on-die SRAM), a TX data path 434, and a TX storage device 434-1 (e.g., on-die SRAM). The hardware replay architecture can replay packets lost during transmission under a lossy protocol (such as TTP). Optionally, the TTP media access control (MAC) block 410 of the node 400 may also include a TTP MAC RDMA address encoding block 438 that may receive the RDMA transmit data 418 from the system on chip (SoC) 420 and encode it.

[0149] In some embodiments, the hardware replay architecture of the node 400 for replaying packets may include at least a TTP tag block 436, an RX data path 432, an RX storage device 432-1, a TX storage device 434-1, and a circuit system of the TX data path 434. As described above, the hardware replay architecture can utilize physical storage devices and data structures to store packets sent and / or received in different links, and maintain the order of the sent packets, particularly when replay occurs. In some embodiments, the physical storage device used by the hardware replay architecture can be any suitable type of local storage device or cache (e.g., a low-level cache) that can store, buffer, and / or save packets associated with one or more links. The size of the physical storage device can be limited, such as a size of megabytes (MB) or kilobytes (KB). In some examples, the physical storage device can be deployed as part of the TX data path 434, or more specifically, as part of the TX storage device 434-1. The physical storage device can also be deployed as part of the RX data path 432, or more specifically, as part of the RX storage device 432-1. For example, the physical storage devices may be RX storage devices 432-1 and TX storage devices 434-1, where the size of RX storage devices 432-1 and TX storage devices 434-1 used by the hardware replay architecture associated with each of the RX data path 432 and the TX data path 434 may be 256KB. In other examples, the physical storage devices may be deployed within or as part of the TTP tag block 436 (e.g., as local storage devices deployed within the TTP tag block 436).

[0150] It should be noted that the hardware replay architecture within the TTP media access control (MAC) block 410 of the node 400 may employ any other suitable size of physical storage. In some embodiments, the data structures used by the hardware replay architecture (e.g., within the TTP tag block 436) may include one or more linked lists, wherein each linked list may record and / or track the order of packets sent for a corresponding link established between a first communication node and a second communication node. In some embodiments, the TTP tag block 436 may utilize linked lists and physical storage devices (e.g., RX storage device 432-1 and TX storage device 434-1) to maintain and manage the stored packets to replay packets sent over multiple links.

[0151] Figure 8 and Fig. 9 FIG. 1 illustrates a node (eg, node 400 or node 400) in an Ethernet-based network according to some embodiments of the present disclosure. Figure 3B Example physical storage devices and data structures (e.g., TX linked list 952) used by device A) of the network implementing a TTP for replaying or retransmitting packets. Figure 8 and Fig. 9 Can be combined Figure 3B To understand, Figure 3B It is shown that in response to receiving a negative acknowledgement packet (eg, TTP_NACK ID=3) notifying that a packet (TTP_PAYLOAD ID=3) was not received, device A replays two packets (TTP_PAYLOAD ID=3 to 4).

[0152] refer to Figure 8 , Figure 3B Device A may store packet 1 (eg, packet physical cache 802) for transmission and / or playback. Figure 3B Packet TTP_PAYLOAD ID = 1), Packet 2 (for example, Figure 3B Group TTP_PAYLOAD ID = 2), Group 3 (for example, Figure 3B Packet TTP_PAYLOAD ID = 3), Packet 4 (for example, Figure 3B Packet TTP_PAYLOAD ID = 4) and Packet 5 (for example, Figure 3B As described above, the packet physical cache 802 may be the TX storage device 434-1 and / or may be a physical storage device deployed within the TTP tag block 436. In some embodiments, the packet physical cache 802 may have two storage spaces—packet physical tags 804 and packet physical data 806. For each packet (e.g., packets 1 to 5), the packet physical tag 804 may include a physical address pointer that points to a physical address in the packet physical data where the packet is stored. For example, the physical address pointer 808 associated with packet 4 stored in the entry of the packet physical tag 804 may point to a physical address where packet 4 (e.g., Figure 3B The entry of the packet physical data 806 of the packet TTP_PAYLOAD ID=4). Figure 8 As shown, device A may send packet 1, packet 2, packet 3, packet 4, and packet 5 in order 820 (e.g., packet 1 is sent first and packet 5 is sent last). However, device A may not store packets 1 to 5 in packet physical data 806 based on order 820. Specifically, although device A sends packet 3 before packets 4 and 5, address 810 in packet physical data 806 storing packet 3 may be located after addresses 812 and 814 in packet physical data 806 storing packets 4 and 5, respectively.

[0153] Fig. 9 TX linked list 952 is shown, which may be provided by node 400 and / or Figure 3BDevice A is used to maintain the packet transmission order between the previous transmission and the replay. TX linked list 952 can be part of TTP tag block 436 of node 400. As discussed above Figure 8 As stated, Figure 3B Device A can store packets 1 to 5 at various addresses of packet physical data 806, which do not reflect the transmission order 820 of packets 1 to 5. Nevertheless, device A can use TX linked table 952 to track and maintain the transmission order from packets 1 to 5. Fig. 9 As shown, the TX linked list 952 includes five elements 960 , 962 , 964 , 968 , 970 , each of which corresponds to or is associated with one of packets 1 to 5 . Fig. 9 The TX linked table 952 is illustrated to track and maintain the transmission order 820 from packet 1 to packet 5. For example, in the TX linked table 952, the element 964 corresponding to packet 3 is located before the element 968 corresponding to packet 4, and the element 968 corresponding to packet 4 is located before the element 970 corresponding to packet 5. Therefore, by utilizing the TX linked table 952, device A can maintain the packet transmission order during the previous transmission and replay, where the replay can be triggered in response to receiving the TTP_NACK ID=3 packet, which notifies Figure 3B Device B in does not receive the packet with TTP_PAYLOAD=3. According to any suitable principles and advantages disclosed herein, replay may be triggered in response to a timeout or a negative acknowledgement.

[0154] like Fig. 9 As shown, Figure 3B Device A may also use one or more pointers 972, 974, and 976 stored in memory to determine which packet(s) to replay. Figure 3BAs shown in (3) and (4) of FIG. 1 , device A sends four packets (e.g., TTP_PAYLOAD ID=1 to 4) and receives three packets (e.g., TTP_ACK ID=1 to 2, TTP_NACK ID=3) to confirm the reception of two packets (ID=1 to 2) sent by device A, but to notify that the packet of TTP_PAYLOAD ID=3 was not received. In response, device A can set pointer 972 to point to element 964 corresponding to packet 3 to indicate that device A is to replay packets starting from packet 3. Device A can also set pointer 974 to point to element 968 corresponding to packet 4 to indicate that device A will replay packet 4 in addition to packet 5. The device can also set pointer 976 to point to element 970 corresponding to packet 5 to indicate that device A can send packet 5 after replaying packets 3 and 4. In addition, device A can set element 960 and element 962 of TX linked list 952 to empty to indicate that packet 1 and packet 2 can be received from the address ( Figure 8 (not shown) is removed, thereby freeing up more storage space to store packets sent or received by device A.

[0155] Thereafter, based on TX linked list 952, pointer 972, and pointer 974, device A can replay packets 3 and 4, as shown in FIG. Figure 3B Then, device A can receive confirmation of the receipt of packets 3 and 4, as shown in (5). Figure 3B In response, based on the TX linked list 952, device A may send a packet 5 corresponding to element 970 of the TX linked list 952 (e.g., Figure 3B 5) to complete the transmission and replay of packets 1 to 5. Additionally and / or optionally, after all packets corresponding to elements of TX linked list 952 have been sent and replayed, device A may release the storage occupied by packets 1 to 5. In some embodiments, device A may indicate that the addresses in packet physical tag 804 and the addresses in packet physical data 806 have been released and are free to be used in conjunction with other linked lists corresponding to other packets by setting free list entry 832 and free list entry 834, respectively, to specific values.

[0156] Fig.10 The present invention illustrates some embodiments of the present invention. Figure 4 FIG. 4 is an example block diagram of a TTP tag block 436 of FIG. 4 , wherein the TTP tag block 436 is part of a hardware replay architecture for replaying packets sent over multiple links. Fig.10As shown, the TTP tag block 436 may include a memory storing a TX linked list 1020, and logic circuit systems 1012, 1014, 1016, and 1018 operating in pipeline stages 1002, 1004, 1006, and 1008, respectively. The logic circuit systems 1012, 1014, 1016, and 1018 may be implemented by any suitable physical circuit system. In some examples, some or all of the logic circuit systems 1012, 1014, 1016, and 1018 may be implemented by a dedicated circuit system, such as in the form of an application-specific integrated circuit (ASIC). In some examples, some or all of the logic circuit systems 1012, 1014, 1016, and 1018 may be implemented by programmable logic gates or general-purpose processing circuit systems, such as in the form of a field programmable gate array (FPGA) or a digital signal processor (DSP). In operation, the TX linked list 1020 may be similar to Fig. 9 In some embodiments, TX linked list 1020 tracks the order of N packets including packet 1022, packet 1024, and packet 1026, wherein node 400 may send the N packets tracked by TX linked list 1020 over a particular link. TTP tag block 436 also includes pointers 1032, 1034, and 1036 pointing to packet 1022, packet 1024, and packet 1026, respectively. TTP tag block 436 may store pointers 1032, 1034, and 1036 in any suitable storage element ( Fig.10 In some applications, the N packets including the packets 1022, 1024, and 1026 of the TX linked list 1020 may be stored in a physical storage device, such as the TX storage device 434-1 of the TX data path 434 of the node 400. In such applications, the TX linked list 1020 may include pointers to the packets 1022, 1024, 1026. In other applications, the N packets including the packets 1022, 1024, and 1026 may be part of the TX linked list 1020 stored in a physical storage device within the TTP tag block 436.

[0157] In some embodiments, node 400 may store N packets (including packet 1022, packet 1024, and packet 1026) sent to a second node using a link established under the TTP in TX storage device 434-1 (or other physical storage device of node 400), where N is any positive integer that may be limited by the size of TX storage device 434-1. Node 400 may continuously send some or all of the N packets to the second node as long as the constraints of the TTP and / or network conditions permit. In order to accommodate the replay of N packets including packet 1022, packet 1024, and packet 1026, TX storage device 434-1 may continue to store one or more packets (e.g., packet 1022) that have been sent until a confirmation of receipt of one or more packets is received from the second node. Packets may be stored until receipt of previously sent packets is confirmed. When a confirmation of receipt of a packet is received, TX storage device 434-1 may discard the packet to make room to store packets to be sent on the link or other links between node 400 and the second node and / or one or more other nodes. Conversely, if a negative acknowledgement of the packet is received (e.g., the second node notifies node 400 that the packet was not received) or a timeout occurs without receiving an acknowledgement or negative acknowledgement of receipt of the packet from the second node, node 400 may replay the packet still stored in TX storage device 434-1 (e.g., retransmit the packet to the second node). In association with replaying the packet, node 400 may discard other packets for which an acknowledgement of receipt has been received. In some embodiments, TX linked list 1020 may coordinate with TX storage device 434-1 to maintain the order between previous transmissions and any subsequent replays of some or all of the N packets, including packet 1022, packet 1024, and packet 1026. As Fig.10 As shown, the TX linked list 1020 includes N elements, each of which corresponds to or includes each of the N packets and a reference to a next element corresponding to the next packet.

[0158] When N packets are transmitted and / or replayed, the TTP tag block 436 may also use pointers 1032, 1034, and 1036, which respectively point to three elements in the TX linked list 1020, to determine whether the packet is to be retained for replay or can be discarded by the TX storage device 434-1 to save storage resources. Taking N as 9 (e.g., 9 packets transmitted from the node 400 to the second node) as an example, in the TX linked list 1020, the first element corresponds to the first packet (e.g., packet 1022) and the first reference, wherein the first reference points to the second element; the second element corresponds to the second packet and the second reference, wherein the second reference points to the third element; the eighth element corresponds to the eighth packet (e.g., packet 1024) and the eighth reference, wherein the eighth reference points to the ninth element; and the ninth element corresponds to the ninth packet (e.g., packet 1026). The TTP tag block 436 may maintain and update three pointers 1032 , 1034 , and 1036 , which point to the first element (eg, packet 1022 ), the eighth element (eg, packet 1024 ), and the ninth element (eg, packet 1026 ), respectively.

[0159] Assuming further that the node 400 has sent the first to ninth packets and has received confirmation of the reception of the first to seventh packets from the second node but has not received confirmation of the reception of the eighth and ninth packets, the pointer 1032 then points to the first element of the TX linked list 1020 (e.g., packet 1022), the pointer 1034 then points to the eighth element of the TX linked list 1020 (e.g., packet 1024), and the pointer 1036 then points to the ninth element of the TX linked list 1020 (e.g., packet 1026). Therefore, the TTP tag block 436 can cause the TX storage device 434-1 to discard some or all of the N packets including the packets 1022, 1024, and 1026, and replay some or all of the N packets based on the pointers 1032, 1034, and 1036. More specifically, the TX storage device 434-1 may replay the packet 1024 pointed to by the pointer 1034 through the packet 1026 pointed to by the pointer 1036 (in this case, only the packet 1024 and the packet 1026 are replayed). The TX storage device 434-1 may also discard the remaining packets (e.g., the packet 1022 pointed to by the pointer 1032 and other packets previously transmitted before the packet 1024; in this case, seven packets including the packet 1022 may be discarded).

[0160] like Fig.10As shown, some or all of TTP tag block 436 (e.g., logic circuitry 1012, 1014, 1016, and 1018) may operate in a pipelined manner to improve throughput of node 400. Logic circuitry 1012, 1014, 1016, and 1018 may operate in conjunction with TX linked list 1020 to determine whether a packet should be replayed or discarded / withdrawn from TX storage device 434-1 or other physical storage device of node 400 storing the packet. Fig.10 As shown, logic circuit systems 1012, 1014, 1016, and 1018 can operate at corresponding pipeline stages according to the clock operated by TTP tag block 436. Specifically, logic circuit system 1012 operates at initial pipeline stage 1002 (labeled as "Q0"), and logic circuit system 1014 operates at first pipeline stage 1004 (labeled as "Q1"). Logic circuit system 1016 operates at second pipeline stage 1006 (labeled as "Q2"), and logic circuit system 1018 operates at third pipeline stage 1008 (labeled as "Q3").

[0161] In operation, the logic circuit system 1012 can select one of the data streams to be processed in the TTP link label pipeline. As shown in the initial pipeline stage 1002, the logic circuit system 1012 can select one of the transmission stream ("TX queue"), the reception stream ("RX queue"), or the confirmation stream ("ACK queue") to be processed in the TTP link label pipeline based on a control signal (e.g., "select"). In the TTP link label pipeline, the logic circuit system determines whether to replay one or more packets of the selected data stream or to withdraw one or more packets of the selected data stream. The TTP link label pipeline can also determine to reject the confirmation of a packet sent after another packet determined by the TTP label pipeline to be replayed.

[0162] Assuming that logic circuitry 1012 selects to send a stream in preparation for replaying packets, then at first pipeline stage 1004, logic circuitry 1014 determines which link to evaluate for replay. This may involve reading a tag associated with the link. Fig.10 As shown, logic circuit system 1014 can select one of two links (e.g., "MOOSE" and "CAT") for possible replay, where each link can be established between the same endpoint or different endpoints. For example, both link "MOOSE" and "CAT" can be established between node 400 and a second node; alternatively, link "MOOSE" can be established between node 400 and the second node, while link "CAT" can be established between node 400 and a third node. Logic circuit system 1014 can select a link (e.g., "CAT") for replay based on a link pointer that points to the selected link.

[0163] Then, at the second pipeline stage 1006, the logic circuit system 1016 may determine which packet(s) sent on the link “CAT” are to be replayed or withdrawn. In some embodiments, the logic circuit system 1016 determines that some packets sent over the link “CAT” are to be replayed, while other packets may be withdrawn based on whether an acknowledgement or negative acknowledgement of receipt is received. For example, if a negative acknowledgement of receipt of packet 1024 is received or an acknowledgement of receipt of packet 1024 is not received within a time period that triggers a timeout, the logic circuit system 1016 may determine that packet 1024 is to be replayed. Conversely, in response to the receipt of an acknowledgement of packet 1022, the logic circuit system 1016 may determine that packet 1022 is to be withdrawn. Additionally and / or alternatively, the logic circuit system 1016 may also determine that other packets sent over the link “CAT” are to be replayed and / or withdrawn based on the TX linked list 1020. For example, based on the order of packets transmitted over link "CAT" specified by TX linked table 1020 (which shows packet 1026 transmitted after packet 1024), in response to receiving a negative acknowledgment of packet 1024, logic circuitry 1016 may determine that packet 1026 is to be replayed as well as replaying packet 1024. Logic circuitry 1016 may also cause TX memory device 434-1 to withdraw packets transmitted between packet 1022 and packet 1024, assuming that acknowledgments for packets transmitted between packet 1022 and packet 1024 have been received, to make more available storage space in TX memory device 434-1. In second pipeline stage 1006, acknowledgments for packets may be rejected in association with a determination that an earlier transmitted packet is to be replayed. Withdrawing packets may involve allowing other data to be written to memory in place of the packets and / or deleting packets from memory.

[0164] Thereafter, in the third pipeline stage 1008, the logic circuit system 1018 may update the link pointer pointing to the link "CAT" to point to another link (e.g., link "MOOSE"). Therefore, in the next round of pipeline operation, the logic circuit systems 1012, 1014, 1016, and 1018 may update the link pointer pointing to the link "CAT" to point to another link (e.g., link "MOOSE"). Fig.10 434-1) to determine whether to replay the packet(s) associated with the link "MOOSE", the TX linked list including, referencing or corresponding to the packet sent via the link "MOOSE". Advantageously, the use of TX storage 434-1 and TX linked list 1020 to implement the replay function enables the node 400 to communicate with the second node using TTP under limited hardware resources without the assistance of a software-controlled mechanism.

[0165] Hardware link timer

[0166] Fig.11An example block diagram of a hardware link timer 1100 is shown that implements a timeout check mechanism for replaying packets without software assistance. In some embodiments, the hardware link timer 1100 may be Figure 4 Some or all of the hardware link timer 1100 may be deployed in Figure 4 436. As described above, in contrast to other Ethernet protocols (e.g., TCP or UDP) where software is typically used to track timeouts on multiple links via multiple timers (e.g., one timer per link), the hardware link timer 1100 can allow the node 400 to determine which packet(s) sent on which link(s) to replay, and if replay is required, when to replay under limited hardware resources (e.g., when large resource pools of virtual and / or physical address space and computing resources are not available). In some embodiments, the hardware link timer 1100 can periodically perform timing checks on established links (e.g., active links) used by the node 400 to communicate with one or more other nodes according to the TTP.

[0167] like Fig.11 As shown, the hardware link timer 1100 may include a first-in-first-out (FIFO) memory 1104, a timer 1102, and logic circuit systems 1120, 1112, 1114, 1116, and 1118, wherein the logic circuit systems 1112, 1114, 1116, and 1116 may be part of the TTP tag block 436 for replaying packets. The FIFO memory 1104 may store timing and status information associated with each active link. The hardware link timer 1100 may check the timing and status associated with each active link stored in the FIFO memory 1104 in a polling manner. More specifically, the hardware link timer 1100 may start checking the timing and status information associated with the first link stored in the first entry of the FIFO memory 1104, and the timing and status information associated with the Nth link stored in the Nth entry of the FIFO memory 1104, and then check the timing and status information associated with the first link stored in the first entry of the FIFO memory 1104 again. The hardware link timer 1100 can schedule a time point using the timer 1102 to read out timing and status information associated with multiple active links and / or packets. The read out timing and status information can be used to determine whether to replay packets associated with the link or to withdraw and / or discard these packets through further information lookup. It should be noted that Figure 4 The node 400 may include Fig.11 Similar multiple hardware link timers are shown, where each hardware link timer may be capable of determining whether there is a timeout associated with multiple links.

[0168] In some embodiments, the FIFO memory 1104 may store timing information associated with one or more links established between the node 400 and (multiple) other nodes. For example, the node 400 may include a hardware link timer 1100 that uses the FIFO memory 1104 to store timing information associated with M links established between the node 400 and one or more other nodes, where M is a positive integer greater than 1. The hardware link timer 1100 may utilize a timer 1102 (e.g., a hardware clock that ticks once within a programmable time period) to track and / or update the timing information of each of the M links by accessing the FIFO memory 1104 in a polling (e.g., cyclic) manner, rather than using M timers (where each timer tracks the timing information of the corresponding link). Specifically, when the timer 1102 ticks once, the hardware link timer 1100 may access the entries of the FIFO memory 1104 one at a time in a polling manner, where each access entry in the FIFO memory 1104 corresponds to one of the M links.

[0169] In some embodiments, the time period of each tick of the timer 1102 can vary and can be between hundreds of microseconds to single-digit microseconds. For example, the time period of the tick of the timer 1102 can be as high as 100 microseconds or as low as 1 microsecond. In addition, the hardware link timer 1100 can adjust the time period of the tick of the timer 1102 based on the number of links (e.g., M) represented by the entries of the FIFO memory 1104. For example, when M increases (e.g., more links are represented by the entries of the FIFO memory 1104), the time period of the tick of the timer 1102 can decrease; and when M decreases (e.g., fewer links are represented by the entries of the FIFO memory 1104), the time period of the tick of the timer 1102 can increase. Therefore, if the time period of the tick of the timer 1102 changes disproportionately with the number of links represented by the entries of the FIFO memory 1104, the time interval in which the status and / or timing information of the link is checked can remain unchanged.

[0170] In some embodiments, the timing and / or status information associated with one of the M links may indicate how long the link has not received an acknowledgment of receipt of a transmitted packet. Assuming that node 400 has transmitted N packets to a second node via a link, an entry of FIFO memory 1104 may store timing and / or status information that, when accessed via polling during a particular time period of a tick of timer 1102, indicates that an acknowledgment of receipt of any of the N packets has not been received within a predetermined duration (e.g., 20 microseconds, 50 microseconds, 100 microseconds, 200 microseconds, 300 microseconds, 400 microseconds, 500 microseconds, and / or any duration therebetween). When accessing an entry of the FIFO memory 1104, the hardware link timer 1100 can utilize the logic circuit systems 1120, 1112, 1114, 1116 and 1118 to check the timing and / or status information stored in the entry, and look up N packets that may be stored in a local storage device of the node 400 (e.g., TX storage device 434-1 or other local storage device) to replay the N packets.

[0171] Alternatively, timing and / or status information associated with one of the M links may be stored in an entry of FIFO memory 1104 to indicate that the link may be shut down (e.g., all packets sent by the first node have been received by the second node). When accessing an entry of FIFO memory 1104, hardware link timer 1100 may utilize logic circuitry 1120, 1112, 1114, 1116, and 1118 to check the timing and / or status information stored in the entry and find packets that may still be stored in the local storage device (e.g., TX storage device 434-1) of node 400 and discard these packets because the timing and / or status information stored in the entry of FIFO memory 1104 indicates that the link may be shut down. Advantageously, by utilizing a single timer (e.g., timer 1102) that ticks over an adjustable period of time for multiple links and / or packets and a FIFO memory 1104 that stores timing and / or state information for multiple links, the node 400 can replay packets at appropriate timing to achieve low latency and free up hardware resources occupied by inactive links (e.g., closed links) for use by active links with limited computing and storage resources.

[0172] like Fig.11 As shown, logic circuit systems 1120, 1112, 1114, 1116, and 1118 may operate in different pipeline stages, similar to Fig.10 The logic circuit systems 1012, 1014, 1016 and 1018 are shown. Fig.11As shown, logic circuitry 1120, 1112, 1114, 1116, and 1118 may operate in conjunction with timer 1102 and FIFO memory 1104 to determine when packets sent over one or more links need to be replayed or may be withdrawn / discarded from local storage (such as TX storage 434-1), or whether one or more links may be shut down. Fig.11 As shown, logic circuit systems 1120, 1112, 1114, 1116, and 1118 may operate at corresponding pipeline stages according to a clock operated by hardware link timer 1100. Specifically, logic circuit systems 1120 and 1112 may operate at an initial pipeline stage (labeled as "Q0"), logic circuit system 1114 may operate at a first pipeline stage (labeled as "Q1"), logic circuit system 1116 may operate at a second pipeline stage (labeled as "Q2"), and logic circuit system 1118 may operate at a third pipeline stage (labeled as "Q3").

[0173] In operation, at an initial pipeline stage Q0, logic circuitry 1120 may select timing and state information to be used for a timing and state information lookup (eg, a TIMER link lookup) of logic circuitry 1112. Fig.11 As shown, the timing and status information may come from entries in the FIFO memory 1104 (e.g., the oldest entry that entered the FIFO memory 1104 earlier than all other entries) or from entries from other sources (e.g., alternate priority link lookup information). Fig.11 As shown, at the initial pipeline stage Q0, the timing and status information associated with "Link A" in FIFO memory 1104 is selected by logic circuit system 1112 based on a control signal (e.g., "Pick Up") that selects "Timer Link Lookup" instead of "TX Traffic" or "RX Traffic." "TX Traffic" may correspond to packets sent via a link established by node 400 (e.g., "Link B"), while "RX Traffic" may correspond to packets received via another link established by node 400 (e.g., "Link D").

[0174] At the first pipeline stage Q1, logic circuitry 1114 determines which link is being queried based on the timing and status information received from the initial pipeline stage Q0. Fig.11As shown, the logic circuit system 1114 determines that "Link A" is being queried to determine later whether "Link A" needs to be replayed or can be closed. Then, in the second pipeline stage Q2, the logic circuit system 1116 determines whether "Link A" can be closed based on the timing and status information associated with "Link A" accessed from the FIFO memory 1104. If the timing and status information associated with "Link A" shows that "Link A" can be closed, the logic circuit system 1116 can trigger the packets associated with "Link A" to be withdrawn / discarded from the local storage device (e.g., TX storage device 434-1). If the timing and status information associated with "Link A" shows that "Link A" is still in an active / open state, the operation of the hardware link timer 1100 continues to the third pipeline stage Q3, where the logic circuit system 1118 determines whether to replay the packets sent through "Link A" or how to update the timing and status information associated with "Link A".

[0175] At the third pipeline stage Q3, the logic circuit system 1118 may determine that at least some of the packets associated with "Link A" are to be replayed based on the state and timing information associated with "Link A" accessed from the FIFO memory 1104. For example, the state and timing information associated with "Link A" may include a "timer bit" that, when set (e.g., set to a logic 1), may indicate that the node 400 has not received an acknowledgment of receipt of at least one of the packets associated with "Link A" within a threshold duration for replaying packets. In some embodiments, the threshold duration may be adjustable and may be 20 microseconds, 50 microseconds, 100 microseconds, 200 microseconds, 300 microseconds, 400 microseconds, 500 microseconds, and / or any suitable duration therebetween. The threshold duration may be in the range of 20 microseconds to 500 microseconds. In some embodiments, the "timer bit" associated with "Link A" (and / or other links) may be set based on the number of times "Link A" is queried from the FIFO memory 1104 and the time period of the timer 1102.

[0176] If the "timer bit" is asserted, the logic circuit system 1118 may cause the packets associated with "Link A" to be replayed. The "timer bit" being asserted may indicate that a timeout associated with one or more packets has occurred (e.g., a threshold duration has been reached without receiving an acknowledgment or negative acknowledgment). In addition, in response to the replay of "Link A", the logic circuit system 1118 may update the timing and status information associated with "Link A" stored in the FIFO memory 1104. For example, the logic circuit system 118 may clear the "timer bit" (e.g., set the "timer bit" from a logic 1 to a logic 0). On the other hand, if the status and timing information associated with "Link A" indicates not to replay one or more packets on "Link A" (e.g., the "timer bit" is not asserted, which corresponds to Fig.11 110), then logic circuit system 1118 may not cause "Link A" to be replayed. In this case, if the timing and status information associated with "Link A" indicates that "Link A" should be replayed the next time it is queried, then logic circuit system 1118 may also set the "Timer Bit" to a logic 1.

[0177] Example Methods for Replay and Link Timing

[0178] Now go to Fig.12 , will describe a method for replaying a slave node (such as node 400 or Figure 3B The packet replay process 1200 may be performed, for example, by the TTP tag block 436 or Figure 4 The process 1200 begins at block 1202, where the TTP tag block 436 may store a linked list including packets sent from the node 400 to the second node via the first link using the Ethernet protocol. For example, the linked list may be the TX linked list 1020, which includes or references packets 1022, 1024, and 1026 to maintain the order of the packets 1022, 1024, and 1026 for sending to the second node.

[0179] At block 1204, the TTP tag block 436 may determine that a first packet in the packets is to be replayed in response to at least one of: (a) receiving a negative acknowledgment of the first packet from the second node, or (b) a timeout associated with the first packet. For example, the TTP tag block 436 may determine that packet 1024 is to be replayed in response to (a) receiving a negative acknowledgment of packet 1024 from the second node, or (b) a timeout associated with packet 1024, indicating that an acknowledgment of packet 1024 was not received within a threshold time period.

[0180] In response to receiving an acknowledgment of a second one of the packets from the second node, TTP label block 436 may withdraw the second packet at block 1206. For example, in response to receiving an acknowledgment of packet 1022 from the second node, TTP label block 436 may withdraw packet 1022.

[0181] Fig.13 400 or 410. Figure 3B Example link timeout process 1300 for one or more links associated with device A). Link timeout process 1300 may be performed, for example, by Fig.11 The process 1300 is implemented by the hardware link timer 1100 or node 400 of the embodiment of the present invention. The process 1300 begins at block 1302, where the hardware link timer 1100 or node 400 stores timing and status information associated with multiple links in a FIFO memory, and the node 400 sends a packet to one or more other nodes through the multiple links using an Ethernet protocol. For example, the hardware link timer 1100 may store the timing and status information associated with the multiple links in the FIFO memory 1104.

[0182] At block 1304, the hardware link timer 1100 or the node 400 may access entries of the FIFO memory based on corresponding ticks of hardware timers deployed within the hardware link timer 1100 or the node 400. For example, the hardware link timer 1100 may access entries of the FIFO memory 1104 based on corresponding ticks of the timer 1102.

[0183] At block 1306, the hardware link timer 1100 or the node 400 may determine, based on the timing and status information associated with the first link of the plurality of links, that at least one packet associated with the first link is to be replayed. For example, the hardware link timer 1100 may determine, based on the timing and status information associated with “Link A,” that at least one packet associated with or sent over “Link A” is to be replayed.

[0184] in conclusion

[0185] The above disclosure is not intended to limit the present disclosure to the precise form disclosed or to the specific field of use. Therefore, in light of the present disclosure, it is contemplated that various alternative embodiments and / or modifications to the present disclosure (whether explicitly described or implied herein) are possible. Having thus described the embodiments of the present disclosure, it will be appreciated by those of ordinary skill in the art that changes may be made in form and detail without departing from the scope of the present disclosure. Therefore, the present disclosure is limited only by the claims.

[0186] It should be understood that not all objects or advantages may be achieved according to any particular example described herein. Thus, for example, those skilled in the art will recognize that some examples may operate in a manner that achieves or optimizes one or a group of advantages taught herein without necessarily achieving other objects or advantages taught or suggested herein.

[0187] All processes described herein can be embodied in software code modules executed by a computing system including a computer or processor, and fully automated by these modules. The code modules can be stored in any type of non-transitory computer readable medium or other computer storage device. Some or all methods can be embodied in dedicated computer hardware.

[0188] It will be apparent from this disclosure that there are many other variations besides those described herein. For example, according to the examples, some actions, events, or functions of any algorithm described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the algorithm). In addition, in some examples, actions or events may be performed concurrently, such as through multithreading, interrupt handling, or multiple processors or processor cores, or on other parallel architectures, rather than sequentially. In addition, different tasks or processes may be performed by different machines and / or computing systems that may work together.

[0189] The various illustrative logic blocks and modules described in conjunction with the examples disclosed herein may be implemented or executed by a machine designed to perform the functions described herein, such as a processing unit or processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The processor may be a microprocessor, but in an alternative, the processor may be a controller, a microcontroller or a state machine, a combination thereof, or the like. The processor may include a circuit system for processing computer executable instructions. In some examples, the processor includes an FPGA or other programmable device that performs logic operations without processing computer executable instructions. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, a microprocessor combined with a DSP core, or any other such configuration. Although primarily described herein with respect to digital technology, the processor may also primarily include analog components. The computing environment may include any type of computer system, including but not limited to a microprocessor-based computer system, a mainframe computer, a digital signal processor, a portable computing device, a computing engine within a device controller or an appliance, and the like.

[0190] The elements of the methods, processes, routines or algorithms described in conjunction with the embodiments disclosed herein may be directly embodied in hardware, software modules executed by a processor device, or a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-transient computer-readable storage medium. An exemplary storage medium may be coupled to a processor device so that the processor device can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integrated into the processor device. The processor device and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor device and the storage medium may reside in a user terminal as discrete components.

[0191] The processes described herein or shown in the drawings of the present disclosure may be initiated in response to an event, such as starting on a predetermined or dynamically determined schedule, starting on demand when initiated by a user or system administrator, or starting in response to some other event. When such a process is initiated, a set of executable program instructions stored on one or more non-transient computer-readable media (e.g., hard disk drivers, flash memory, removable media, etc.) may be loaded into a memory (e.g., RAM) of a server or other computing device. The executable instructions may then be executed by a hardware-based computer processor of the computing device. In some embodiments, such a process or portions thereof may be implemented serially or in parallel on multiple computing devices and / or multiple processors.

[0192] Unless expressly stated otherwise, conditional language such as "can," "could," "might," or "may" is understood in the context of being generally used to convey that certain examples include and other examples do not include certain features, elements, and / or steps. Thus, such conditional language generally does not indicate that features, elements, and / or steps are in any way examples, nor does it indicate that examples necessarily include logic for deciding whether to include or perform such features, elements, or steps in any particular example, with or without user input or prompting.

[0193] Unless specifically stated otherwise, disjunctive language, such as the phrase "at least one of X, Y, or Z," should be understood with the context generally used to indicate that an item, term, etc. can be X, Y, Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language generally does not and should not indicate that some examples require that at least one of X, at least one of Y, or at least one of Z are all present.

[0194] Any process description, element or block in the flowcharts described herein and / or shown in the drawings should be understood to represent a code module, segment or portion, which includes executable instructions for implementing specific logical functions or elements in the process. Alternative examples are included within the scope of the examples described herein, in which elements or functions can be deleted, executed out of the order shown or discussed, including substantially concurrently or in reverse order, depending on the functions involved, as understood by those skilled in the art.

[0195] It should be emphasized that many changes and modifications may be made to the above examples, and its elements should be understood as one of other acceptable examples. All such modifications and changes are intended to be included within the scope of the present disclosure.

[0196] Any process description, element or block in the flowcharts described herein and / or shown in the drawings should be understood to represent a code module, segment or portion, which includes executable instructions for implementing specific logical functions or elements in the process. Alternative implementations are included within the scope of the examples described herein, in which elements or functions can be deleted, executed out of the order shown or discussed, including substantially concurrently or in reverse order, depending on the functions involved, as understood by those skilled in the art.

[0197] Unless expressly stated otherwise, articles such as "a" or "an" should generally be interpreted as including one or more of the items. Thus, phrases such as "a device configured to..." are intended to include one or more of the referenced devices. Such one or more referenced devices may also be collectively configured to perform the references. For example, "a processor configured to perform references A, B, and C" may include a first processor configured to perform reference A, the first processor working in conjunction with a second processor configured to perform references B and C.

Claims

1. A first node for sending a packet in an Ethernet-based network, wherein the first node include: One or more processors, including: a first-in-first-out (FIFO) memory configured to store timing and status information associated with a plurality of links, wherein the first node is configured to send packets to one or more other nodes over the plurality of links using an Ethernet protocol; a timer configured to tick according to a time period, wherein the timer is associated with the plurality of links; and A logic circuit system is configured to: accessing entries of the FIFO memory based on corresponding ticks on the timer; and determining to replay at least one packet associated with a first link of the plurality of links based on the timing and status information associated with the first link, The Ethernet protocol described therein is lossy.

2. The first node of claim 1, wherein the logic circuitry is configured to access the entries of the FIFO memory in a polling manner. 3 . The first node of claim 1 , wherein the timer is configured to adjust the time period based on a number of active links associated with the entry of the FIFO memory, wherein the active links are included in the plurality of links.

4. The first node of claim 1, wherein the logic circuitry is configured to determine to withdraw packets associated with a second link of the plurality of links based on the timing and status information associated with the second link.

5. A first node according to claim 4, wherein the packet associated with the second link is stored in a local storage device of the first node, and wherein in response to determining that the packet associated with the second link is to be withdrawn, the logic circuit system causes the local storage device to discard the packet associated with the second link.

6. The first node of claim 1, wherein the logic circuitry is configured to determine to shut down a second link of the plurality of links based on the timing and status information associated with the second link.

7. A first node according to claim 1, wherein the timing and status information associated with the first link among the multiple links indicates that within a threshold duration for replaying packets, an acknowledgment of receiving the at least one packet associated with the first link has not been received by the first node.

8. A first node for Ethernet-based communication, the first node include: One or more processors configured to implement a transport layer hardware-only Ethernet protocol, The transport layer hardware-only Ethernet protocol is lossy, and Wherein the one or more processors include a hardware link timer configured to determine packets sent for replay under the transport layer hardware-only Ethernet protocol.

9. The first node of claim 8, wherein the first node sends a first plurality of packets over a first link and sends a second plurality of packets over a second link according to the transport layer hardware-only Ethernet protocol, and wherein the hardware link timer include: A first-in-first-out (FIFO) memory is configured to store timing and status information associated with the first link in a first entry of the FIFO memory and to store timing and status information associated with the second link in a second entry of the FIFO memory.

10. The first node of claim 9, wherein the hardware link timer comprises a timer associated with a plurality of links that ticks according to a time period, wherein the hardware link timer ticks accesses entries of the FIFO memory in a polling manner of the timer, wherein the entries comprise the first entry and the second entry.

11. The first node of claim 10, wherein the hardware link timer is configured to adjust the time period based on a number of active links associated with an entry of the FIFO memory, and wherein the active links include the first link and the second link.

12. The first node according to claim 10, wherein the hardware link timer is configured to: determining to replay at least some of the first plurality of packets based on the timing and status information associated with the first link stored in the first entry of the FIFO memory; and A determination is made to withdraw the second plurality of packets based on the timing and status information associated with the second link stored in the second entry of the FIFO memory.

13. The first node of claim 12, wherein the second plurality of packets are stored in a local storage device of the first node, and wherein in response to determining that the second plurality of packets are to be withdrawn, the hardware link timer causes the local storage device to discard the second plurality of packets.

14. The first node of claim 12, wherein the timing and status information associated with the first link indicates that an acknowledgement of receiving one of the first plurality of packets has not been received by the first node within a threshold duration for replaying packets.

15. A computer-implemented method implemented at a first node in an Ethernet-based network, the computer-implemented method include: storing timing and status information associated with a plurality of links in a first-in-first-out (FIFO) memory of the first node, wherein the first node is configured to send packets to one or more other nodes over the plurality of links using an Ethernet protocol; accessing entries of the FIFO memory based on respective ticks of a hardware timer; as well as determining to replay at least one packet associated with a first link of the plurality of links based on the timing and status information associated with the first link, The Ethernet protocol described therein is lossy.

16. The computer implemented method of claim 15, wherein the entries of the FIFO memory are accessed in a polling manner.

17. The computer-implemented method of claim 15, further comprising: include: A time period of the hardware timer is adjusted based on a number of active links associated with the entry of the FIFO memory, wherein the active link is included in the plurality of links.

18. The computer-implemented method of claim 15, further comprising: include: A determination is made based on the timing and status information associated with a second link of the plurality of links to withdraw packets associated with the second link.

19. The computer-implemented method of claim 15, further comprising causing the at least one packet associated with the first link to be replayed.

20. A computer-implemented method according to claim 15, wherein the timing and status information associated with the first link of the plurality of links indicates that an acknowledgment of receiving the at least one packet associated with the first link has not been received by the first node within a threshold duration for replaying packets.