Best effort hardware RDMA transmission
By improving the hardware RDMA transmission scheme, the problem of low effective throughput in data centers and cloud networks is solved, achieving efficient RDMA operation and resource optimization, and improving transmission efficiency and scalability.
Patent Information
- Application Number
- CN202480018847.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-15
- Filing Date
- 2024-03-14
- Publication Date
- 2025-11-07
AI Technical Summary
Existing RDMA transmission protocols suffer from low effective throughput in data centers and cloud networks, mainly due to network congestion, improper retransmission mechanisms, out-of-order delivery, and unreasonable resource allocation, which leads to decreased transmission efficiency.
A hardware-based RDMA transmission scheme is adopted, which achieves efficient RDMA operation through improved congestion control, out-of-order tracking resource management and resource allocation mechanisms, combined with a smart network interface controller (NIC). It supports operations such as SEND, WRITE, and READ, and optimizes resource utilization.
It improves the effective throughput of RDMA transmission, ensures high performance even under extreme conditions, reduces on-chip complexity and resource consumption, and supports scalability and compatibility.
Smart Images

Figure CN120917437A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure generally relates to hardware remote direct memory access (RDMA) performed by an intelligent network interface controller (NIC). BACKGROUND Hardware-based RDMA transport (single-path) effective throughput is impacted by its deployment in a best-effort datacenter or cloud network for the following reasons: 1. Best-effort networks allow packet loss. These "lost" packets are usually caused by network congestion. Network congestion in datacenter networks is difficult to control due to the heterogeneous mix of transport and workloads deployed on the network, and the uncoordinated congestion control mechanisms deployed for each of them. Some RDMA transports (e.g., RoCEv2 Reliable Connections (RC) or Extended Reliable Connections (XRC), hereinafter often referred to as RoCEv2 or RoCE) are designed for "lossless" L2 networks, and as such, the common mechanism to address the best-effort loss problem in Ethernet LANs is to deploy IEEE 802.1Q Priority Flow Control (PFC), but this introduces large-scale head-of-line blocking, congestion spreading, and can lead to deadlocks. Further, converged Ethernet is extremely difficult to operate in networks with thousands to tens of thousands of servers, which is common in cloud environments. Further, PFC and other link control mechanisms are problematic on very high bandwidth links such as 100 Gb.
[0003] 2. Known transports using GoBackN reliability protocol (e.g., RoCEv2) force retransmission of all correctly received packets after a lost packet.
[0004] 3. Known transports using out-of-order placement (e.g., iWARP's Direct Data Placement (DDP)) and selective acknowledgement reliability protocol (e.g., TCP-based iWARP with SACK) have no signaling to inform the sender that the receiver has run out of out-of-order tracking resources. As a result, the uninformed transmitter continues to transmit packets that must be discarded at the receiver even when correctly received by the receiver.
[0005] 4. Some asymmetric transports (e.g., RoCEv2) provide reliability signaling (ACK / NAK) only for the requester traffic (request packet stream of the associated RC QP), and have no protocol-mandated signaling to protect the responder traffic (data-carrying response packet stream of the associated RC QP). This results in asymmetric effective throughput for the two types of traffic (e.g., the responder-only RDMA node receives no other transport signaling except CNP for congestion control).
[0006] 5. Static / Fixed Retransmission Timeout (RTO), each packet has a maximum number of attempts (e.g., 7 attempts for RoCEv2), which degrades effective throughput when tail drop occurs.
[0007] Previous solutions 1. RoCEv2: suffers from all the effective throughput loss issues mentioned above (and dependency on PFC).
[0008] 2. iWARP: TCP / IP based RDMA transport with out-of-order placement and selective acknowledgement. iWARP suffers from limited reliability protocol (TCP ACK and SACK) and lacks the signaling needed to inform of out-of-order tracking resource exhaustion. Moreover, the on-chip complexity of iWARP is increased due to its dependency on TCP.
[0009] 3. 1RMA (1RMA: Re-envisioning Remote Memory Access for Multi-tenant Datacenters, available at research.google / pubs / pub49377 / or ACM Special Interest Group on Data Communication Annual Conference Proceedings on Applications, Technologies, Architectures, and Protocols for Computer Communications, Association for Computing Machinery, (2020), 708-721). 1RMA is not a fully HW based transport. Reliability, ordering, congestion control, and other transport elements that determine effective throughput are handled by SW (lower performance, higher power, higher host cycles consumed). Moreover, 1RMA transport does not support SEND operation (at least) and Verb interface (at least in kernel space), which takes a significant amount of time to stabilize in mainstream OS Linux / Windows for other standard RDMA transports. Finally, 1RMA is not a transport with publicly available specification, to the best of our knowledge, it is only deployed internally at Google and there is only one publication that defines its concept.
[0010] 4. IRN (Improved RoCE NIC, available at dl.acm.org / doi / 10.1145 / 3230543.3230557 or arxiv.org / pdf / 1806.08159.pdf). It is an extension to RoCEv2 RC for best effort networks, where one of its goals is “to make minimal changes to the RoCE NIC design to eliminate its PFC requirement”, which addresses part (not all) of the effective throughput optimization gap at high cost.
[0011] a. IRN proposes ACK / NACK signaling for the requestor traffic (and similarly, receive-ACK / receive-NACK for the responder traffic), which can cause unnecessary retransmissions as it retransmits all inferred missing packet sequence numbers (PSNs). Since an out-of-order received PSN is signaled only once by a single (receive-)NACK packet, when such a signaling packet is lost, it causes the sender to falsely infer that such a PSN is missing (never received by the receiver). This causes unnecessary retransmission of previously received packets, degrading the effective throughput.
[0012] b. The "recovery mode" proposed by IRN is triggered by every (receive-)NACK that arrives at the sender, and this mode will stop transmission of new packets while it retransmits every PSN that is marked as missing (inferred and truly missing PSNs). Thus, every new incoming (receive-)NACK causes every PSN that is marked as missing to be retransmitted. Since multiple (receive-)NACKs are received within the RTT (when the traffic encounters multiple drops within the RTT, e.g., tail drop in a switch), this can cause the same PSN to be retransmitted multiple times within the round trip time (RTT). These multiple retransmissions are unnecessary and degrade the effective throughput.
[0013] c. IRN also extends the transmission with basic end-to-end flow control (i.e., BDP-FC) by assuming a constant bandwidth delay product (BDP) for all connections and dividing it by the maximum transmission unit (MTU) to obtain the maximum outstanding packet number for each connection (this outstanding packet number is tracked by the sender for its own flow control). This end-to-end flow control will not be fine-tuned for each connection, or it does not need to be characterized for each connection's path in order to be fine-tuned. The un-tuned BDP can cause over-subscribed connections or under-subscribed connections, and in the over-subscribed case, it can cause increased congestion (i.e., an IRN with an over-estimated BDP will heavily rely on congestion control to maintain the effective throughput), while an IRN with an under-estimated BDP will not be able to achieve the maximum available effective throughput.
[0014] d. IRN requires large state expansion for each connection on both the sender side and the receiver side, which does not solve the feasibility of the scalability challenge as the number of connections expands to the public cloud level.
[0015] i. IRN needs a bitmap to track the receive / acknowledgment status of each packet in the BDP in-flight packets, which is one of the most state-consuming options to solve the effective throughput problem.
[0016] ii.IRN requires a timer on the responder side. Typically, the transmission uses a timer only on the requester side, because the timer is an expensive resource to extend, so IRN transmission replicates this extension cost.
[0017] e.Finally, IRN is not a transmission with a publicly available specification, and as far as is known, it is only described in a single publication that defines its concept.
[0018] 5. QUIC transport protocol: QUIC is a UDP-based SW transport protocol that lacks RDMA operations by design (the transport was not designed for HW implementation, it was initially implemented as a user-space library, with kernel implementation as a future option). QUIC defines a Probe Timeout (PTO) feature as a tail loss avoidance feature that in turn reduces the likelihood of a retransmission timeout (RTO), which is often the main cause of reduced effective throughput. PTO triggers sending a probe datagram when the packet that elicited an ACK is not acknowledged within an expected period of time. The implementation of the QUIC PTO algorithm is based on the reliability features of the TCP Tail Loss Probing (TLP), RTO, and F-RTO algorithms.
[0019] 6. Current RoCE / IRN and iWARP Verb have limited resource allocation granularity for RNICs. QP-related HW resources are allocated during QP creation or QP modification (before the QP is ready to exchange bidirectional traffic) by force. HW resources include SQ, RQ, CQ, RSPQ (a responder resource used to process responses for inbound RDMA READ / ATOMIC requests), inbound request tracking, and inbound response tracking (to support out-of-order data placement).
[0020] When a QP is created (via CreateQP Verb), a pair of queues (SQ and RQ) is always created and associated with their CQ(s). Note that the CQ(s) need to be created before the QP is created via CreateCQ Verb. On ModifyQP (when changing state as follows: Init -> Init, Init -> RTR, RTR -> RTS, RTS -> RTS, SQD -> SQD, SQD -> RTS), incoming RDMA READ, RDMA WRITE, and ATOMIC operations can be independently enabled / disabled.
[0021] IB Verbs allow an end user to query attributes of a specified Host Channel Adapter (HCA) via the QueryHCA Verb. The maximum responder resources per QP and the maximum responder resources per HCA are among the attributes that can be queried. On ModifyQP, one can also specify the responder resources (number of responder resources to use for processing incoming RDMA READ and ATOMIC operations). This cannot exceed the maximum value allowable for QPs on that HCA.
[0022] From a resource perspective, the imposed limitations are listed as follows: a. An RQ is always created with the CreateQP Verb. If the node on the connection expects not to receive Send / RDMA WRITE w / IMMEDIATE messages, the RQ will never be used. The CQ entry allocated for the RQ will also be wasted.
[0023] b. An SQ is always created with the CreateQP Verb. If the node on the connection expects to be a receiver-only (or responder-only) node, the SQ will never be used. The CQ entry allocated for the SQ will also be wasted.
[0024] c. Existing Verbs do not provide an explicit way to not allocate an RSPQ when it is not needed (i.e., the node on the connection expects not to receive RDMA READ / ATOMIC). An intelligent application can figure out that the QP will not receive any inbound RDMA READ / ATOMIC requests (when its peer advertises the sender depth as 0), and specify 0 responder resources via the ModifyQP Verb to instruct the adapter not to allocate an RSPQ. Whether an RSPQ will be allocated still depends on the device, as this is not a strict requirement of the Verb.
[0025] d. Existing Verbs do not cover the out-of-order tracking resources, as RoCEv2 is proposed for PFC-enabled lossless networks. Therefore, the user has no control over any out-of-order tracking resources implemented by RoCE / IRN or iWARP, which can result in over-allocation, even when it is not necessary.
[0026] The described limitations for each known RDMA transport result in reduced effective throughput. It is desirable to improve the effective throughput over what is provided by the RDMA transport protocol. BRIEF DESCRIPTION OF DRAWINGS
[0027] For the purposes of illustration, certain examples described in the present disclosure are shown in the drawings. Like numbers refer to like elements throughout. The full scope of the application disclosed herein is not limited to the precise arrangements, sizes, and instrumentalities shown in the illustrations. In the drawings: FIG. 1is a block diagram of a cloud service provider environment where virtual machines at one cloud service provider location are connected to various virtual machines in another cloud service provider location.
[0028] FIG. 2A is a block diagram of a processor board of a server present in a cloud service provider environment according to an example of the present invention.
[0029] FIG. 2B is a block diagram of a network interface card (NIC) of a server present in a cloud service provider environment according to an example of the present invention.
[0030] FIG. 2C is a ladder diagram of RDMA operations of sender and receiver hosts and NICs according to an example of the present invention.
[0031] FIG. 2D is a flow diagram of NIC and host RDMA initialization operations according to an example of the present invention.
[0032] FIG. 3 illustrates all roles of a BE transfer according to an example of the present invention.
[0033] FIG. 4 illustrates nodes A and B with the requester role only of a BE transfer according to an example of the present invention.
[0034] FIG. 5 illustrates nodes A and B with the requester role and node B with the responder role of a BE transfer according to an example of the present invention.
[0035] FIG. 6A and FIG. 6B illustrates a sender packet level state machine according to an example of the present invention.
[0036] FIG. 7A and FIG. 7B illustrates a receiver packet level state machine according to an example of the present invention.
[0037] FIG. 8A and FIG. 8B illustrates a sender connection level state machine according to an example of the present invention.
[0038] FIG. 8C is a flow diagram of operations of a sender connection state machine and a receiver connection state machine when receiving a packet or transmitting a packet.
[0039] FIG. 9A and FIG. 9B illustrates a receiver connection level state machine according to an example of the present invention.
[0040] FIG. 9Cis a flowchart of the process of determining the inf_pid value by the receiver connection state machine when entering the RECOVERY superstate.
[0041] FIG. 10 is an illustration of the base transport header (BTH) as used in RDMA operations.
[0042] FIGS. 10A-10F illustrates the packet format of a transaction packet using message level interleaving according to an example of the invention.
[0043] FIGS. 11A-11C illustrates the packet format of a reliability protocol packet using message level interleaving according to an example of the invention.
[0044] FIG. 11D and FIG. 11E illustrates the packet format of a transaction packet using packet level interleaving according to an example of the invention.
[0045] FIG. 11F and FIG. 11G illustrates the packet format of a reliability protocol packet using packet level interleaving according to an example of the invention.
[0046] FIG. 12A is an illustration of certain header values in an example packet for message level interleaved RDMA operations according to an example of the invention.
[0047] FIG. 12B is an illustration of certain header values in an example packet for packet level interleaved RDMA operations according to an example of the invention.
[0048] FIG. 12C is an illustration of certain header values in an example packet for separate PID space interleaved RDMA operations according to an example of the invention.
[0049] FIG. 13A is a ladder diagram of an RDMA WRITE operation according to an example of the invention.
[0050] FIG. 13B is a ladder diagram of an RDMA READ operation according to an example of the invention.
[0051] FIG. 14 illustrates various alternatives for the transmission of packets and the reception of packets.
[0052] FIG. 15 illustrates the relationship of FIGS. 15A-15E
[0053] FIGS. 15A-15E illustrates the connection level and packet level state machine operations for the sender and receiver for the alternative (b) for the reception of packets.
[0054] FIG. 16 illustrates FIGS. 18A-18G the relationship.
[0055] FIGS. 16A-16G illustrates connection-level and packet-level state machine operations for the sender and receiver for alternative (c) for packet reception.
[0056] FIG. 17 illustrates FIGS. 17A-17G the relationship.
[0057] FIGS. 17A-17G illustrates connection-level and packet-level state machine operations for the sender and receiver for alternative (d) for packet reception.
[0058] FIG. 18 illustrates FIGS. 18A-18H the relationship.
[0059] FIGS. 18A-18H illustrates connection-level and packet-level state machine operations for the sender and receiver for alternative (e) for packet reception. DETAILED DESCRIPTION
[0060] A viable HW-based RDMA transport that addresses the (above) effective throughput problem should guarantee at least the same performance as a specification transport (e.g., RoCEv2 or iWARP) in extreme (pathological) cases.
[0061] A viable HW-based RDMA transport that addresses the (above) effective throughput problem should minimize added on-chip complexity and resources by avoiding the use of a data reordering buffer and minimizing connection state (including out-of-order tracking resources) in order to make the transport scalable.
[0062] It is desirable for a HW-based RDMA transport to decouple from (not impose restrictions on) congestion control solutions so that known good congestion control algorithms can be used with the new transport and even future innovations in the congestion control algorithm space can be able to minimize imposed restrictions.
[0063] It is preferable for a HW-based RDMA transport to define support for all commonly used RDMA operations (SEND, WRITE, READ, ATOMIC, etc.) and it is also preferable for a HW-based RDMA transport to support the standard Verb interface.
[0064] Reference is now made to FIG. 1FIG. 1 illustrates a cloud service provider environment. A first cloud service provider 100 is connected to a second cloud service provider 104 through a lossy Ethernet 102, such as the Internet or other wide area network (WAN). The first cloud service provider 100 is illustrated as having a first server 106A and a second server 106B. Each server 106A, 106B includes three virtual machines (VMs) 108 and a hypervisor 110A, 110B, with server 106A including VMs 108A-108C and server 106B including VMs 108D-108F. While VMs are used as an example environment in this specification, it should be understood that containers or other similar items operate similarly such that VMs are used in this specification to represent any of VMs, containers, or similar entities. The second cloud service provider 104 is illustrated as having three servers 112A, 112B, and 112C. Each server 112A-112C includes three virtual machines 114, with server 112A including VMs 114A-114C, server 112B including VMs 114D-114F, and server 112 including VMs 114G-114I; and hypervisors 116A-116C. More details of server 112A are provided as an example of servers 106A, 106B, 112B, and 112C. Server 112A includes a processing unit 118, a NIC 120, a RAM 122, and a non-transitory storage 124. The RAM 122 includes virtual machines 114A-114C in operation and a hypervisor 116 in operation. The non-transitory storage 124 includes a host operating system 126, virtual machine images 128, and a stored version of a hypervisor 130.
[0065] Servers 106A and 106B are connected within the first cloud service provider 100 through a lossy Ethernet network that allows for connection to network 102 and the second cloud service provider 104. Similarly, servers 112A, 112B, and 112C are connected within the second cloud service provider 104 through a lossy Ethernet network to allow for access to network 102 and the first cloud service provider 100.
[0066] FIG. 1A complex environment of VMs is illustrated. VM 108A in server 106A is connected to VM 114A and VM 114B in server 112A, VM 114D in server 112B, and VM 114G in server 112C. VM 108E in server 106B is connected to VM 114E in server 112B. VM 108F in server 106B is connected to VM 114H in server 112C. Thus, VM 108A is connected to four different VMs in three different servers, while servers 112B and 112C each have two VMs connected to two different servers. For the purposes of this description, any VM-to-VM connection can include an RDMA connection. This is a very simple example for explanation purposes. In a common environment, hundreds of VMs are running on a single server, and a single VM is connected to thousands of remote VMs on thousands of servers. This means that any given VM can have a large number of RDMA connections. This environment is the target environment for the HW RDMA transport based on a best effort (BE) network described herein.
[0067] For the purposes of this description, each NIC 120 is a smart NIC, which can perform RDMA operations as described below.
[0068] Terminology RDMA NIC (RNIC): A network interface controller (NIC) with RDMA transport capabilities.
[0069] Packet ID (PID): A unique packet identifier for all non-signaling packets. A generalization of the packet sequence number (PSN) concept.
[0070] Requestor role: The role of a connected network node that sends request messages (e.g., SEND, WRITE, READ REQ, etc.) to a connected peer node. In response to a signaling packet received from a connected peer node, this role also handles request packet retransmissions as required by the reliability protocol.
[0071] Responder role: The role of a connected network node that sends response messages (e.g., READ RSP, etc.) to a connected peer node. In response to a signaling packet received from a connected peer node, this role also handles response packet retransmissions as required by the reliability protocol.
[0072] Completer role: The role of a connected network node that sends signaling packets (as required by the reliability protocol) to a connected peer node. The signaling packets are sent in response to received packets that belong to a request message or a response message. This role can be further subdivided into two sub-roles: 1. Inbound Request Tracking Sub-Role: This sub-role handles the reliability protocol for inbound request traffic.
[0073] 2. Inbound Response Tracking Sub-Role: This sub-role handles the reliability protocol for inbound response traffic.
[0074] Initiator Node: A connected network node that initiates a transaction by sending a request message to a target node over a network transfer.
[0075] Target Node: A connected network node that services a transaction associated with a request message received from an initiator node over a network transfer.
[0076] Request Traffic: A sequence of packets that transfer a request message (e.g., SEND, WRITE, READ REQ, ATOMIC REQ, etc.) from a requester role of an initiator node to a completer role of a target node (i.e., in connected peer nodes).
[0077] Response Traffic: A sequence of packets that transfer a response message (e.g., READ RSP, ATOMIC RSP, etc.) from a responder role of a target node to a completer role of its initiator node (i.e., in connected peer nodes).
[0078] Sender: A connected network node that transmits messages (broken into non-signaling packets) and receives signaling packets in response. The requester role of an initiator node is the sender of a request message. The responder role of a target node is the sender of a response message.
[0079] Receiver: A connected network node that receives messages (broken into non-signaling packets) and transmits signaling packets in response.
[0080] 1. The completer role of an initiator node is the receiver of a response message.
[0081] 2. The completer role of a target node is the receiver of a request message.
[0082] Signaling Packet: A non-data carrying packet generated by a receiver node of request traffic or response traffic that conveys the status of the receipt of packets in the traffic.
[0083] Single-Phase Transaction (1P_Transaction): A transaction composed of a single request message from a requester role of an initiator node to a completer role of a target node, including associated signaling returned from the completer to the requester.
[0084] Two-Phase Transaction (2P_Transaction): A transaction consisting of a request message from a requester role at the initiator node to a completer role at the target node (including associated signaling from the completer back to the requester), and a response message from a responder role at the target node to a completer role at the initiator node (including associated signaling from the completer back to the responder).
[0085] One-Sided Transaction: A transaction that requires host process participation only at the initiator node with the requester role (e.g., RDMA WRITE without IMMEDIATE - no host process participation at the target side is required).
[0086] Two-Sided Transaction: A transaction that requires host process participation at both the initiator node and the target node (e.g., SEND and RDMA WRITE with IMMEDIATE require host process participation at both the initiator node and the target node).
[0087] Response Queue (RSPQ): A work queue within the RNIC that holds WQEs that represent response messages to be transmitted to the initiator node for an associated two-phase transaction. RSPQ-WQEs are automatically generated within the RNIC in response to a request message for a received two-phase transaction.
[0088] Packet Received (Pkt-Rcvd): Transmission status of request and response (non-signaling) packets. For request traffic, this status indicates that a request packet has been successfully received (in-order or out-of-order) at the target node. For response traffic, this status indicates that a response packet has been successfully received (in-order or out-of-order) at the initiator node. Pkt-Rcvd does not imply any posting, execution, delivery, or completion. This status applies to non-signaling packets for all transaction types (one-phase and two-phase, one-sided and two-sided), all traffic types (request traffic / response traffic), and on both the sender and receiver nodes.
[0089] Packet Delivered (Pkt-Dlvr): Transmission status of a request or response (non-signaling) packet. For request traffic, this status indicates that the request packet has been received in-order at the target node (i.e., all previous packets have also been received), which permits it to execute and / or complete at the target node. For response traffic, this status indicates that the response packet has been received in-order at the initiator node. A prerequisite for this status is that the packet must have been in the Pkt-Rcvd status before it can be in the Pkt-Dlvr status (this condition covers the immediate transition from Pkt-Rcvd to Pkt-Dlvr when the packet is received strictly in-order), thus delivery implies reception. This status applies to packets of all transaction types (single-phase and two-phase, single-sided and double-sided), all traffic types (request traffic / response traffic), and on both the sender and receiver nodes.
[0090] Packet Placed (Pkt-Plcd): Transmission status of the data payload carried by a request or response packet. The data payload of the packet has been successfully placed in memory at the target node. A prerequisite for this status is that the packet carrying the data payload must have been in the Pkt-Rcvd status (out-of-order placement) or Pkt-Dlvr status (in-order placement) before the data payload is placed. Placement does not imply delivery or completion of the associated packet and message. This status applies to the data payload of the following packet types: 1. Single-sided single-phase transaction request traffic on the receiver (target) node.
[0091] 2. Double-sided single-phase transaction request traffic on the receiver (target) node.
[0092] 3. Single-sided two-phase transaction response traffic on the receiver (initiator) node.
[0093] Packet Lost (Pkt-Lost): Transmission status of a request or response (non-signaling) packet. It indicates that the request or response packet has not been received by the receiver node at the time the next packet (out-of-order) is received by the receiver node. This status applies to packets of all transaction types (single-phase and two-phase, single-sided and double-sided), all traffic types (request traffic / response traffic), and on both the sender and receiver nodes.
[0094] Message Executed (Msg-Exec): Transmission status associated with a request message from the perspective of the receiving target node. This status indicates that the request message has been executed at the target node. A prerequisite for this status is that 1) the complete message must have been delivered (i.e., all packets that make up the message must be in the Pkt-Dlvr status), and 2) the requested operation must have been fully executed by the target node. This status applies to the following types of messages: 1. Single-sided single-phase transaction request message: Once the complete message has been delivered (all packets associated with the request message are in Pkt-Dlvr state), and optionally (for some operation types, e.g., persistent memory operations), once the complete message payload has been placed (all data payloads associated with the request message are in Pkt-Plcd state).
[0095] 2. Double-sided single-phase transaction request message: Once the complete message has been delivered (all packets associated with the request message are in Pkt-Dlvr state), and the complete message payload has been placed (all data payloads associated with the request message are in Pkt-Plcd state). A completion (CQE) can be issued to the associated host process on the target node.
[0096] 3. Single-sided two-phase transaction request message: Once the requested operation has been performed at the target node (e.g., a READ, FETCHADD, etc. has been performed, and the corresponding response transfer job has been queued for transmission scheduling).
[0097] Transaction Complete (Xact-Cmpl): From the initiator node’s perspective, the transmission state associated with a transaction. The transaction has completed. When a transaction enters the Xact-Cmpl state, a completion (CQE) can be issued to the associated host process on the initiator node. This state applies to all transaction types. This state applies to transactions as follows: 1. Single-sided single-phase transactions enter the Xact-Cmpl state after the outbound request message has been completely signaled back for delivery (all packets in the request message are in Pkt-Dlvr state).
[0098] 2. Double-sided single-phase transactions enter the Xact-Cmpl state after the outbound request message has been completely signaled back for delivery (all packets in the request message are in Pkt-Dlvr state).
[0099] 3. Single-sided two-phase transactions enter the Xact-Cmpl state after the following conditions are met: a. The outbound request message has been completely signaled back for delivery (all packets in the request message are in Pkt-Dlvr state).
[0100] b. The associated inbound response message has been completely delivered (all packets in the response message are in Pkt-Dlvr state). The transition to this state can (but does not have to) wait until the delivery of the associated response message has been signaled to the target node (i.e., the sender of the inbound response message).
[0101] c. Optional - for some operation types (e.g., persistent memory operations), a third condition can be that the associated inbound response message has been completely placed (all data payloads associated with the response message are in Pkt-Plcd state).
[0102] This example of a best effort connection HW RDMA transfer (hereinafter BE transfer or simply BE) is described in two steps: 1. The "baseline" transfer is described as an integration of pre-existing transfer features to address some but not all of the aforementioned issues.
[0103] 2. An "extended" transfer is described with features (to extend the baseline transfer) to further improve effective throughput beyond that achievable by the integration of pre-existing transfer features.
[0104] Baseline BE transfer Semantics Transfers can be categorized into three categories: operations, reliability, and interleaving.
[0105] 1. Operation semantics define what an operation is and how it should be handled.
[0106] a. 1P requests have two subcategories: single-sided and double-sided.
[0107] i. Single-Sided (RDMA WRITE): A data transfer operation that allows the local peer to transfer data to a previously advertised buffer (iWarp calls it a tagged buffer). This advertised buffer is carried in the packet header as a remote key. The local peer can collect from multiple local source buffers and scatter to a single remote sink buffer.
[0108] (a). Note that RDMA WRITE w / IMMEDIATE is similar to the RDMA WRITE operation, but it carries IMMEDIATE data that needs to consume the RQ WQE. Therefore, RDMA WRITE w / IMMEDIATE is categorized as a double-sided 1P request.
[0109] ii. Double-Sided (Send): A data transfer that allows the local peer to transfer data to a buffer that is not explicitly advertised (iWarp calls it an untagged buffer). The local peer can collect from multiple local source buffers and scatter to multiple remote sink buffers.
[0110] b. 2P Request (RDMA READ, ATOMIC) and 2P Response (RDMA READ Response, ATOMIC Response): Data transfer operations that allow a local peer to retrieve data from a previously advertised buffer (tagged buffer). This advertised buffer is carried in the packet header as a remote key. The local (initiator) peer can collect from a single remote (target) source buffer and scatter to multiple local (initiator) receive buffers.
[0111] c. In all cases, memory buffers are distributed as memory keys that need to be pre-registered. Key verification is required before any data transfer to / from a memory buffer.
[0112] d. RoCE supports exactly the same operation semantics. iWarp supports the same 1P request operation semantics, but the 2P request operation semantics are different as it can only have one receive buffer.
[0113] 2. Reliability semantics define how reliability is achieved with signaling packets.
[0114] a. All signaling packets carry a packet-level PID.
[0115] b. ACK accumulative acknowledges all packets up to the PID it carries.
[0116] c. SACK selective acknowledges received packets and implicitly signals packet loss. Selective retransmission can be triggered based on implicit packet loss information.
[0117] d. NAK accumulative acknowledges all packets up to PID-1 it carries. It explicitly signals packet loss at PID and requires GoBackN retransmission to be triggered at PID.
[0118] 3. Interleaving semantics define how outbound request traffic and outbound response traffic are interleaved on the same connection.
[0119] a. In message-level interleaving, one message of either traffic must complete before the next message of either traffic can continue. For example, an outbound RDMA READ Response multi-packet message must complete before a new request message such as a SEND request can be issued.
[0120] b. RoCE has no interleaving as outbound response traffic uses the same PID space as inbound request traffic on the same connection.
[0121] The baseline BE transport integrates the following pre-existing transport features with the described modifications: 1. The baseline BE transport takes the following features from InfiniBand RoCE-RC: a. Verb as an application programming interface (API) to interact with RNICs according to the requirements of the OpenFabrics Interface / OpenFabrics Enterprise Distribution (OFI / OFED) Open Fabric software distribution package. This API requires the use of the concepts of queue pair (QP), send queue (SQ), receive queue (RQ), and completion queue (CQ) as the basis for the HW / SW interface for work submission and completion reporting.
[0122] b. Messages are split into packets (each packet has the MTU size except the last packet in the message), and reliability is applied at the packet level (not at the byte level, nor at the message level). Each packet is identified by a unique PID (Packet ID), which is a generalization of the concept of packet sequence number (PSN).
[0123] c. NAK signaling packet with the same reliability semantics as in RoCE-RC.
[0124] 2. The baseline BE transport takes the following features from RoCE-RC and iWARP: a. Single-sided single-phase transactions (e.g., RDMA WRITE) have the same semantics data packing and acknowledgments as in RoCE-RC (RoCE Reliable Connected) and iWARP.
[0125] b. Double-sided single-phase transactions (e.g., SEND) have the same operational semantics as in RoCE-RC and iWARP.
[0126] c. RoCE-RC except clauses: i. Single-sided two-phase transactions (e.g., RDMA READ request / response pair) have the same operational semantics as in RoCE RC, but the reliability semantics in the BE transport are not taken from RoCE RC because RoCE RC requires a data-carrying response message with signaling properties (i.e., the data-carrying response packet also has an implicit acknowledgment property, not an ACK packet), so the acknowledgment mechanism is different.
[0127] ii. If RoCE compatibility is desired from the Verb perspective, a different / unique mnemonic / opcode can be defined for each pre-existing single-sided two-phase operation, e.g., RDMA PULL (used as an equivalent of RDMA READ for BE), RDMA PullAdd (used as an equivalent of RDMA FetchAdd for BE), etc.
[0128] 3. The baseline BE transport takes the following features from iWARP: a. Single-sided two-phase transactions require responses that can carry data (e.g., non-zero length read RDMA READ responses) or no data (e.g., zero length read RDMA READ responses) with the same reliability semantics as iWARP: i. Response traffic packets do not have implicit acknowledgement / signaling aspects.
[0129] ii. Response traffic packets are covered by the reliability protocol (same as the one used by the request traffic), so the receiver of the response packet (i.e., the initiator node that sent the initial request, e.g., RDMA READ request) must acknowledge its receipt by signaling to the sender of the response packet (i.e., the target node).
[0130] iii. Responses are committed to a response queue (RSPQ) for scheduling egress, so the outbound request traffic and response traffic associated with the same BE connection must be interleaved for transmission by the same node. For reference, iWARP only supports request traffic and response traffic interleaving semantics at message boundaries (i.e., all packets of a message must be transmitted before a message of opposite traffic can initiate transmission).
[0131] b. Response and request interleaving includes message level interleaving semantics according to iWarp.
[0132] 4. Baseline BE transport combines a mix of iWARP's Direct Data Placement (DDP) messaging and construction with RoCE-RC scatter placement messaging and construction to enable out-of-order packet reception and out-of-order scatter data placement (with in-order completion). This feature is supported in both request traffic and response traffic.
[0133] a. Baseline BE transport supports in-order and out-of-order packet reception by employing a reliability protocol with selective signaling inspired by iWARP's reliability semantics for ACK and Selective ACK (SACK) for signaling.
[0134] b. Baseline BE transport supports out-of-order placement inspired by iWARP's DDP messaging and construction but combined with RoCE-RC scatter placement messaging and construction (by using concepts like memory regions, memory windows, memory keys). The key difference from iWARP DDP is: i. BE transport does not include local receive address and length in the request message (iWARP requires these in some instances).
[0135] ii. BE transport includes control information in the request message, which is echoed in the response to the initiator node. The initiator node uses this control information to find the local placement information of the response, i.e., the information must allow the initiator node to find the associated SQ-WQE and obtain the local placement information from it.
[0136] c. Enabling the out-of-order and scatter data placement features requires adding the following new control information to the packet header, i.e., the following control information set is different from that specified for iWARP and RoCE RC protocols: i. For all request and response traffic packets: add a Packet ID (PID) to each packet, allowing the receiver of the traffic to explicitly detect holes (missing packets) in the packet stream of the traffic. Different PID formats can be supported (allowing different tradeoffs for interleaving of request and response traffic on the same connection, as described below). In the first format, the PID is equivalent to the regular PSN. In the second format, the PID is a combination of a message ID (MID) and an intra-message packet number (IMPN). Note that in the specific deployment of BE transport, all packets should use a single (selected) PID format.
[0137] ii. For one-sided single-phase transactions: RDMA addressing information (to the target node) for placement is added to each request packet. IMPN can optionally be added to facilitate message tracking. This option is connection-based, not per-packet. Note that IMPN can be inferred from the PID of the packet (depending on the PID format selected for the BE deployment), in which case it does not need to be duplicated.
[0138] iii. For two-sided single-phase transactions other than RDMA WRITE with IMMEDIATE: use RQMSN and IMPN in each request packet to enable out-of-order placement to the RQ. Note that IMPN can be inferred from the PID of the packet (depending on the PID format selected for the BE deployment), in which case it does not need to be duplicated. For RDMA WRITE with IMMEDIATE, RQMSN is needed, and IMPN is optional.
[0139] iv. For one-sided two-phase transaction requests: use RSPQ message sequence number (RSPQMSN) and RDMA addressing information (to the target node) in each request message to enable out-of-order placement to the RSPQ.
[0140] v. For single-sided two-phase transaction response: RSPQMSN and In-Msg-Num (IMPN) are used in each response packet to enable out-of-order placement of response data and to track completion of the transaction. The requester maps RSPQMSN to the associated SQ-MSN (SQ-WQE) to retrieve the scattered placement information (converts IMPN to a scattered RDMA address on the initiator node). Note that IMPN can be inferred from the packet's PID (depending on the PID format chosen for the BE deployment), in which case it does not need to be copied.
[0141] Note: Due to this out-of-order placement, memory write consistency is lost (as in iWARP).
[0142] 5. As an optimization, the baseline BE transport can optionally incorporate a Flush feature based on Flush Timeout (FTO) for early detection of tail packet loss to effectively reduce the probability of RTO. The inspiration for FTO comes from the PTO concept of QUIC. The Flush feature works as follows: a. Introduce a Flush packet (FLSH) as a special type of transport signaling packet.
[0143] i. The FLSH packet indicates the PID of the expected tail packet in the BE connection that is being flushed back.
[0144] b. The sender needs to maintain a Flush timer and a Flush counter, which are used in the Flush loop as follows: i. An FTO threshold (based on the Flush counter reaching an upper limit value) can be configured to exit the Flush loop.
[0145] ii. The sender starts the Flush timer when it transmits a tail packet of the BE connection (i.e., whenever the BE connection is in a tail drop condition); if the Flush timer is currently running, it is restarted.
[0146] (a). The tail drop condition for the BE connection is whenever the connection is susceptible to RTO due to potential tail drop.
[0147] iii. Upon receiving a signaling packet associated with a tail packet, the Flush timer is cleared.
[0148] iv. If the Flush timer expires, a FLSH packet is sent (with the PID of the tail packet of the BE connection) on the same network path as the tail packet was sent. The Flush counter is incremented.
[0149] (a). If FTO threshold is not reached, the Flush timer is restarted.
[0150] (b). If FTO threshold is reached, the connection now relies on RTO.
[0151] v. If RTO transmission timer is restarted (due to receiving an ACK packet), the Flush counter is reset.
[0152] (a). If BE connection is in tail-drop condition, the Flush timer is automatically restarted.
[0153] vi. The sender can share a single timer for both FTO Flush timer and RTO transmission timer as follows: (a). When there is no tail-drop condition, the timer is used as RTO transmission timer.
[0154] (b). When there is tail-drop condition, the timer is used as FTO Flush timer. Once FTO threshold expires, the timer is used as RTO transmission timer (set to the "remaining" RTO time, i.e. time after all FTO timeout repeats). If RTO transmission timer is reset (due to receiving an ACK), the timer can be reset again and used as FTO Flush timer (if there is tail-drop condition).
[0155] c. Receiver responds upon receiving FLSH packet.
[0156] i. Receiver sends appropriate signaling packet to indicate the status of the tail packet (PID carried by FLSH packet) being queried or the overall status of the connection.
[0157] In summary, the features used in BE baseline come from RoCE-RC, iWarp and QUIC.
[0158] 1. From RoCE-RC: a. Standard Verb API for interacting with RNIC.
[0159] b. Reliability is applied at packet level. iWarp applies reliability at byte level because it runs on top of TCP, which works on byte stream.
[0160] c. NAK is employed to explicitly signal packet loss. iWarp does not support NAK and sender needs to infer packet loss from ACK / SACK.
[0161] d. Scatter capability on local initiator host memory, i.e. Read response data can be scattered to multiple local buffers. iWarp can only scatter read response data to a single local buffer.
[0162] 2. From iWarp: a. DDP to enable out-of-order data placement on receiver side.
[0163] b. Receiver sends ACK / SACK to signal in-order and out-of-order received packets so that the sender can do selective retransmission.
[0164] c. Symmetric reliability protocol for request and response traffic. Response traffic is covered by the same reliability protocol as request traffic.
[0165] i. Response can initiate selective retransmission of response packets upon receiving SACK.
[0166] ii. In RoCE-RC, response packets are not acknowledged and response packet loss needs to rely on request packet retransmission by the requestor.
[0167] 3. From QUIC: a. Probe packet and Probe Timeout (PTO) are renamed as Flush packet and FTO.
[0168] FIG. 2A and FIG. 2B Figure illustrates a processor board 200 and NIC 202 used in a cloud service environment according to the present invention. The processor board 200 includes one or more processors 203 to perform computations. Each processor 203 includes a CPU or core 205, a PCIe root port module 204 including a phy 206 to allow the processor 203 to communicate with a peripheral device such as the NIC 202. The processor 203 includes an enhanced memory controller 208. The enhanced memory controller 208 includes a crypto block 210 to perform encryption and decryption of data residing in a RAM memory 212. An IOMMU 218 is present to monitor and control I / O operations between the NIC 202 and the VM memory of the processor board 200.
[0169] The processor board 200 includes a program memory 220, such as the program memory shown in the non-transitory memory 124. The RAM memory 212 is divided into a host kernel space 216 and a host user space 226. The host kernel space 216 includes an operating system and hypervisor 221, a NIC driver 219, an RDMA stack 240, and a BE RDMA physical function (PF) driver 242. The host user space 226 includes a VM1 222 and a VM2 224. The VM1 222 has a kernel space 229 and a user space 232. The VM1 kernel space 229 includes a guest operating system 223, an RDMA stack 227, and a BE RDMA virtual function (VF) driver 230. The VM1 user space 232 includes an encryption memory region 228, a communication program 1 231 that cooperates with a remote VM, an RDMA stack 225, a BE RDMA library 233, and an RDMA memory 235. The VM1 RDMA memory 235 includes a BE1-1 SQ 246, a BE1-1 RQ and CQ 247, a BE1-2 SQ 248, and a BE1-2 RQ and CQ.
[0170] BE1-1 represents a first BE RDMA connection for the VM1 222, and BE1-2 represents a second BE RDMA connection for the VM1 222. As FIG. 3 illustrated, the connection BE1-1 illustrates a full bi-directional transfer configuration. As FIG. 4 illustrated, the connection BE1-2 illustrates a requestor-only 1P_Transaction transfer configuration.
[0171] Similarly, the VM2 224 includes a kernel space 236 and a user space 238. The kernel space 236 includes a guest operating system 215, an RDMA stack 239, and a BE RDMA VF driver 241. The VM2 user space 238 includes an encryption memory 234, a communication program 2 237, an RDMA stack 243, a BE RDMA library 244, and a VM2 RDMA memory 245. The VM2 RDMA memory 245 includes a BE2-1 SQ 251 and a BE2-1 RQ and CQ 253. As FIG. 4 illustrated, the connection BE2-1 also illustrates a requestor-only 1P_Transaction transfer configuration.
[0172] NIC 202 includes a SmartNIC system on chip (SOC) 201, off-chip RAM 287, and SmartNIC firmware 288. While shown as a single chip, SOC 201 can also be multiple interconnected chips that operate together. SOC 201 includes a PCIe endpoint module 250, which includes a phy 249 and a host interface 256. The host interface 256 performs translation of data payloads between PCIe and ingress packet processing hardware 257, egress buffer memory 268, and hardware state machines in a pool of hardware state machines 283. A DMA controller 260 is provided in SOC 201 to perform automatic transfers between processor board 200 and NIC 202 and within NIC 202. Egress buffer memory 268 and ingress packet processing and routing hardware 257 are connected to PCIe endpoint module 250. While shown as separate egress and ingress buffer memories, this is a logical illustration, and egress buffer memory 268 and ingress buffer memory 266 can be contained in a single memory. Ingress packet processing and routing hardware 257 is connected to ingress buffer memory 266, which receives packets from an Ethernet module 267 on SOC 201. Ingress packet processing and routing hardware 257 strips underlying headers from received packets and provides metadata for routing the packets to the appropriate PF or VF. Egress buffer memory 268 is connected to egress packet build hardware 259, which builds packets provided to Ethernet module 267. Ethernet module 267 includes a phy 269, which is connected to a lossy Ethernet cloud network.
[0173] In RDMA transport using the Verb API, as used in examples according to the present application, only data payloads are exchanged between the VM and the NIC 202, so connection context information and packet specific information must be provided for each packet, whether data, request or response. The egress packet building hardware 259, along with the header processing module 261 connected to the egress packet building hardware 259, uses the connection context information and packet specific information, which are used and updated by the hardware state machine 283 during processing, which in turn obtains the connection context information and packet specific information based on the connection and traffic information associated with the packet. With the connection context information and packet specific information, the egress packet building hardware 259 and the header processing module 261 develop the appropriate RDMA header stack for the packet, such as the Ethernet header, IP header, UDP header, BTH header and BE header. The developed header stack is combined by the egress packet building hardware 259 with any data being carried in the packet and then provided to the Ethernet module 267 for transmission. The above header stack with a single Ethernet header and the like is suitable for conventional data centers. In a cloud network environment, external or underlay headers must be added that are used by the cloud network to route the packet in the cloud network. The header processing module 261 utilizes the connection context information to build the second layer of headers in the cloud network environment. Then, when the packet is assembled for provision to the Ethernet module 267, the egress packet building hardware combines the two sets of headers, the RDMA headers and the cloud network headers.
[0174] The header processing module 261 is illustrated as including a processor 252, RAM 254 and header hardware 255. A header processing operating system 258 is provided in the RAM 254 to control the operation of the processor 252. While a conventional processor and operating system are illustrated, in some examples, a special function processor can be used that is configured to directly perform the required header operations in response to input commands. The RAM 254 includes various header tables 262 that are used by the header hardware 255 to determine the appropriate external headers. The header hardware 255 acts as a match / action block, matching on header values and performing the resulting action. While the header tables are illustrated as being contained in the RAM 254, the header tables 262 can also be stored in off-chip RAM 287 to reduce the size of the RAM 254. The header processing module 261 can also receive packets directly from the Ethernet module 267 and provide packets to the Ethernet module 267 for communication with the cloud network for management operations.
[0175] In the ingress operation, the packet is provided to ingress packet processing and routing hardware 257. The ingress packet processing and routing hardware 257 analyzes the packet to check for validity. The ingress packet processing and routing hardware 257 provides the packet header to header processing module 261. The header processing module 261, using header hardware 255, uses a virtual network identifier (VNI) or virtual subnet identifier (VSID) and underlying header values in a cloud network environment, or RDMA headers in a data center environment, to determine the appropriate PF or VF and connection context information. The RDMA headers and BE headers are processed to develop packet-specific information that is used by the hardware state machine pool 283 to manage the BE protocol for the packet. The header processing module 261 returns the data payload, PF or VF metadata, connection context information, and packet-specific information to the ingress packet processing and routing hardware 257. The connection context information and packet state are provided to a state machine in the state machine pool 283 to process the packet, as described below. The packet-specific information is used to manage the BE protocol for the packet, as described below. After BE protocol processing, the ingress packet processing and routing hardware 257 then uses the PF or VF metadata and connection context information for internal queuing information; and places the packet in the appropriate output queue of the PCIe endpoint module 250. Payload or message information from the PCIe endpoint module 250 is provided to the designated host memory buffer or queue.
[0176] VM1 block 282 and VM2 block 284 are contained in on-chip RAM 263, although the VM1 block 282, VM2 block 284, and other VM blocks can be stored in off-chip RAM 287. Ingress buffer memory 266 and egress buffer memory 268 can be contained in on-chip RAM 263, or can be dedicated memory regions in the SOC 201.
[0177] The SOC 201 includes a processor 286 to perform processing operations as controlled by firmware 288, which is a non-transitory memory including an operating system 290, basic NIC functionality module 292, RDMA functionality 293, requester functionality 294, responder functionality module 296, and completer functionality 298. The tasks of the requester, responder, and completer are described below. The SOC 201 further includes a hardware state machine pool 283. Individual state machines in the hardware state machine pool 283 are assigned to packets in processing, such as packets that have been transmitted but not yet acknowledged, or packets that have been received but not yet signaled for completion.
[0178] The VM1 block 282 includes a BE1-1 portion 2202 and a BE1-2 portion 2204. The BE1-1 portion 2202 includes a requester instantiation 2206, a completer instantiation 2212 with a response table 2214 and a request table 2216, and a state machine state and context 2218, as described in more detail below. The request table 2216 includes an allocated amount of resources provided for out-of-order tracking. In operation, when a packet is to be processed, the given state machine state and context are retrieved and combined with selected packet header information and provided to an allocated hardware state machine to determine the next state and appropriate output. Various state machines in examples according to the present application are described below. The BE1-2 portion 2204 includes a requester instantiation 2220, a completer instantiation 2222 with a request table 2224, and a state machine state and context 2226. The request table 2224 includes an allocated amount of resources provided for out-of-order tracking.
[0179] It should be appreciated that this combination of using a pool of hardware state machines with state and context data and packet header data is one example of state machine design and use. Other state machine designs can be utilized in light of the teachings provided herein, such as using firmware with state, context, and header information, more fixed designs that do not use a pool of hardware state machines, pipelining of hardware state machines, etc.
[0180] The VM2 block 284 includes a BE2-1 portion 2230. The BE2-1 portion 2230 includes a requester instantiation 2232, a completer instantiation 2234 with a request table 2236, and a state machine state and context 2238.
[0181] The existence of the three BE portions BE1-1 2202, BE1-2 2204, and BE2-1 2230 illustrates that a BE portion is developed for each BE RDMA connection.
[0182] While the header processing module 261 is shown as having a separate processor 252 and RAM 254, in some examples, the processor 286 and on-chip RAM 263 can be used. Thus, in an integrated example, FIG. 2B The illustrated physical separation becomes logical separation of the processor, RAM, and firmware. The header hardware 255 remains unchanged and is only accessible by the header processing module 261. The tradeoff is program development time versus silicon cost of the processor 252 and RAM 254.
[0183] FIG. 2CFigure 2 illustrates the initialization of an RDMA connection and the operation of a READ operation using a BE RDMA NIC, such as the NIC 202 described above. In operation 2300, host 1 and host 2 communicate to determine the parameters of the RDMA connection, such as MAC / IP addresses, buffer keys, etc. These values form part of the RDMA connection context. In operation 2302, host 1 configures the connected RDMA NIC 1 to include the appropriate functionality and queues. In FIG. 2D This operation is illustrated in more detail in Figure 23. In operation 2304, host 2 configures the connected BE RDMA NIC 2. In operation 2306, an application in host 1 places an RDMA READ WQE in the SQ, such that RDMA NIC 1 is the initiator and RDMA NIC 2 will be the target. The WQE includes the RDMA opcode and buffer locations. For SEND, the initiator local gather buffer is included in the WQE, as the target remote buffer will be developed by a WQE in the RQ of the remote or target node. For RDMA WRITE, both the initiator local gather buffer and the target remote buffer are included in the WQE. For RDMA READ, both the initiator local scatter buffer and the target remote buffer are included in the WQE. In operation 2308, the application in host 1 pulls the RDMA READ WQE from the SQ by ringing the doorbell of BE RDMA NIC 1. In operation 2310, BE RDMA NIC 1 pulls the RDMA READ WQE from the SQ. The requester, such as the requester 2206 of BE 1-1, obtains the RDMA READ WQE and parses the WQE to configure the requester and completer for the read operation. This parsing includes determining the RDMA opcode, the buffers applicable to the opcode, the transfer length, and any data or immediate data. For an egress message, the requester uses the data provided by the WQE and the connection context to set up the header processing. For example, the transfer length can be used to program the header processing to break the transfer into FIRST, MIDDLE, and LAST values used in the BTH opcode field. For SEND and RDMA WRITE, the requester constructs the DMA commands to be used to transfer the data. The completer, such as the completer 2212, is configured to generate and provide a CQ CQE when the message is completed. In operation 2312, RDMA NIC 1 sends the RDMA READ packet to RDMA NIC 2.
[0184] RDMA NIC 2 receives the RDMA READ packet. In the case of a message that requires a response, such as the example RDMA READ, the completer, such as completer 2212, can develop an RSPQ WQE that includes the necessary information to cause the read operation, such as the known opcode, message number, buffer and key carried in the RDMA READ request packet. In operation 2314, the completer places the RSPQ-WQE in the RSPQ. The RDMA READ target or destination QP and other information in the RSPQ-WQE is used by the responder, such as responder 2208, to obtain the appropriate connection context for configuring the DMA controller and header processing. If this is a SEND or RDMA WRITE operation, the completer will use the QP and opcode in the initial packet of the message to obtain the appropriate connection context. For a SEND, the completer obtains the relevant WQE from the RQ based on the message number in the SEND packet. With the WQE, the completer can program the DMA controller to place the incoming data in the appropriate buffer. For a RDMA WRITE, the completer uses the buffer in the received RDMA WRITE packet to program the DMA controller to transfer the incoming data to the appropriate buffer.
[0185] In operation 2316, the responder of BE RDMA NIC 2, such as responder 2208, obtains the RSPQ-WQE just placed in the RSPQ in a similar manner to how the requester processes WQEs from the SQ and programs the DMA controller to pull the requested data from host 2 memory based on the buffer in the WQE. The responder also configures header processing to properly develop the packet headers. In operation 2318, RDMA NIC 2 sends the requested data in a series of RDMA READ RESPONSE packets. In operation 2320, RDMA NIC 1 receives the RDMA READ RESPONSE packets. The completer programs the DMA controller using the packet header information to transfer the RDMA READ data to host 1 memory. In operation 2322, when the response is complete, the completer develops a CQE to be placed in the CQ and the RDMA READ operation is complete.
[0186] While the above example has described message level operations as being performed using firmware requesters, responders, and completers, in other examples, message level operations can also be configured into hardware associated with hardware state machines to perform the same operations.
[0187] In FIG. 2DIn the middle, the operation of the CreateQP() Verb is illustrated. In the example according to the application, the CreateQP() Verb has a number of possible arguments. An example syntax is CreateQP (ReqMode = FULL, 1P, 2P or NA, SQ = 0 or 1, RQ = 0 or 1, RSPQ = 0 or 1, REQT = 0 or 1, RSPT = 0 or 1). ReqMode = FULL indicates a full requester that supports both single-phase transactions and two-phase transactions. ReqMode = 1P indicates a 1P requester that supports only single-phase transactions. ReqMode = 2P indicates a 2P requester that supports only two-phase transactions. ReqMode = NA indicates no requester support (SQ must not be instantiated). SQ = 1 indicates that an SQ is to be created. RQ = 1 indicates that an RQ is to be created. RSPQ = 1 indicates that an RSPQ is to be created. REQT = 1 indicates that out-of-order inbound request packets need to be tracked. RSPT = 1 indicates that out-of-order inbound response packets need to be tracked.
[0188] CreateQP() 2350 begins at step 2367 where it is determined whether RQ, SQ, and RSPQ are all zero. This is an improper condition, so if true, the operation proceeds to step 2365 where the operation ends and an error condition is returned. If there is no error in step 2367, then in step 2369 it is determined whether both REQT and RSPT are equal to zero. This is an error condition, and if satisfied, the operation proceeds to step 2365. If there is no error in step 2369, then in step 2371 it is determined whether SQ is equal to 1 and ReqMode is NA. This is an error condition, and if satisfied, the operation proceeds to step 2365. If there is no error in step 2371, then in step 2352 the SQ value is evaluated and a determination is made as to whether SQ is equal to 1. If so, then in step 2354 SQ is created and associated with a CQ. Each newly created SQ or RQ must be associated with a CQ, which can be a pre-existing CQ or a newly created CQ for this purpose. In step 2356, the NIC 202 is instructed to instantiate the requester according to the ReqMode value, and to allocate connection and packet state machines, and to set the connection context using the instantiated requester. After step 2356, or if SQ = 0 in step 2352, then in step 2358 the RSPQ value is evaluated and a determination is made as to whether RSPQ is equal to 1. If so, then in step 2360 the NIC 202 is instructed to instantiate the responder, and to allocate connection and packet state machines (if not already done), and to set the connection context using the instantiated responder (if not already done). After step 2360, or if RSPQ = 0 in step 2358, then in step 2362 the RQ value is evaluated and a determination is made as to whether RQ is equal to 1. If so, then in step 2364 RQ is created and associated with a CQ, as with SQ. In step 2366, the NIC 202 is instructed to allocate connection and packet state machines, and to set the connection context using instances of the requester and responder (if not already done). After step 2366, or if RQ = 0 in step 2362, then in step 2368 the NIC 202 is instructed to instantiate the completer, and to allocate connection and packet state machines (if not already done), and to set the connection context using an instance of the completer (if not already done).
[0189] After step 2368, in step 2370, the REQT value is evaluated and a determination is made as to whether REQT equals 1. If so, in step 2372, the NIC 202 is instructed to activate request tracking in the completer. After step 2372, or if in step 2370, RSPQ = 0, in step 2374, the RSPT value is evaluated and a determination is made as to whether RSPT equals 1. If so, in step 2376, the NIC 202 is instructed to activate response tracking in the completer. After step 2376, or if in step 2374, RSPT = 0, the operation completes at step 2378 with a return of a success response. It will be appreciated that this is one example flowchart and other flowcharts providing the same end functionality can be readily developed.
[0190] Extended BE Transfers Full or extended BE transfers extend the baseline BE transfer with the following transfer features: Transfer Roles 1. Decoupled Transfer Roles: The functionality of the BE transfer can be split into multiple distinct roles: the completer, the requester, and the responder, as shown in FIG. 3 .
[0191] a. Requester Role: Handles the transfer functionality required to execute the SQ WQE for the BE RDMA connection. Note that the SQ-WQE held in the SQ is produced by an external entity to the BE RDMA NIC (or the BE transfer IP block integrated into a larger SoC), such as FIG. 2C Host 1 in issues the SQ-WQE to the BE RDMA NIC (or the SoC integrated accelerator / GPGPU issues the SQ-WQE to the integrated BE transfer IP). The requester role duties are: i. Generate the outbound request message for all transaction types (source of request traffic).
[0192] ii. Split the outbound request message into request packets with unique PIDs and track all outstanding request packets based on the signaling received for the request traffic.
[0193] iii. Handle retransmission of request packets based on the signaling received for the request traffic, i.e., track the status of each packet belonging to the outbound request message (Pkt-Rcvd, Pkt-Lost, Pkt-Dlvr, etc.).
[0194] iv. Handle the retirement of the SQ WQE that has been completely delivered and placed (i.e., popped from the SQ) when all packets forming the outbound request message reach the Pkt-Dlvr state.
[0195] v. For single and two-phase transactions initiated by the requester, once the transaction reaches the Xact Cmpl state, the local completer is instructed to complete (generate the associated CQE) the SQ WQE associated with the transaction.
[0196] (a). In single-phase transactions, the CQE can be generated as soon as the associated SQ-WQE has been retired from the SQ.
[0197] (b). In two-phase transactions, the CQE can be generated as soon as the associated SQ-WQE has been retired from the SQ and the associated inbound response message has been completely delivered.
[0198] vi. The requester role functionality can be split into two capabilities: (a). Single-phase requester capability: enables support for single-phase transactions.
[0199] (b). Two-phase requester capability: enables support for two-phase transactions.
[0200] b. Responder role: handles the transport functionality required to execute the RSPQ-WQE for the BE connection. Note that the RSPQ-WQE saved in the RSPQ is generated by the completer role for the same BE connection in response to the request message of a two-phase transaction received. The responder role duties are: i. Generate the outbound response message (source of response traffic) for two-phase transactions.
[0201] ii. Split the outbound response message into response packets with unique PIDs and keep track of all outstanding response packets based on the signaling of the response traffic received.
[0202] iii. Handle retransmission of response packets based on the signaling of the response traffic received, i.e. keep track of the status of each packet (Pkt-Rcvd, Pkt-Lost, Pkt-Dlvr, etc.) belonging to the outbound response message.
[0203] iv. Handle the retirement of the RSPQ-WQE (i.e. pop the completed WQE from the RSPQ) that has been completely delivered and placed when all packets forming the outbound response message reach the Pkt-Dlvr state.
[0204] c. Completer role: handles the reception of the request and response traffic and the reliability protocol signaling functionality (this implements the BE need for out-of-order reception). The completer role duties are: i. Reassemble the inbound request / response messages from the request / response packets received.
[0205] ii. Generate the signaling packets (according to the reliability protocol used) for the inbound request / response traffic.
[0206] iii. Handling execution of one-sided single-phase / two-phase transactions at the target node (bring the inbound request message to the Msg-Exec state).
[0207] (a). Note that for two-phase transactions, generating and submitting the RSPQ-WQE is the action required at the target node to complete execution of the inbound request message.
[0208] iv. Handling execution of two-sided single-phase transactions at the target node (bring the inbound request message to the Msg-Exec state) and completion (generate CQE), i.e. local completion of the inbound request message is placed and generated.
[0209] v. Handling placement of one-sided two-phase transactions at the initiator node, i.e. placement of the inbound response message.
[0210] vi. Handling completion (generate CQE) of one-sided single-phase / two-phase transactions at the initiator node at the direction of the local requestor, and handling completion (generate CQE) of two-sided single-phase transactions at the initiator node at the direction of the local requestor.
[0211] (a). Single-phase transactions can be completed immediately at the direction of the requestor.
[0212] (b). Two-phase transactions can only be completed at the direction of the requestor and after the associated inbound response message has been fully placed.
[0213] vii. The completer role functionality can be split into two capabilities: (a). Inbound request tracker: enables support for request traffic signaling for single-phase and two-phase transactions.
[0214] (b). Inbound response tracker: enables support for response traffic signaling for two-phase transactions.
[0215] 2. Full or extended BE transport can include an optional enhanced Verb interface that allows the end user to explicitly specify which of the 15 node configurations for BE connections listed in Table 1 are supported on each created BE connection (see FIG. 4 and FIG. 5). Table 1 is intended to convey resource requirements ('=0' indicates no resource is required, '=1' indicates a resource is required). An extension has been added to allow the requestor mode to be specified: ReqMode: enum{FULL, 1P, 2P, NA}. ReqMode = FULL indicates a full requestor that supports both single phase transactions and two phase transactions. ReqMode = 1P indicates a 1P requestor that only supports single phase transactions. ReqMode = 2P indicates a 2P requestor that only supports two phase transactions. ReqMode = NA indicates no requestor support (SQ must not be instantiated). Table 1 does not imply the actual Verb API implementation.
[0216] Table 1 - BE Connection Configuration Enumerations
[0217] Referring to FIG. 3 , the CreateQP() command to produce a full capability would be CreateQP(ReqMode=Full, SQ=1, RQ=1, RSPQ=1, REQT=1, RSPT=1) provided by the host to the BE RDMA NIC at each peer node in the connection, which would result in a full completer, responder, and full requestor. FIG. 4 A simplified 1P requestor, request completer, bidirectional fabric would be formed using the CreateQP (ReqMode=1P, SQ=1, RQ=1, RSPQ=0, REQT=1, RSPT=0) command provided by each host. For a 2P requestor, full completer, bidirectional, the CreateQP(ReqMode=2P, SQ=1, RQ=1, RSPQ=0, REQT=1, RSPT=1) command at node A, and for a 1P requestor, inbound request only completer, responder, the CreateQP (ReqMode=1P, SQ=1, RQ=0, RSPQ=1, REQT=1, RSPT=0) command at node B, would produce FIG. 5 A fabric. According to FIG. 2D , using the enhanced CreateQP() Verb and operation, full flexibility in instantiating roles in the BE RDMA NIC is provided.
[0218] Resources are allocated on demand, as FIG. 2D illustrated. During BE connection setup, as FIG. 2CAs shown in step 2300, two BE-capable nodes need to communicate and negotiate the resources required for the BE connection (note that BE connection establishment can be in-band, i.e., handled by the RoCE Communication Manager (CM) interacting with the peer node via QP1, or out-of-band, i.e., setting up the BE connection through some alternative method, such as using LAN UDP packet switching).
[0219] a. The requester role is optional. When needed, there are three types to choose from, depending on the need. In any case, SQ resources are needed to support the requester role.
[0220] i. 1P requester: Enables support for initiating one-phase transactions (e.g., Send, RDMA WRITE) by transmitting request messages.
[0221] ii. 2P requester: Enables support for initiating two-phase transactions (e.g., RDMA READ / ATOMIC) by transmitting request messages.
[0222] iii. Full requester: Enables support for initiating both one-phase and two-phase transactions by transmitting request messages.
[0223] b. The responder role is optional. When needed, RSPQ is needed to transmit response messages for two-phase response transactions.
[0224] c. The completer role is mandatory. There are three types of completer. In any case, at least one CQ and one or two inbound tracking resources are needed to support the completer role. When a BE node is set up to accept inbound bi-directional transactions, RQ resources are also needed.
[0225] i. Inbound request-only completer: Supports only local completion, remote completion for bi-directional transactions, and inbound request tracking and its associated outbound signaling.
[0226] ii. Inbound response-only completer: Supports only local completion, and inbound response tracking and its associated outbound signaling.
[0227] iii. Full completer: Supports local completion, remote completion for bi-directional transactions, inbound request tracking, inbound response tracking, and their associated outbound signaling.
[0228] d. REQT: Inbound request tracking resource.
[0229] e. RSPT: Inbound response tracking resource.
[0230] Enhancing the operation of the CreateQP() command allows full control over the functionality of a specified RDMA connection, allowing each RDMA connection to be optimized to minimize the NIC resources used.
[0231] Flow interleaving Full or extended BE transport can include BE connection flow interleaving, where BE transport interleaves request and response flow transmissions on the same connection protected by the same reliability protocol. This interleaving is only supported if the local node's connection is set to include both the requester and responder roles. BE transport supports the following interleaving options: 1. Message boundary interleaving (legacy): This interleaving mode is taken from iWARP. FIG. 12A An example is provided. It leverages the shared PID space between the two flows. The PID format for this mode is equivalent to the legacy PSN (Packet Sequence Number), thus only a single packet stream needs to be tracked at the receiver of the peer node. Note that in this mode, packets belonging to the bilateral single-phase request message and unilateral two-phase response message will need to carry the IMPN field to enable out-of-order placement (unilateral single-phase / two-phase request messages will not require the IMPN).
[0232] 2. Packet boundary interleaving (extended): This interleaving mode allows one request flow message and one response flow message to flow out simultaneously by interleaving the request and response flow messages at the packet level while sharing the PID space. FIG. 12B An example is provided. This mode requires higher complexity at the receiver node of the flow as it needs to track two interleaved packet streams on the same PID space and de-multiplex them. Given that the PID space is shared, the cumulative signaling packets (i.e., ACK and NAK) can carry independent PIDs for the request and response flows. The PID format for this mode requires three data fields: a. MID: Message Identifier. Each new message to be transmitted (whether a request or response message) should be assigned the next immediate MID. (MID is incremental).
[0233] b. IMPN: In-Message Packet Number.
[0234] i. For valid packets, IMPN must be less than PKTCOUNT.
[0235] c. PKTCOUNT: Message Packet Count of the message.
[0236] Note that in this mode, packets will not need to carry additional control information to enable out-of-order placement (i.e., bilateral single-phase request message and unilateral two-phase response message).
[0237] Note that in alternative examples, this mode can be modified by removing the PKTCOUNT from each request packet and including the PKTCOUNT only in the first packet of the message (i.e., IMPN = 0). This would result in a bandwidth saving, but can imply a longer latency to detect the loss of the last packet of the message (i.e., IMPN = PKTCOUNT - 1) when the first packet is lost (e.g., the device needs to request retransmission of the first packet of the message in order to figure out the packet count of the message, and thus eventually verify whether the last packet of the message was dropped).
[0238] 3. Independent PID space: This interleaving mode allows one request traffic message and one response traffic message to flow out simultaneously by interleaving the request traffic messages and the response traffic messages at the packet level, while each traffic uses an independent PID space. FIG. 12C Examples are provided. This replicates the state tracking required when both traffic (request traffic and response traffic) are supported in a node, either as a sender or as a receiver. Given that the PID space is independent, the cumulative signaling packets (i.e., ACK and NAK) must carry independent PIDs for the request traffic and the response traffic. The PID format for this mode utilizes the PSN and infers the FlowID from the BTH.OPCODE: PID: {FlowID, PSN}, where the FlowID can be inferred from the BTH.OPCODE, as the selected opcode is a request-only opcode, while the selected opcode is a response-only opcode.
[0239] Note that a single (selected) interleaving option should be used in any particular BE connection. In some examples, the interleaving option can differ between the various BE connections of a BE RDMA NIC, while in other examples, the interleaving option is the same for all the individual BE connections of a BE RDMA NIC.
[0240] By using additional header fields and alternative meanings of other header fields, the interleaving of request packets and response packets can be done at the packet level (rather than the message level), thus reducing the blocking latency between the outbound request traffic messages and the outbound response traffic messages.
[0241] Reliability protocol The full or extended BE transport includes a new reliability protocol based on four signaling packets (feedback from the receiver to the sender), namely, acknowledgment (ACK), receipt acknowledgment (RACK), negative acknowledgment (NAK), and selective negative acknowledgment (SNAK): 1. NAK is traditional, and signals implicit (cumulative) delivery acknowledgement to GoBackN, which is for one PID before the indicated PID. NAK can only be sent once (and relies on RTO if lost).
[0242] 2. SNAK signals a hole (packet loss) in the received PID stream, i.e., transitions to Pkt-Lost state, without any implicit acknowledgement context, and can be sent multiple times to prevent RTO. Pkt-Lost state implies that the associated PID is eligible for retransmission.
[0243] a. SNAK packet can carry a limited number of PID fields, which can be flexibly allocated to indicate multiple disjoint holes in the PID stream. SNAK packet can signal each hole in the following three modes: point hole (single PID hole, which consumes one PID field in the SNAK packet), range hole (multiple contiguous PIDs in the hole, which consumes two PID fields in the SNAK packet), and infinite hole, which represents a range hole of a left-closed right-open interval (consumes the PID / PSN field of BTH to indicate the left-closed boundary of the infinite hole, i.e., the first PID in the infinite hole, and another field explicitly indicates that this PID field should be treated as the infinite hole start PID), and consumes a single PID field in the SNAK packet. Note that this document uses the term "non-infinite hole" to refer to a hole of point hole type or range hole type.
[0244] i. SNAK must indicate a single latest hole in the PID stream when the receiver is in ACTIVE superstate, and must indicate an infinite hole when in RECOVERY superstate. The receiver can include as many older holes in the PID stream as possible in the SNAK. The number of holes indicated in the SNAK packet, in addition to the single latest hole in ACTIVE superstate or the infinite hole in RECOVERY superstate, can be set based on the policy present in the BE RDMA NIC.
[0245] ii. The sender can selectively retransmit packets associated with a new non-infinite hole in the PID stream upon receiving a SNAK for that hole.
[0246] b. SNAK with infinite hole indication signals to the receiver that the out-of-order (OOO) tracking resources for its inbound (request or response) traffic have been exhausted. The out-of-order tracking resources are limited because the SOC 201 has limited memory that can be allocated for this purpose.
[0247] i. When the receiver's out-of-order tracking resources are exhausted, the receiver must send a SNAK with an infinite hole, indicating the first PID of the infinite hole.
[0248] ii. The sender must stop tracking PIDs within the infinite hole (as these PIDs are not tracked by the receiver and would cause packets to be discarded if received by the receiver): (a). New PIDs (i.e., traffic that has not been transmitted before) always belong to the infinite hole and thus should not be eligible for first transmission while the infinite hole exists.
[0249] (b). Previously transmitted PIDs that are within the infinite hole should not be tracked in the Pkt-Lost state as they should not be eligible for retransmission.
[0250] c. A SNAK packet can signal a point hole, a range hole, an infinite hole, or a mix of them, i.e., the SNAK packet fields can be flexibly assigned for any case.
[0251] i. Note that the same PID can be signaled differently in SNAKs at different times, e.g., it can first be signaled as a point hole, then as either endpoint of a range hole, or be implicitly covered by a range hole (without being either endpoint of the range hole).
[0252] 3. ACK is conventional and signals multi-packet (cumulative) in-order delivery, i.e., transition to Pkt-Dlvr state (which implies reception, i.e., first implicit transition to Pkt-Rcvd state, followed by immediate transition to Pkt-Dlvr state).
[0253] 4. RACK signals multi-packet (non-cumulative) out-of-order reception, i.e., only transition to Pkt-Rcvd state.
[0254] a. RACK should signal an OOO-received packet once and only once, and with a byte count greater than zero. RACK can repeatedly signal an OOO-received packet, but the repeated signaling should have a byte count of zero.
[0255] b. RACK indicates the amount of bytes that have been OOO-received, which informs the sender about the amount of traffic that has terminated in-flight (i.e., has left the network).
[0256] i. The sender's congestion control (CC) can use the RACK information to enable the transmission to proceed forward (i.e., inform the CC about packets that have successfully left the network, e.g., to supplement the congestion window). There are no requirements on the CC mechanism.
[0257] c.RACK signaling is best effort (may be lost in the network). Even when RACK can be lost, once the previous hole is delivered, the final ACK will implicitly signal the same receipt (since delivery implies receipt). This guarantees that the sender will (eventually) receive notification of all packets that left the network, even if RACK is lost.
[0258] i.If the sender's CC is configured to use byte count of RACK (e.g., to supplement the congestion window), then it must also use the ACK information to prevent congestion window leakage due to RACK loss, while avoiding duplicating the congestion window supplementation.
[0259] d.RACK packets can optionally indicate multiple OOO-received PIDs (a larger packet is needed to carry this information). Again, this option is connection-based, not per-packet. There are two modes of encoding PIDs in RACK packets: point PID receipt (a single received packet, consuming one PID field in the RACK packet) or range PID receipt (multiple consecutive received packets, consuming two PID fields in the RACK packet). Note that a single RACK packet can provide a mix of modes, e.g., the PID fields for OOO-received packets can be flexibly allocated for either case.
[0260] i.The PIDs carried by RACK allow the sender to track which packets have been OOO-received by the receiver. At the sender side, this optional feature has two (non-mandatory) uses: (a). By transitioning the tracked PID's state from Pkt-Lost to Pkt-Rcvd, optimize selective retransmission (stop retransmitting a packet once it has been RACKed), ignoring later received SNAKs with non-infinite holes that cover the same PID (these can be protocol errors or network traffic reordering issues).
[0261] (b). Allow updating the tracked PID of the tail packet, i.e., if the old tail packet's PID has been RACKed and the previous PID has been SNAKed, then the SNAKed PID becomes the new tail packet's PID.
[0262] ii.The multi-packet nature of RACK allows the receiver (the sender of RACK packets) to consolidate OOO-received packets (if desired) before sending RACK, in order to reduce the bandwidth consumption for signaling.
[0263] e.The sender is free to decide when to selectively retransmit packets that are signaled (SNAKed) as holes in the PID stream. Some options are provided as examples: i. The sender decides to retransmit each signaled hole upon reception of each SNAK packet.
[0264] (a). This option does not require any new state tracking at the sender side.
[0265] (b). This option provides the highest possible redundancy that can be driven by protocol signaling only, increasing the likelihood of hole delivery in the PID stream.
[0266] (c). This option leads to high likelihood of bandwidth waste due to repeated retransmission of the same packet.
[0267] ii. The sender can decide to retransmit the signaled holes received in the SNAK packet up to a configurable number of times per RTT.
[0268] (a). This option uses a new state to track the PIDs at the sender side that have outstanding retransmissions.
[0269] (b). The elapsed RTT interval can be approximated by receiving a signaling packet (ACK / RACK / SNAK) of the highest transmitted PID before the first retransmission of the hole.
[0270] (i). This option provides limited redundancy driven by protocol signaling only.
[0271] (ii). This option provides an upper bound on the wasted bandwidth through repeated retransmission of the same packet.
[0272] (c). The elapsed RTT interval can be approximated by using a timer or a time wheel.
[0273] (i). This option provides deterministic redundancy independent of protocol signaling.
[0274] (ii). This option provides deterministic bandwidth utilization through repeated retransmission of the same packet.
[0275] (iii). This option requires additional time to maintain state or mechanisms.
[0276] f. Optionally, the sender and receiver can negotiate the number of tracking resources (hole tracking at the receiver, retransmission tracking at the sender) at each peer in order to choose the best packet format for RACK and SNAK (carrying enough PIDs but not more than the range that can be utilized at both ends).
[0277] Packet format Examples of packet formats for request packets and response packets as well as reliability signaling packets are provided in FIGS. 10A-10F and FIGS. 11A-11F FIGS. 12A-12C is an example packet containing various header field related values to illustrate the interleave option.
[0278] FIG. 10 The details of the Base Transport Header (BTH) are illustrated in
[0279] Referring to FIGS. 10A-10F and FIGS. 11A-11C a legacy or message level interleave header is illustrated. Referring to FIG. 10A three new headers are provided: IMETH, RQETH, and RSPQETH. IMETH (In Message Extended Transport Header) is required for each multi-packet message packet whose current header fields cannot provide the "offset" hint for a certain packet within the message, and carries the IMPN. RQETH (Receive Queue Extended Transport Header) is required for each operation that needs to consume an RQ WQE on the target (such as SEND packets and RDMA WRITE packets with IMMEDIATE), and carries the RQMSN. RSPQETH (Response Queue Extended Transport Header) is required for any request message that needs a response (e.g., RDMA READ packets and ATOMIC packets), and carries the RSPQMSN. In addition, RETH (RDMA Extended Transport Header) is required for each RDMA WRITE and RDMA READ request packet, and carries the virtual address, remote key, and DMA length of the RDMA operation.
[0280] FIG. 10B Three variants of the SEND packet are illustrated. The RQETH header is required in all three variants, but the IMETH header is optional for single-packet messages. Again, this option is connection based, not per-packet. FIG. 10C and FIG. 10D Two variants of the RDMA WRITE packet are illustrated. FIG. 10C The packet in does not include the IMETH header, while FIG. 10D The packet in includes the IMETH header. The use of the IMETH header is optional. As mentioned above, the IMETH header is optional in single-packet messages. Again, this option is connection based, not per-packet. FIG. 10E RDMA READ request and RDMA READ response packets are illustrated. The RSPQETH header is used to provide the RSPQMSN, which is used to match the response packet with the request packet. FIG. 10FThe ATOMIC request and response packets are illustrated. Again, the RSPQETH header is used here to provide the RSPQMSN, which is used to match the response packet to the request packet.
[0281] Reference is made to FIG. 11A , three new headers used in the reliability signaling packets are illustrated: RAETH, RAPETH, and SNETH. The RAETH (Receive ACK Extended Transport Header) is required for RACK. It carries the mandatory Ack-Byte-Count and up to a single OOO-received Point PID in the PSN field of the BTH (this should be negotiated during connection setup). The Ack-Byte-Count indicates the amount of payload bytes that have been OOO-received at the receiver node. This byte count can combine the payloads of multiple OOO-received packets, typically received consecutively in bursts. The RAPETH (Receive ACK PID Extended Transport Header) is a connection-based optional header that carries the extended OOO-received PID(s). Its optional use and size are determined during connection setup. The header length is 4*N bytes, where N is the number of OOO-received PID(s) carried. Each Status(n) [3-0] field provides {Valid(l), Rsvd(l), PIDType(2)}, where PIDType: enum{POINT, RANGE-START, RANGE-END}. The statusO field relates to the PID carried in the BTH field of the header. The SNETH (Selective ACK Extended Transport Header) is required for RACK. Its size is determined during connection setup. The header length is 4*N bytes, where N is the number of hole PID(s) carried. Each Status(n) [3-0] field provides {Valid(l), Rsvd(l), PIDType(2)}, where PIDType: enum{POINT, RANGE-START, RANGE-END, INFINITE-START} for the associated PID. The statusO field relates to the PID carried in the BTH field of the header.
[0282] FIG. 11A The regular ACK / NAK packets are also illustrated. FIG. 11B Two versions of the RACK are illustrated. The upper version provides single packet acknowledgments, while the lower version is used to acknowledge a series of packets (including either point OOO-received PIDs or range OOO-received PIDs). FIG. 11CSNAK is illustrated. SNAK always indicates the latest hole as well as as many previous holes or ranges as needed. In most examples, SNAK is also used to indicate an infinite hole, whose PID value is the first packet of the infinite hole. In some examples, a different dedicated packet type (other than SNAK) can be used to indicate an infinite hole, such as the value of IHAK (infinite hole acknowledgment) along with the relevant start PID.
[0283] Reference FIGS. 11D-11F , illustrating header changes for packet level interleaving. A new header, RQPETH, is utilized. RQPETH (Receive Queue Packet Level Interleaving Extended Transport Header) is mandatory for any multi-packet message carrying data (whether request or response traffic), whose packets do not already include message size information, e.g., SEND packets of a multi-packet SEND message. For single-packet SEND messages, the header is optional. Again, this option is connection-based, not per-packet. It carries a 24-bit PKTCOUNT field to provide message size in number of packets. In addition, the meaning of fields in the existing headers has changed. PID becomes {MID, IMPN}, which is 48 bits instead of 24 bits. MID is a 24-bit message ID, which is carried in BTH.PSN. IMPN is a 24-bit intra-message packet ID, which is carried in IMETH. IMETH is required for all packets of a multi-packet message. For packets of a single-packet message (e.g., RDMA READ request), IMETH is optional. When it is not present, PID is implied to be {MID, 0}. Operations that do not include message size or length information in pre-existing headers (such as SEND packets of a multi-packet SEND message) must carry RQPETH. All signaling packets require IMETH to correctly indicate the first PID signaled as {BTH.PSN, IMETH.IMPN}. RAPETH and SNETH grow with the PID from 24 bits to 48 bits.
[0284] FIG. 11E Changes to the three types of SEND packets are illustrated. The inclusion of the RQPETH header is also evident in comparison to FIG. 10B . FIG. 11F Changes to the ACK / NAK and SNAK packets are illustrated. Both include the IMETH header to provide packet numbering in the message, as the PSN field is message numbering and does not change throughout the message. SNETH has grown to 49 bytes to incorporate the change of the PID value to {MID, IMPN}. FIG. 11G Two types of RACK are illustrated. Each includes IMETH to provide packet numbering. RAPETH has grown to 49 bytes to reflect the change in PID value.
[0285] FIGS. 12A-12C An example is provided from the perspective of node A of an outgoing SEND message, an incoming RDMA READ request message and its associated outgoing RDMA READ response messages, and finally a second outgoing SEND message. FIG. 12A is message level interleaved, so the second SEND message must be delayed until after the RDMA READ response messages are complete. FIG. 12B is packet level interleaved, and the second SEND message is moved after the first RDMA READ RESPONSE packet, and the RDMA READ RESPONSE messages are completed after the second SEND message. FIG. 12C uses separate PID space, so the second SEND message is again moved after the first RDMA READ RESPONSE packet.
[0286] In more detail FIG. 12A , the first operation is a two packet SEND. For the first packet, the PSN (which represents the full PID) is 0154A7, the IMPN is 0, and the RQMSN is 421537. The second packet has a PID of 0154A8, the IMPN is 1, both of which are incremented from the first packet, with the RQMSN remaining the same. The second operation is an RDMA READ, with the RDMA READ request packet having a PID of 5D793F, a DMA Length of 1000, and an RSPQMSN of 75873E. The next four packets are the RDMA READ response packets. The PIDs start at 0154A9 and increment to 0154AC. The IMPN values start at 0 and increment to 3. The RSPQMSN for all four response packets is 75873E. The RSPQMSN allows the requester to match the responses to the request. For the sender, the PIDs are incremented in the normal flow, incrementing with each packet. The final operation is a SEND ONLY. The PID is 0154AD, incremented from the last PID of the READ response. The RQMSN is 421538, incremented from the previous SEND message. This illustrates the PID, IMPN, and MSN operations in message level interleaving.
[0287] Referring to FIG. 12B , packet level interleaving is shown. The first packet of the first SEND operation uses the same PID as the first packet of the FIG. 12AExample identical PSN (MID component of PID) IMPN and RQMSN values, and note that the first packet's PID is given by {PSN, IMPN} ({0154A7, 0}). The PKTCOUNT value of 2 is added to the packet's RQPETH header. For the second packet of the SEND operation, the PSN or MID value remains unchanged (not incremented), but the IMPN value is incremented to 1. The other values remain unchanged. Next, the RDMA READ request is received, and this request is identical to the one in FIG. 12A . The next packet is the first packet of the RDMA READ RESPONSE. The PSN or MID is incremented to 0154A8, its IMPN value is 0, and the RSPQMSN is the RSPQMSN value of the RDMA READ request. To illustrate packet interleaving, next the second SEND operation occurs. The PSN or MID value is incremented to 0154A9, and the RQMSN is identical to FIG. 12C . The remaining three packets of the RDMA READ response follow the second SEND. For each packet, the PSN or MID remains 0154A8, and the IMPN value is incremented. The use of the same PID for multiple packet operations, along with the incrementing of the IMPN, allows the receiver to link the packets, even though they are separated, and allows detection of packet loss.
[0288] FIG. 12A Figure illustrates operations using different PID spaces. The packets of the initial SEND operation are identical to those in FIG. 12A . Similarly, the RDMA READ request is identical to that in FIG. 13A . The first packet of the RDMA READ response has a PID of F4903A, rather than a PID value that is consecutive to the PID sequence from the other packets from the sender, as these packets belong to response traffic that uses a separate PID space. Next is the second SEND operation, and uses the PID value 0154A9, which is incremented from the second packet of the first SEND operation. Next are three RDMA READ response packets, whose PIDs are incremented from the F4903A value of the first RDMA READ response packet, and whose IMPN values are also incremented. The use of separate PID spaces for request traffic and response traffic packets removes any confusion that might be caused by any intervening operations from the sender, allowing packet-level interleaving.
[0289] FIG. 13B and FIG. 13A Figure illustrates acknowledgment operations in single-phase and two-phase operations. FIG. 13B Figure illustrates two WRITE messages and one combined ACK, although separate ACKs can be provided. FIG. 6AThe diagram illustrates a READ operation with an initial READ request and its ACK. Two READ responses are provided with ACKs, although a single merged ACK may be provided. These examples provide further indication of the operation of the requester, responder, and completer functions.
[0290] By using SNAK, RACK packets and the additional data provided in those packets (including an indication of out-of-order resource exhaustion at the receiver), the use of additional reliability signaling allows for improved efficiency of high BE RDMA communications by minimizing retransmissions.
[0291] state machine 1. FIG. 6B and FIG. 6A The diagram illustrates a packet-level hierarchical state machine description of the reliability protocol in the sending node. Such a packet-level state machine applies to the sending node for both request and response traffic on the BE connection. For each PacketID 'p' tracked by the sending node, there exists an instance of this state machine (denoted as 'PktSndrState[p]'). The diagrams of the presented packet-level state machine illustrate the superset case, where all optional sending characteristics are included in the BE transmission. In this state machine, the values in the upper half of the ellipse are the state names, while the values in the lower half are the "entry:" statements within the state, indicating the entry action performed when entering that state. The values on the lines connecting the ellipses are the inputs to the state machine, used to cause state transitions. Dashed ellipses represent optional characteristics.
[0292] As an example, FIG. 6A The top central ellipse in the diagram represents the Pkt-Pending state. Upon entering the Pkt-Pending state, a schedule_tx(p) message is provided to indicate that packet p will be transmitted. When the message TX(pid=p) is received (indicating that the sender has scheduled the packet with PID p for transmission), the state machine transitions to the Pkt-Outstanding state. In the Pkt-Outstanding state, if a message indicating NAK has been received, where the packet's PID is less than or equal to p, or if the retransmission timer (RTO) has expired, the state machine returns to the Pkt-Pending state and reschedules packet p for transmission. The symbol || indicates logical OR, while the symbol && indicates logical AND. The symbol ! is the inverse OR NOT symbol; for example, !enable_retx means the condition is true when enable_retx is not set. The Pkt-Pending state is returned from the Pkt-Rcvd, Pkt-Lost, and Pkt-ReTx states for the same reasons as returning from the Pkt-Outstanding state. This is a comparison of... FIG. 6AExplanation of the state machine in the Pkt-Pending state. FIG. 6B , FIG. 7A , FIG. 7B and FIG. 6A The operation of the state machine in the Pkt-Pending state operates in this way and specific transitions and actions are presented in the state machine in the Pkt-Outstanding state and are not further explained here. FIG. 6B , FIG. 7A , FIG. 7B and FIG. 6A The operation of the state machine in the Pkt-Pending state operates in this way and specific transitions and actions are presented in the state machine in the Pkt-Outstanding state and are not further explained here.
[0293] a. Two copies of the state machine are shown, each encapsulated within a different connection superstate value (i.e. ACTIVE or RECOVERY, as described below). Whenever the connection of the sender node transitions between these connection superstate values, each packet-level state machine also transitions while maintaining the same substate it had within the previous connection superstate.
[0294] b. All signaling actions are inputs to the state machine (i.e. signaling packets received by the sender node), namely ‘ACK()’, ‘NAK()’, ‘RACK()’, and ‘SNAK()’.
[0295] c. All transmitted packet notifications (request traffic or response traffic) represent inputs to the state machine (i.e. data packets transmitted by the sender node), labeled as ‘TX()’ in the figure.
[0296] i. Note that the transmission of Flush packets (i.e. ‘TX_FLSH()’) is not controlled by the packet-level hierarchical state machine and is therefore not represented in the figure.
[0297] ii. Note that ‘enable_retx’ represents a Boolean state variable used to enable selective retransmission of packet p, as described above. The update() function represents the algorithm or policy used by the sender to decide when to disable / enable the retransmission schedule of a particular packet (i.e. p).
[0298] iii. Note that ‘schedule_tx()’ represents the action of scheduling a packet for transmission, where the sender transmits the packet after a scheduling delay.
[0299] d. There are four mandatory states, namely ‘Pkt-Pending’, ‘Pkt-Outstanding’, ‘Pkt-Dlvr’, ‘Pkt-Lost’. ‘Pkt-Rcvd’ and ‘Pkt-ReTx’ are optional (in the Pkt-Pending state and the Pkt-Outstanding state, respectively). FIG. 6B and FIG. 7APkt-Rcvd is needed only if RACK is configured to indicate the PID of the OOO received packet. Pkt-ReTx is needed if the sender wants to reduce the frequency of retransmitting the same packet within a short period of time (e.g., one RTT). Note that when Pkt-ReTx is not supported, the packet state will transition from Pkt-Lost to Pkt-Outstanding as soon as the lost packet is retransmitted (i.e., the incoming TX() arrow from the Pkt-Lost state will point to Pkt-Outstanding).
[0300] Pkt-Pending indicates that a packet has been issued to the transmit queue (i.e., SQ for request traffic packets or RSPQ for response traffic packets). Pkt-Outstanding then indicates that the packet has been transmitted, but the sender side has not yet received explicit or implicit signaling for that packet. Pkt-Dlvr means that the packet has been received in-sequence. Pkt-Rcvd means that the packet has been received, but at least one packet with a prior PID value has not been received. Using these states in combination with RACK and SNAK signaling can allow the goal of retransmitting only explicitly known lost packets to be achieved. This reduces the number of packets that are retransmitted compared to other transmissions. In some cases, a packet can be in Pkt-Rcvd, but it will be retransmitted due to an infinite hole being signaled (that infinite hole starts at a PID value prior to the PID value of the packet received OOO).
[0301] e. Note that the argument pid represents the PID field of the transmitted packet (request or response), or the PID field on the received (signaling) packet. For example, i. ACK(pid ≥ p) represents an ACK received for a PID that is greater than or equal to the PID of the packet being tracked by the state machine instance (i.e., p).
[0302] ii. TX(pid = p) represents transmission of the packet being tracked by the state machine instance (i.e., p).
[0303] f. The main difference between the ACTIVE state machine of the sender and the RECOVERY superstate is that in the RECOVERY superstate, no new packets will be transmitted (i.e., no packets in Pkt Pending state will be transmitted), only lost packets will be transmitted (i.e., only packets in Pkt-Lost state will be retransmitted). This is appropriate because the RECOVERY superstate is entered when an infinite hole is signaled, indicating that the receiver can no longer track more packets in Pkt-Lost state whose PID is after the infinite hole (hole). This maximizes the rate of recovery of lost packets.
[0304] 2. FIG. 7B and FIG. 8A A packet-level hierarchical state machine description of the reliability protocol in the receiver node is shown. Such a packet-level state machine applies to the receiver node for both request and response traffic on BE connections. There is one instance of this state machine (referred to as 'PktRcvrState[p]') for each Packet ID 'p' being tracked by the receiver node. The figures presented of the packet-level state machine show the superset case, i.e. where all optional receiver features are included in the BE transport. In this state machine, the upper half of the ovals are the state names, while the lower half of the ovals are the "entry:" statements within the state that represent the entry actions performed upon entering the state. The values on the lines between the ovals are the inputs to the state machine that cause the state machine to transition.
[0305] a. Two copies of the state machine are shown, each encapsulated within a different connection superstate value (i.e. ACTIVE or RECOVERY, see below for a description of the connection superstate). Whenever the connection of the receiver node transitions between these connection superstate values, each packet-level state machine also transitions while maintaining the same substate it had within the previous connection superstate.
[0306] b. All signaling actions are outputs of the state machine (i.e. signaling packets transmitted by the receiver node), represented by 'RequestACK()', 'RequestRACK()', and 'RequestSNAK()'. These outputs should be understood as requests to generate the actual request signal, which does not always result in the generation of a signaling packet(s).
[0307] i. ACKs are cumulative. When multiple packet-level hierarchical state machines track packets with consecutive PID values and request an ACK, only the highest PID will be included in the ACK packet.
[0308] ii. The number of PIDs in a RACK is limited. The receiver will select PIDs from the multiple packet-level hierarchical state machines that are requesting to be sent in a single RACK based on a policy (e.g. latest received PIDs, as many oldest PIDs as possible).
[0309] (a). Note that 'RequestRACK()' has an additional length argument, / , that passes the amount of bytes that have been OOO received.
[0310] iii. The number of PIDs in a SNAK is limited. The receiver will select PIDs from the multiple packet-level hierarchical state machines that are requesting to be sent in a single SNAK based on a policy (e.g. single latest hole, as many oldest holes as possible).
[0311] c. All received message packets (i.e. non-signaling packets) (request traffic or response traffic) represent inputs to the state machine (i.e. data packets received by the receiver node), denoted in the figures as 'RX()'. Received Flush packets represent inputs to the machine, denoted by 'RX_FLSH()'.
[0312] Note that on 'RX_FLSH()', the receiver can optionally signal a RACK with l = 0, which indicates that no new traffic has left the network, but the original tail_pid has been received.
[0313] d. Note that the argument pid represents the PID field of a received packet (request or response), or the requested PID field of a transmitted signaling packet. For example, i. RX(pid > p) represents a received packet with a PID greater than the PID (i.e. p) tracked by the state machine instance.
[0314] ii. RequestACK(pid = p) represents a request for an acknowledgement for a packet (i.e. p) tracked by the state machine instance.
[0315] e. Note that epid is a state variable tracked in the connection superstate. It represents the PID of the oldest unreceived packet.
[0316] f. Note that next_epid is a state variable tracked in the connection superstate. It represents the PID of the second-oldest unreceived packet. As an example, if the most recently delivered PID is pid = 5, then epid = 6 and next_epid = 7.
[0317] g. Note that tail_pid is a state variable tracked in the connection superstate. It represents the PID that the sender node believes to be the tail of the flow.
[0318] h. Note that inf_pid is a state variable tracked in the connection superstate. It represents the PID at which the existing infinite hole begins.
[0319] i. Note that adjust_inf_pid is a Boolean event that informs that inf_pid was updated to the PID of an older received packet due to resource exhaustion.
[0320] 3. Reliability protocol connection superstate: Active superstate and Recovery superstate a. The goal of the Active superstate of a connection is to minimize the impact on traffic throughput by preventing retransmission of OOO received packets.
[0321] i. Allow the sender to inject new traffic into the network while prioritizing retransmission of signaled (SNAKed) holes in the outbound PID stream.
[0322] ii. The receiver signals new SNAKs under the following conditions: (a). Detect a new hole in the inbound PID stream, i.e. when a packet is OOO received, thus there must be a PID hole before the PID of the received packet.
[0323] (b). Fill a hole in the PID stream while leaving other holes, i.e. when a selectively retransmitted packet covers a hole currently tracked by the receiver, and there are more holes remaining to be covered.
[0324] iii. The receiver signals RACK for each point or (contiguous) range of PIDs OOO received.
[0325] iv. The receiver signals ACK for each in-sequence received PID.
[0326] b. The goal of the connected Recovery super-state is to quickly recover from exhausted hole tracking resources and limit the bandwidth wasted on discarded packets (i.e. beyond the infinite hole PIDs, further OOO received packets will be discarded as they cannot be tracked as the PIDs increase).
[0327] i. Do not allow the sender to inject new traffic into the network and only allow retransmission of non-infinite hole PIDs that are signaled (SNAKed).
[0328] ii. During this super-state, the receiver signals SNAK for each received message (non-signaling) packet (merging still possible as an optimization based on the SNAK-based policy).
[0329] (a). If the received PID belongs to the interior of an infinite hole, discard the packet and notify the tracked holes via the emitted SNAK packet, including the infinite hole notification.
[0330] (b). If the received PID belongs to the interior of a previously notified range hole, accept the packet and signal accordingly (ACK or RACK), and emit a SNAK indicating the updated list of tracked holes (including the infinite hole notification).
[0331] (i). Note that this received PID can split a single PID range hole into at most two sub-ranges (each sub-range can be a PID point hole or a range hole), depending on the position of the received PID within the PID range of the range hole.
[0332] (c). If the received PID corresponds to a previously announced PID point hole, then fill in the hole and signal accordingly (ACK or RACK), and issue a SNAK indicating the updated list of tracked holes (if there are any tracked holes left, including the infinite hole notification).
[0333] iii. The receiver signals RACK for each point or (contiguous) range of PIDs received by the OOO receiver that (partially or completely) fill in a previously announced PID point hole or range hole.
[0334] iv. The receiver signals ACK for each point or (contiguous) range of PIDs received in-sequence. The ACK should cumulatively acknowledge all received packets up to the leftmost (oldest) outstanding PID hole.
[0335] In most examples, the transition from the Active superstate to the Recovery superstate is triggered by exhaustion of the out-of-order tracking resources at the receiver, and signaled to the sender with an infinite hole SNAK. In some examples, the sender and receiver can negotiate the available OOO tracking resources, and the sender (rather than the receiver) can track the available resources and provide a change to the recovery state message to initiate the recovery operation for both the sender and receiver, and return to the active state message upon recovery.
[0336] d. The transition from the Recovery superstate to the Active superstate is triggered by delivery (i.e. filling in) of each outstanding PID hole prior to the infinite hole tracked at the receiver, and signaled to the sender with a NAK for the PID at the left edge of the infinite hole. To reduce the likelihood of retransmission timeouts at the sender (in case the NAK is lost), there are the following optimizations: i. Mandatory: The receiver can issue an ACK (at the last outstanding PID hole prior to filling in the infinite hole), followed by a NAK, both cumulatively acknowledging to the same PID.
[0337] ii. Optional: The sender can automatically transition to the Active superstate upon receiving an ACK that cumulatively acknowledges everything prior to the infinite hole.
[0338] It should be understood that the packet-level state machines of the sender and receiver represent logical flow. The specific implementation of the flow can differ from the state machines illustrated, but the logical flow will remain as illustrated.
[0339] While in most examples, the transition from Recovery state to Active state occurs when all holes are filled and full in-order tracking resources are available, in some examples, the transition from Recovery state to Active state can occur when a limited number of holes are available and some in-order tracking resources are still in use, but a sufficient number of tracking resources have been recovered to allow new packet transmission to occur, with the remaining holes to be filled in normal Active operation with a minor impact on transmission rate. This early transition can occur at a settable level, or can occur based on measured recovery rate and expected remaining time to full recovery.
[0340] 4. The BE sender node's connection-level state machine (handling the superstate of the sender) is as shown in FIG. 8B and FIG. 8B . This state machine, together with the packet-level hierarchical state machines, provides the complete view of the BE sender.
[0341] a. If the BE sender supports the optional Flush feature, this state machine is responsible for the Flush feature (transmitting Flush packets, maintaining the Flush timer and Flush counter), since only the connection-level state machine can detect the tail-dropped condition. When a RACK for tail_pid is received while the BE sender is in the tail-dropped condition, the tail_pid must be updated to the hole with the highest PID.
[0342] b. In this state machine: i. The 'entry:' action of each state indicates the action performed by the BE sender each time it transitions to such a state.
[0343] ii. The 'do:' and 'do-first:' actions of each state indicate the actions taken by the BE sender while remaining in such a state, in reaction to certain events (these are the self-arguments on the left side of the → operator).
[0344] iii. A certain event can trigger multiple 'do:' actions. There is no order requirement for these actions, except that the 'do-first;' action is completed before any other 'do:' action.
[0345] iv. update(PktStateMachines[], <event type>) evaluates the indicated packet-level state machines as triggered by 'event_type' to perform, which can result in instantiation of a state machine for a new untracked PID, state invariance of an existing state machine instance, state transition of an existing state machine instance, or destruction of a pre-existing state machine.
[0346] The v. update(tail_pid) function sets the tail_pid to be the highest PID tracked that has been transmitted (i.e., not in Pkt-Pending state) from the BE sender's perspective.
[0347] c. The BE sender needs several variables in addition to its main connection super state (ConnSuperState). All variables are listed in FIG. 8A Table 1, indicating the variable type, data type, and description, in the following format: i. <var_type> <var_name>: <data_type> [ <description>] ii.<var_type> can be either ‘+’ for state variables updated by the state machine, or ‘*’ for auxiliary variables (used to convey certain events (triggered under specific conditions)) or auxiliary computations tracked by the state machine.
[0348] d. PktStateMachines[N_MAX_PID] is a hash table of type PktSndrState indexed by the PIDs of the packets being tracked.
[0349] ‘i. N_MAX_PID’ denotes the maximum number of packets tracked by the BE sender node.
[0350] e. The BE sender connection state machine is responsible for the RTO timer.
[0351] f. The BE sender connection state machine is responsible for resolving all schedule_tx() requests received from the associated set of packet state machines (i.e. if multiple packet state machines request transmission at the same time, the connection state machine must resolve the order in which they will be scheduled for TX). This resolution is policy-based and is not illustrated in FIG. 8A .
[0352] g. FIG. 8C The policy-based outputs of any connection state machine (i.e. outputs generated based on policies that can differ between different implementations) are not explicitly depicted, but it is assumed that any implementation of this state machine must support all of them (according to its specific policy). For example, the transmission order of data-carrying message packets across the sender’s PktStateMachines is all policy-based (i.e. the order in which requested retransmissions and requested first transmissions are scheduled for transmission), and is not illustrated in this figure. Thus, the “TX(pid)” reference in any action of this connection state machine does not imply an output, but rather a notification that a data-carrying message packet has been scheduled for transmission based on the local policy.
[0353] h. When a SNAK signaling packet is received indicating the existence of an infinite hole, the BE sender connection state machine transitions from the Active superstate to the Recovery superstate. When a NAK signaling packet is received signaling an infinite hole PID, or when an ACK signaling packet is received acknowledging everything before the infinite hole, the BE sender connection state machine transitions from the Recovery superstate to the Active superstate.
[0354] 5. FIG. 8B is a flowchart of the operations of the BE sender connection state machine and the BE receiver connection state machine whenever a packet is received or transmitted, as indicated by step 800. In step 802, auxiliary variables are computed (as described in FIG. 9A and 9B The value indicated by <...> is determined in step 804. If so, the exit action of the current superstate is evaluated and executed if indicated in step 806. The state machine transitions to the new superstate in step 808. The entry action of the new current superstate is evaluated and executed if indicated in step 810. All do-first actions of the current superstate are evaluated and executed if indicated in step 812 (operation flows here after step 810, or when step 804 is NO). All do actions of the current superstate are evaluated and executed if indicated in step 814. The auxiliary variables are recomputed in step 816. Based on the recomputation of the auxiliary variables, it is determined in step 818 whether a connection superstate transition is required. If so, operation returns to step 806.
[0355] 6. The connection-level state machine of the BE receiver node (handling the superstates of the receiver) is as shown in FIG. 9B , FIG. 9C and FIG. 9B This state machine, together with the packet-level hierarchical state machines, provides the complete view of the BE receiver.
[0356] a. In this state machine: i. The 'entry:' action of each state indicates the action performed by the BE receiver each time it transitions to such a state.
[0357] ii. The 'do:' and 'do-first:' actions of each state indicate the actions taken by the BE receiver while staying in such a state, in reaction to certain events (these are the arguments on the left side of the → operator).
[0358] iii. A certain event can trigger multiple 'do:' actions. There is no order requirement for these actions, except that the 'do-first;' action is done before any other 'do:' action.
[0359] iv. The update(PktStateMachines[], <event_type>) function evaluates the indicated packet-level state machine to be executed triggered by <event_type>, which can result in instantiation of a state machine for a PID not being tracked, state invariance of an existing state machine instance, state transition of an existing state machine instance, or destruction of a pre-existing state machine.
[0360] v. The 'exit:' action of each state indicates the action performed by the BE receiver each time it transitions away from such a state.
[0361] vi. The update(next_epid) function sets next_epid to the second lowest un- received PID (i.e., the packet is not in the Pkt-Rcvd or Pkt-Dlvr state) from the receiver's perspective.
[0362] vii. The update(max_ooo_rcv_pid) function sets max_ooo_rcv_pid to the highest out-of-order received PID (i.e., the packet is in the Pkt-Rcvd state) from the receiver's perspective.
[0363] viii. The update(inf_pid) function sets inf_pid to the lowest untracked PID (i.e., the packet state machine does not exist) from the receiver's perspective. The update(inf_pid) function is described in more detail below.
[0364] b. The BE receiver uses several variables in addition to its main connection super state (ConnSuperState). All variables are listed in Table 1 below, indicating the variable type, data type, and description, in the following format: FIG. 9A i. <var_type> <var_name>: <data_type> [ <description> ] <description>] ii.<var_type> can be a'+'indicating a state variable updated by the state machine, or a'*'indicating an auxiliary variable (used to convey certain events (triggered under specific conditions)) or auxiliary calculations tracked by the state machine.
[0365] c. PktStateMachines[N_MAX_PID] is a hash table of type PktRcvrState indexed by the PIDs of the packets being tracked.
[0366] i. 'N_MAX_PID' indicates the maximum number of packets tracked by the BE receiver node.
[0367] d. The BE receiver connection state machine is responsible for resolving all ACK(), RACK() and SNAK() requests from the associated set of packet state machines (i.e. if multiple packet state machines request the same signaling at the same time, the connection state machine must resolve which PIDs are included in the RACK / SNA packet, or which PIDs are merged into the same ACK packet). This resolution is policy-based and is not illustrated in FIG. 9A .
[0368] e. FIG. 8C The policy-based output of any BE receiver connection state machine (i.e. output generated based on policies that can differ between implementations) is not explicitly depicted, but it is assumed that any implementation of this state machine must support all of these outputs (according to its specific policy). For example, the generation of a multi-PID RACK(), a multi-PID SNAK() or an accumulated ACK() packet (all of which provide signaling associated with multiple PIDs) in response to a request issued by the receiver's PktStateMachines are all policy-based (i.e. which PIDs are included on each RACK or SNAK, or how many PIDs are merged into each ACK), and are therefore not illustrated in this figure.
[0369] f. When the value of res_depleted is true, the BE receiver connection state machine transitions from the ACTIVE superstate to the RECOVERY superstate, which precedes the execution of any do operations in the ACTIVE state machine. When the value of res_count is equal to the value of res_max, the BE receiver connection state machine transitions from the RECOVERY superstate to the ACTIVE superstate, indicating that the expected number of tracking resources has been recovered. Both superstate transitions do not occur in sequence (i.e. no "do:” operations are executed), as illustrated in FIG. 9C .
[0370] As mentioned above, when the receiver sends an unlimited hole indication, the super state generally changes from ACTIVE to a super state. Generally, the super state changes from RECOVERY to ACTIVE when all holes are filled, which means that the resources in the receiver have recovered, so that the ACTIVE super state starts with the receiver in the maximum available resources.
[0371] When an unlimited hole is determined, the value of inf pid must be determined, which unlimited hole starts at inf pid (inf pid is the known pid where the unlimited hole starts). FIG. 9C A flowchart is provided.
[0372] For FIG. 6A , the connection state machine from the receiver has exhausted its hole tracking resources starts at step 900, which is due to a new hole being discovered or a previously tracked hole being split into multiple holes. In step 902, it is determined whether the pid fills an existing range hole, which then would cause the range hole to be split into two non-contiguous holes. If no, the pid creates a new hole and is discarded in step 904. In step 906, it is determined whether res count is 0. If yes, the value of inf pid is set to the value of max ooo rcv pid plus 1 in step 908. If res count is not 0 in step 906, the pid is used for a new range hole and must be tracked as a new point hole, so it is determined whether this is the first hole in step 910. If yes, the value of inf pid is set to the value of epid plus 1 in step 914. If not the first hole in step 912, the value of inf pid is set to the value of max ooo rcv pid plus 2 in step 916.
[0373] If the pid does fill an existing range hole in step 902, then in step 918 it is determined whether res_count is 0. If so, then in step 920 it is determined whether the value range_split_in_point_and_range_holes is true. If not, then both resources must be released, and in step 922 it is determined whether the highest tracked hole is a point hole. If not, this means that the highest tracked hole is a range hole, so the range hole must be discarded to release both resources, and in step 924 the value of inf_pid is set to the highest range hole start PID, and max_ooo_rcv_pid is set to inf_pid minus 1. If the highest tracked hole is a point hole, then tracking of the highest point hole is stopped and one resource is released, and in step 926 the value of inf_pid is set to the highest point hole PID, and the value of max_ooo_rcv_pid is set to inf_pid minus 1.
[0374] After step 926, if res_count is not 0 in step 918, or if the value range_split_in_point_and_range_holes is true, both of which indicate that one resource needs to be released, then in step 928 it is determined whether the highest tracked hole is a point hole. If so, then tracking of the point hole is stopped, and in step 930 the value of inf_pid is set to the highest point hole PID, and the value of max_ooo_rcv_pid is set to inf_pid minus 1. If the highest tracked hole is not a point hole in step 928, then the highest tracked hole is a range hole, which must be tracked as a point hole and one resource is released, and in step 932 the value of inf_pid is set to the highest range hole start PID plus 1, and the value of max_ooo_rcv_pid is set to inf_pid minus 1.
[0375] This algorithm for generating the value of inf_pid is the most efficient algorithm, as it never "wastes" available tracking resources, however it is not the only possible update algorithm that is functionally correct. Other algorithms can be used in some examples. For example, from a complexity perspective, when two resources are needed (e.g., because a new range hole was discovered), but only one resource is available, a simpler algorithm would waste the single available tracking resource.
[0376] While a series of six different state machines have been illustrated, FIG. 6B the sender packet state machine in FIG. 7A the receiver packet state machine in FIG. 7B the sender packet state machine in FIG. 8A the receiver packet state machine in FIG. 8B the sender packet state machine in FIG. 9A sender connection state machine in FIG. 9B and FIGS. 10A-10E receiver connection state machine in ), but it is understood that this is a logical representation and that actual implementations can vary by combining sender or receiver state machines as needed for the particular implementation of the state machine.
[0377] FIGS. 11A-11C An exemplary packet format for a transaction packet is illustrated. FIGS. 11D-11F An exemplary packet format for a reliability protocol signaling packet is illustrated. FIG. 14 Changes to packet level interleaving are illustrated.
[0378] BE transport combines pre-existing features in an innovative way to achieve symmetric effective throughput in request and response traffic by using out-of-order reception, avoiding reliance on GoBackN, and avoiding use of data reordering buffers, while minimizing RTO possibilities. The list of pre-existing features combined to achieve this unique property is as follows: Symmetric reliability protocol for request and response traffic.
[0379] Out-of-order data placement for both request and response traffic.
[0380] Scatter support for inbound response data placement.
[0381] Avoidance of RTO with the help of FLUSH packets triggered by FTO.
[0382] BE transport includes a modular architecture (explicitly defining requestor, responder, and completer roles with independent capabilities, each of which can be enabled or disabled on a per-connection basis) that enables independent customizing of granular resource allocation to the needs of each connection.
[0383] BE transport includes a new reliability protocol with ACK / NAK / RACK / SNAK, mandatory packet level state machines (Pkt-Pending, Pkt-Outstanding, Pkt-Dlvr, Pkt-Lost for the sender, and Pkt-Rcvd, Pkt-Dlvr, Pkt-Lost for the receiver).
[0384] RACK packets prevent unnecessary retransmission of OOO received packets, and enable CCs to avoid overreacting (by over-limiting transmission rate) due to delayed ACKs.
[0385] SNAK packets explicitly signal the hole(s) in the PID stream (i.e., avoid inferential holes) to reduce the likelihood of getting stuck in an infinite hole. Furthermore, SNAK allows multiple explicit signaling of the same hole, as opposed to a single one in NAK, further reducing the likelihood of getting stuck in an infinite hole. In the ACTIVE superstate, signaling the latest hole is mandatory. In the RECOVERY superstate, signaling an infinite hole is mandatory. In both superstates, signaling older holes included in the SNAK is policy-based.
[0386] BE transport includes new effective throughput optimization superstates (Active and Recovery) to prevent wasting bandwidth (receiver-end OOO tracking resources exhaustion) when an infinite hole is detected. While this condition can be infrequent, the Recovery superstate allows the connection to reach a steady point before transitioning back to the Active state (which maximizes effective throughput).
[0387] BE transport includes a request / response traffic packet-level interleaving option to improve QoS between request traffic and response traffic on a BE connection.
[0388] FIG. 15 Various examples of packet traffic and packet traffic problems are illustrated. The example of transmitting five packets is used to simplify the description, it being understood that transmission can vary from a single packet to thousands of packets. In row (a), transmitter 1202 transmits packets 1 through 5 to be received by receiver 1204. The sequential transmission of the five packets by transmitter 1202 is used in the four example packet delivery conditions of rows (b) through (e).
[0389] In row (b), receiver 1204 successfully receives the five packets in sequence. In response, receiver 1204 transmits ACK-1 through ACK-5 response signaling packets to acknowledge successful receipt of packets 1 through 5.
[0390] In row (c), packet 3 is dropped during transmission, so that receiver 1204 receives only packets 1, 2, 4, and 5. In response, receiver 1204 provides ACK-1, ACK-2, SNAK-3, RACK-4, and RACK-5 response signaling packets. Upon receipt of the SNAK-3 response signaling packet, transmitter 1202 retransmits packet 3. This time, receiver 1204 successfully receives packet 3 and responds with an ACK-5 response signaling packet.
[0391] Row (d) illustrates out-of-order packet reception, as receiver 1204 receives packet 1, followed by packet 2, followed by packet 4, followed by packet 5, and finally packet 3. Receiver 1204 responds by providing ACK-1, ACK-2, SNAK-3, RACK-4, and RACK-5 response signaling packets. Then, when receiver 1204 receives out-of-order packet 3, receiver 1204 provides an ACK-5 response signaling packet. Transmitter 1202 will complete providing retransmission packet 3 based on the SNAK-3 response signaling packet, and receiver 1204 responds with ACK-5.
[0392] Row (e) illustrates that the receiver is no longer able to track packet status, creating an infinite hole. Receiver 1204 receives packets 1, 2, and 5, with packets 3 and 4 missing. The loss of these two packets is illustrated as sufficient to exceed the OOO tracking resources (i.e., in this example, the receiver can only track a single point hole at a time), such that a loss of packet status occurs, and an infinite hole is generated upon receiving packet 5. Receiver 1204 provides ACK-1, ACK-2, and SNAK-3 with an infinite hole response signaling packet. When SNAK-3 with an infinite hole response signaling packet is received, transmitter 1202 enters RECOVERY mode and retransmits packet 3. Receiver 1204 receives packet 3, and because the internal tracking resources are flushed upon receiving packet 3, receiver 1204 responds with ACK-3 and NAK-4 response signaling packets, both of which indicate a return to ACTIVE mode. Transmitter 1202 responds by sending packets 4 and 5. Receiver 1204 returns ACK-5, completing the transmission.
[0393] FIGS. 15A-15E illustrates the relationship of FIG. 16 Similarly, FIGS. 16A-16G illustrates the relationship of FIG. 17 illustrates the relationship of FIGS. 17A-17G illustrates the relationship of FIG. 18 illustrates the relationship of FIGS. 18A-18 illustrates the relationship of FIGS. 15A-15E I.
[0394] Referring now to FIG. 14 , this is the state machine operation for the case shown in FIG. 15A Row (b), where all packets are delivered in order. FIG. 6A The diagram illustrates sender 1500, which has a sender connection state machine 1502 for a single BE RDMA connection and a series of sender packet-level state machines 1504. The diagram also illustrates receiver 1501, which has a receiver connection state machine 1506 and a receiver packet-level state machine 1508 for a BE RDMA connection. Sender connection state machines 1502 and 1506 belong to different (peer) nodes and are both in the active connection superstate indicating normal operation. Sender connection state machine 1502 has an unackd_pid value of 1 and an enable_tx value of 1. Its tail_pid value is 0. Receiver connection state machine 1506 has res_count and res_max values of 4, an epid value of 1, and a max_ooo_rcv_pid value of 0. In operation 1, sender connection state machine 1502 provides a creation indication to sender packet-level state machines 1504 and creates sender packet state machine 1. The sender connection state machine 1502 is notified by the host interface 256 or the egress buffer memory 268 that a packet has been received, and obtains the relevant packet context and stores it in the relevant state machine state and context area of the appropriate BE section in RAM 263. The sender packet state machine 1 is in the Pkt-Pending state, as shown below. FIG. 15B As shown. In operation 2, the sender packet state machine 1 provides a schedule_tx message for PID=1 to the sender connection state machine 1502, as this is the entry action for the Pkt-Pending state. The schedule_tx message is received at the sender connection state machine 1502 and forwarded to the egress memory buffer 268 and the egress packet building hardware 259 to cause the packet to be dequeued and placed in the transmission path. In operation 3, the sender connection state machine 1502 provides a create or RX message to the sender packet-level state machine 1504, which creates the sender packet state machine 2, whose state is packet pending. With each message provided to the state machine, connection, or packet, the required context and state are obtained from the relevant state machine state and context area in the appropriate BE portion of RAM 263 and provided with the message to the hardware state machine pool 283 to allow specific state machine operations. In operation 4, the sender packet state machine 2 provides a schedule_tx message for PID=2 to the sender connection state machine 1502. Similar operations 4-10 occur to create sender packet state machines 3, 4, and 5, each in the state of packet pending, and send schedule_tx messages for packets 3, 4, and 5.
[0395] Move to FIG. 8A In operation 11, the sender connection state machine 1502 provides a transmission or TX message to the sender packet state machine 1 to indicate that packet 1 is being transmitted. The sender connection state machine 1502 performs this action based on a transmission indication received from the egress memory buffer 268 and the egress packet construction hardware 259 (indicating that the packet is in the transmission queue of the Ethernet module 267). FIG. 9A In the ACTIVE superstate of the sender connection state machine, the action is the statement do:TX(pid)→{update(PktStateMachines[p=pid], TX)update(tail_pid)}. The sender packet state machine 1 advances to the packet incomplete state. The sender connection state machine 1502 increments the tail PID value to 1. In operation 12, the sender connection state machine 1502 transmits packet 1 (PID=1) to the receiver connection state machine 1506. It should be understood that the sender connection state machine 1502 does not actually send the packet, but rather instructs the packet processing hardware (such as egress buffer memory 268 and egress packet construction hardware 259) to send the packet based on receiving the schedule_tx message or according to its own intention. Similarly, the receiver connection state machine 1506 does not actually receive the packet, but rather receives an indication that the packet has arrived from the ingress packet processing hardware (such as ingress buffer memory 268 and ingress packet processing and routing hardware 257). Using the indication of packet arrival, the context and state of the relevant hardware state machine are retrieved from the relevant BE portion in RAM263.
[0396] Based on the packet reception instruction, the receiver connection state machine 1506 is in FIG. 7A In the ACTIVE state machine, the statement `do-first:RX(pid ≥ epid)→update(PktStateMachines[p ≥ epid], RX)` is executed. In operation 13, the receiver connection state machine 1506 provides an RX message with PID 1 and expected PID 1 to the receiver packet-level state machine 1508. The receiver packet-level state machine 1508 creates receiver packet state machine 1 based on the RX (pid=p==epid) condition, whose state is "packet delivered," as shown below. FIG. 15C The sender connection state machine 1502 provides a transmit message to the sender packet state machine 2, which transitions to the packet not done state. The sender connection state machine 1502 increments the tail PID value to 2. In operation 15, the receiver packet state machine 1 provides a request ACK message for PID = 1 to the receiver connection state machine 1506, as this is the entry action of the Pkt-Dlver state, and in operation 16, the receiver packet state machine 1 terminates operation as it exits the Pkt-Dlver state (known as self-destruct). Upon receiving the request ACK message, the receiver connection state machine 1506 informs the egress packet build hardware 259 to prepare an ACK signaling packet for the indicated PID. The receiver connection state machine 1506 advances to the expected PID or epid value 2 (the previous next epid value), and the next epid value is incremented to 3. In operation 17, packet 2 (PID = 2) is provided by the sender connection state machine 1502 and received by the receiver connection state machine 1506.
[0397] In FIG. 15D In response, in operation 18, the receiver connection state machine 1506 provides the RX message with PID = 2 and expected PID = 2 to the receiver packet level state machine 1508, which creates receiver packet state machine 2 with state packet delivered. The epid value is updated to next epid value 3, and the next epid value is incremented to 4. In operation 19, the sender connection state machine 1502 provides the transmit message to sender packet state machine 3, which advances to the packet not done state. The sender's connection state machine 1502 increments the tail PID value to 3. In operation 20, the receiver packet state machine 2 issues a request ACK message for PID = 2 to the receiver connection state machine 1506, and terminates itself in operation 21. In operation 22, the sender connection state machine 1502 provides packet 3 (PID = 3) to the receiver connection state machine 1506. In operation 23, the sender connection state machine 1502 provides the transmit message to sender packet state machine 4, which advances to the packet not done state. The sender connection state machine 1502 increments the tail PID value to 4. In operation 24, the receiver connection state machine 1506 provides the RX message with PID = 3 and expected PID = 3 to the receiver packet level state machine 1508, which causes the creation of receiver packet state machine 3 with state packet delivered. The epid value is updated to next epid value 4, and the next epid value is incremented to 5. In operation 25, the receiver connection state machine 1506 provides the ACK response signaling packet with PID = 1 to the sender connection state machine 1502. As described above, the sender connection state machine 1502 does not actually receive the packet, but is provided with an indication by the ingress packet processing hardware (such as ingress buffer memory 268 and ingress packet processing and routing hardware 257) that the packet has arrived. With the indication of packet arrival, the context and state of the relevant hardware state machine is retrieved from the relevant BE portion in RAM 263.
[0398] In operation 26, the sender connection state machine 1502 provides the ACK message with PID = 1 to sender packet state machine 1, which advances to the packet delivered state FIG. 8A ). The ACK message is based on the do: ACK(pid ≠ tail_pid) → { reset(RTO Timer); update(PktStateMachines[p≤pid], ACK); set(unackd_pid, pid+1)} statement in FIG. 6A . The sender packet state machine 1 advances to state Pkt-Dlvr is based on the ACK(pid≥p) condition, as in FIG. 15D The sender connection state machine 1502 increments the unacknowledged PID value to 2, as packet 1 has been acknowledged. In operation 27, the receiver packet state machine 3 provides the receiver connection state machine 1506 with a request ACK message for PID = 3.
[0399] With reference to FIG. 15E In operation 28, the receiver packet state machine 3 terminates itself. In operation 29, the sender packet state machine 1 terminates operation. In operation 30, the sender connection state machine 1502 transmits packet 4 (PID = 4) to the receiver connection state machine 1506. In operation 31, the receiver connection state machine 1506 provides the receiver packet level state machine 1508 with an RX message for PID 4 and expected PID 4, which causes the creation of receiver packet state machine 4, which is in the packet delivered state. The receiver connection state machine 1506 changes the epid value to the next_pid value 5, and increments the next_pid value to 6. In operation 32, the sender connection state machine 1502 provides the sender packet state machine 5 with a transmit message, which enters the packet not completed state. The sender connection state machine 1502 increments the tail PID value to 5. In operation 33, the receiver connection state machine 1506 provides the sender connection state machine 1502 with an ACK response signaling packet for PID 2. In operation 34, the sender connection state machine 1502 forwards the ACK message for PID 2 to the sender packet state machine 2, and increments the unacknowledged PID value to 3. The sender packet state machine 2 advances to the packet delivered state. In operation 35, the receiver packet state machine 4 provides the receiver connection state machine 1506 with a request ACK message for PID 4. In operation 36, the receiver packet state machine 4 terminates operation. In operation 37, the sender packet state machine 2 terminates. In operation 38, the sender connection state machine 1502 provides the receiver connection state machine 1506 with packet 5 (PID = 5).
[0400] Turning to FIGS. 16A-16G In operation 39, the receiver connection state machine 1506 provides an RX message with PID 5 and expected PID 5 to the receiver packet-level state machine 1508. This causes the creation of receiver packet state machine 5, whose state is packet delivered. Receiver connection state machine 1506 changes the epid value to the next_pid value 6 and increments the next_pid value to 7. In operation 40, receiver packet state machine 5 provides a request ACK message for PID 5 to receiver connection state machine 1506. In operation 41, receiver packet state machine 5 terminates the operation. In operation 42, receiver connection state machine 1506 provides an ACK response signaling packet with PID 3 to sender connection state machine 1502. In operation 43, sender connection state machine 1502 provides an ACK message with PID 3 to sender packet state machine 3. Sender packet state machine 3 advances to the packet delivered state and terminates the operation in operation 44. Sender connection state machine 1502 increments the unackd_pid value to 4. In operation 45, the receiving connection state machine 1506 provides an ACK response signaling packet for PID 4 to the sending connection state machine 1502. In operation 46, the sending connection state machine 1502 provides an ACK message for PID 4 to the sending packet state machine 4, and the sending packet state machine 4 advances to the packet delivered state. The sending connection state machine 1502 advances the unacknowledged PID value to 5. In operation 47, the sending packet state machine 4 terminates the operation. In operation 48, the receiving connection state machine 1506 provides an ACK response signaling packet for PID 5 to the sending connection state machine 1502. Upon receiving the ACK response signaling packet for PID 5, the sending connection state machine 1502 provides an ACK message for PID 5 to the sending packet state machine 5 in operation 49 and increments the unackd_pid value to 6. The sending packet state machine 5 advances to the packet delivered state and terminates the operation in operation 50.
[0401] Now for reference FIG. 14 This is aimed at FIG. 15A The state machine operation shown in line (c) involves discarding group 3. The operation is related to... FIG. 15B and FIG. 16C Same, up to operation 21. From FIG. 16D Operations 22 begin in the middle, as packet 3 is discarded, and the operations begin to diverge. Operation 22 indicates that packet 3 is discarded. In operation 23, the sender connection state machine 1502 provides a transmit message to the sender packet state machine 4, which then enters the packet not completed state. The sender connection state machine 1502 increments the tail PID value to 4. In operation 24, the receiver connection state machine 1506 provides an ACK response signaling packet to the sender connection state machine 1502 for packet 1. In operation 25, the sender connection state machine 1502 provides an ACK message to the sender packet state machine 1 with PID 1, and the sender packet state machine 1 changes state to packet delivered. The sender connection state machine 1502 increments the unackd_pid value to 2.
[0402] In FIG. 8A In the middle, in operation 26, the sender packet state machine 1 terminates. In operation 27, the sender connection state machine 1502 provides packet 4 (PID=4) to the receiver connection state machine 1506. The receiver connection state machine 1506 performs FIG. 16E The do-first: RX(pid > epid) -> update(PktStateMachines[p > epid], RX) statement causes the transmission of RX messages for receiver packet state machines 3 and 4. In operation 28, a RX message is provided for creating receiver packet state machine 3, and the receiver packet state machine 3 starts in the packet lost state because packet 4 has been received but packet 3 has not been received (as indicated by the RX message's pid = 4, epid = 3 parameters), which causes the state machine to follow the RX(pid > p) condition. In operation 29, a request is provided for creating receiver packet state machine 4. Receiver packet state machine 4 starts in the packet received state. Here, the state machine uses the RX message's pid = 4, epid = 3 parameters to jump to the Pkt-Rcvd state by satisfying the RX(pid = p > epid) condition. Receiver connection state machine 1506 decrements the res_count value to 3 because one out-of-order resource has been used. The next_epid value is incremented to 5, and the max_ooo_rcv_pid value is set to 4. In operation 30, sender connection state machine 1502 provides a transmit message to sender packet state machine 5, which advances to the packet not completed state. Sender connection state machine 1502 increments the tail PID value to 5. In operation 31, receiver connection state machine 1506 provides an ACK response signaling packet for packet 2 to sender connection state machine 1502. In operation 32, an ACK message with a PID value of 2 is provided to sender packet state machine 2. Sender packet state machine 2 advances to the packet delivered state and terminates in operation 33. After operation 32, sender connection state machine 1502 increments the unacknowledged PID value to 3. In operation 34, receiver packet state machine 3 provides a request SNAK message for packet 3 to receiver connection state machine 1506 because that is the entry action for the Pkt-Lost state. In operation 35, receiver packet state machine 4 provides a request RACK message indicating packet 4 to receiver connection state machine 1506 because that is the entry action for the Pkt-Rcvd state.
[0403] In FIG. 16F In the middle, in operation 36, the sender connection state machine 1502 provides packet 5 (PID = 5) to the receiver connection state machine 1506. In operation 37, the receiver connection state machine 1506 provides a packet 5 received message with epid = 3 to the receiver packet state machine 3, and in operation 38 provides a received packet 5 message with epid = 3 to the receiver packet state machine 4, and in both cases the message has no effect on the packet state machines. In operation 39, the receiver packet state machine 5 is created from the message with PID value 5 and expected PID value 3, and the receiver packet state machine 5 opens in the packet received state. The receiver connection state machine 1506 increments the next epid value to 6, and increments the max ooo rcv pid value to 5. In operation 40, the SNAK response signaling packet for packet 3 is provided from the receiver connection state machine 1506 to the sender connection state machine 1502. In operation 41, the RACK response signaling packet for packet 4 is provided from the receiver connection state machine 1506 to the sender connection state machine 1502. In operation 42, the sender connection state machine 1502 provides a SNAK message with PID 3 to the sender packet state machine 3 based on the do: SNAK(pid) -> { update(PktStateMachines[p=pid], SNAK); update(tail_pid)} statement. The sender packet state machine 3 enters the packet lost state based on the SNAK(pid=p) condition. In operation 43, the sender packet state machine 3 provides a schedule tx message for packet 3 to the sender connection state machine 1502 to cause packet 3 to be retransmitted, which is the entry action for the Pkt-Lost state. In operation 44, the sender connection state machine 1502 provides a RACK message with PID 4 to the sender packet state machine 4 based on the do: RACK(pid) -> { update(PktStateMachines[p=pid], RACK); update(tail_pid)} statement. The sender packet state machine 4 enters the packet received state based on the RACK(pid=p) condition. In operation 45, the receiver packet state machine 5 provides a RACK request for packet 5 to the receiver connection state machine 1506. In operation 46, the sender connection state machine 1502 provides a transmit message for packet 3 to the sender packet state machine 3, which enters the packet retransmission state. In operation 47, the RACK response signaling packet for packet 5 is provided from the receiver connection state machine 1506 to the sender connection state machine 1502.
[0404] Reference FIG. 16G In operation 48, the sender connection state machine 1502 provides the RACK message with PID 5 to the sender packet state machine 5, which advances to the packet received state. The tail PID value of the sender connection state machine 1502 is updated back to 3. In operation 49, the sender connection state machine 1502 provides the retransmitted packet 3 (PID = 3) to the receiver connection state machine 1506. Based on the do-first: RX(pid >= epid) -> update(PktStateMachines[p >= epid], RX) statement, the receiver connection state machine 1506 updates the receiver packet state machines 3, 4, and 5. In operation 50, the receiver connection state machine 1506 provides the receive message for packet 3 with PID value 3 and expected PID value 3 to the receiver packet state machine 3, which then advances to the packet delivered state based on the RX(pid = p == epid) condition. In operation 51, the receiver connection state machine 1506 provides the receive message for packet 3 with PID value 3, expected PID value 3, and next_epid of 6 to the receiver packet state machine 4, which enters the packet delivered state based on the RX(pid == epid < p) condition. In operation 52, the receiver connection state machine 1506 provides the received packet 3 indication with PID value 3, expected PID value 3, and next_epid of 6 to the receiver packet state machine 5. The receiver packet state machine 5 advances to the packet delivered state based on the RX(pid == epid < p) condition. The resource counter res_count is incremented to 4 due to the missing packet being received. The epid value is set to the next_epid value of 6, and the next_epid value is incremented to 7. In operation 53, the receiver packet state machine 3 provides the ACK message request for PID 3 to the receiver connection state machine 1506, which is the enter action of the Pkt-Dlvr state. In operation 54, the receiver packet state machine 3 terminates. In operation 55, the receiver packet state machine 4 provides the request ACK message for PID 4 to the receiver connection state machine 1506, which is the enter action of the Pkt-Dlvr state. In operation 56, the receiver packet state machine 4 terminates.
[0405] Reference FIG. 14 In operation 57, the receiver packet state machine 5 provides a request ACK message for PID 5 to the receiver connection state machine 1506, which is the entry action for the Pkt-Dlvr state. In operation 58, the receiver packet state machine 5 completes operation. In operation 59, the receiver connection state machine 1506 provides an ACK response signaling packet for packet 5 to the sender connection state machine 1502, cumulatively acknowledging packets 3, 4, and 5. In operation 60, based on the do: ACK(pid == tail_pid) -> {clear(RTO Timer); update(PktStateMachines[p <= pid], ACK); set(unackd_pid, pid+1)} statement, the sender connection state machine 1502 provides an ACK message for packet 5 to the sender connection state machine 3, which enters the packet delivered state based on the ACK(pid >= p) condition. In operation 61, the sender packet state machine 3 terminates. Based on the do: ACK(pid!= tail_pid) -> {reset(RTO Timer); update(PktStateMachines[p <= pid], ACK); set(unackd_pid, pid+1)} statement, in operation 62, the sender connection state machine 1502 provides an ACK message for packet 5 to the sender packet state machine 4, which enters the packet delivered state and terminates in operation 63. In operation 64, the ACK message for packet 5 is also provided to the sender packet state machine 5, which enters the packet delivered state. The sender connection state machine 1502 increments the unack_pid value to 6. In operation 65, the sender packet state machine 5 terminates. To this point, FIGS. 17A to 17G The scenario of line (c) is complete.
[0406] FIG. 14 Figures 1 1A and 1 IB illustrate FIG. 15A the state machine operations of line (d), where packet 3 is delayed, rather than lost. Again, operations 1 through 21 are the same as in FIG. 15B and FIG. 17C In FIG. 17C , in operation 22, packet 3 (PID=3) is provided by the sender connection state machine 1502, but as from FIG. 17D The exiting line shows that delivery is delayed. In operation 23, the transmission is provided from the sender connection state machine 1502 to the sender packet state machine 4. The sender packet state machine 4 enters the packet not completed state. The sender connection state machine 1502 increments the tail PID value to 4. In operation 24, the ACK response signaling packet for packet 1 is provided from the receiver connection state machine 1506. In operation 25, the sender connection state machine 1502 provides the ACK message for PID = 1 to the sender packet state machine 1. The sender connection state machine 1502 increments the unacknowledged PID value to 2.
[0407] Referring to FIG. 17E Upon receiving the ACK signaling response for packet 1, the sender packet state machine 1 advances to the packet delivered state and terminates in operation 26. In operation 27, the sender connection state machine 1502 provides packet 4 (PID = 4) to the receiver connection state machine 1506. In operation 28, the RX message with PID value 4 and expected PID value 3 is provided by the receiver connection state machine 1506 to create the receiver packet state machine 3, which starts in the packet lost state. In operation 29, the RX message with PID value 4 and expected PID value 3 is provided to the receiver packet level state machine 1508, which causes the creation of the receiver packet state machine 4 and causes it to enter the packet received state. The res_count value is decremented to 3, the next_epid value is 5, and the max_ooo_rcv_pid value is 4. In operation 30, the transmission message is provided by the sender connection state machine 1502 to the sender packet state machine 5, which advances to the packet not completed state. The tail_pid value is incremented to 5. In operation 31, the ACK response signaling packet for packet 2 is provided by the receiver connection state machine 1506 to the sender connection state machine 1502. In operation 32, the ACK message for PID = 2 is provided by the sender connection state machine 1502 to the sender packet state machine 2. The sender connection state machine 1502 increments the unacknowledged PID value to 3. Upon receiving the ACK message, the sender packet state machine 2 advances to the packet delivered state and stops operating in operation 33. In operation 34, the request SNAK response signaling indication with PID value 3 is provided by the receiver packet state machine 3 to the receiver connection state machine 1506. In operation 35, the RACK message request for packet 4 is provided by the receiver packet state machine 4 to the receiver connection state machine 1506.
[0408] Referring to FIG. 17F In operation 36, the sender connection state machine 1502 provides packet 5 (PID = 5) to the receiver connection state machine 1506. In operation 37, the receiver connection state machine 1506 provides a received packet message with received PID value 5 and epid value 3 to receiver packet state machine 3 and in operation 38 to receiver packet state machine 4. In operation 39, the receiver connection state machine 1506 provides an RX message with PID value 5 and expected PID value 3 to cause the creation of receiver packet state machine 5 in the packet received state. The receiver connection state machine 1506 increments next_epid to 6 and increments max_ooo_rcv_pid to 5. In operation 40, the receiver connection state machine 1506 provides an SNAK response signaling packet for packet 3 to the sender connection state machine 1502. In operation 41, the receiver connection state machine 1506 provides a RACK response signaling packet for packet 4 to the sender connection state machine 1502. In operation 42, the sender connection state machine 1502 provides an SNAK message with PID 3 to the sender packet state machine 3, which enters the packet lost state. In operation 43, the sender packet state machine 3 provides a schedule_tx for packet 3 to the sender connection state machine 1502 to cause packet 3 retransmission, which is the entry action of the Pkt-Lost state. In operation 44, the sender connection state machine 1502 provides a RACK message with PID 4 to the sender packet state machine 4, which enters the packet received state. In operation 45, the receiver packet state machine 5 provides a RACK signaling response request for packet 5 to the receiver connection state machine 1506. After operation 45, packet 3 of operation 22 is finally received and packet 3 is provided to the receiver connection state machine 1506. In operation 46, the sender connection state machine 1502 provides a packet 3 transmission indication to the sender packet state machine 3, which enters the packet retransmission state. In operation 47, the receiver connection state machine 1506 provides a RACK response signaling packet for packet 5 to the sender connection state machine 1502.
[0409] Reference FIG. 17G In operation 48, the sender connection state machine 1502 provides the RACK message with PID 5 to the sender packet state machine 5, which enters the packet received state. The sender connection state machine 1502 updates the tail PID value to 3. In operation 49, the receiver connection state machine 1506 provides the received indication for packet 3 to the receiver packet state machine 3, which expects PID 3. The receiver packet state machine 3 advances to the packet delivered state. Similarly, in operation 50, the received indication for packet 3 (expecting PID 6) is provided to the receiver packet state machine 4, which advances to the packet delivered state. In operation 51, the received indication for packet 3 (expecting PID 6) is provided to the receiver packet state machine 5, which enters the packet delivered state. The receiver connection state machine 1506 increments the res_count value to 4, indicating that all tracking resources are available. The next_epid value is transferred to the epid, and the next_epid is incremented to 7. In operation 52, the receiver packet state machine 3 provides the ACK response signaling request for packet 3 to the receiver connection state machine 1506. In operation 53, the receiver packet state machine 3 stops operation. In operation 54, the receiver packet state machine 4 provides the ACK response signaling request for packet 4 to the receiver connection state machine 1506. In operation 55, the receiver packet state machine 4 terminates. In operation 59, the sender connection state machine 1502 provides the retransmitted packet 3 to the receiver connection state machine 1506. In operation 56, the receiver packet state machine 5 provides the ACK response signaling request for packet 5 to the receiver connection state machine 1506. In operation 57, the receiver packet state machine 5 completes operation. In operation 58, the receiver connection state machine 1506 provides the ACK response signaling packet for packet 5 to the sender connection state machine 1502, cumulatively acknowledging packets 3, 4, and 5.
[0410] Reference FIG. 14 In operation 60, the receiver connection state machine 1506 provides an ACK response for the received retransmitted packet 3 with a pid value of 5. Since the ACK with pid 5 is a duplicate, the sender connection state machine 1502 updates the unack_pid value to the same 6, but does not duplicate the update to the sender packet state machines, as the update has already started based on the ACK of operation 58. In operation 61, the ACK message with PID 5 is provided from the sender connection state machine 1502 to sender packet state machine 3, which advances to the packet delivered state and terminates in operation 62. In operation 63, the ACK message with PID 5 is provided to sender packet state machine 4, which enters the packet delivered state and terminates in operation 64. In operation 65, the sender connection state machine 1502 provides the ACK message with PID 5 to sender packet state machine 5, which enters the packet delivered state and terminates in operation 66. Based on the receipt of the ACK of operation 58, the unack_pid value is incremented to 6. This completes the FIGS. 18A-18H scenario of row (d).
[0411] Reference is now made to FIG. 14 , which illustrates FIGS. 16A to 16D the state machine operations of row (e), where the receiver 1501 exhausts the tracking resources and declares an infinite hole. Operations 1-26 are the same as operations 1-26 in FIG. 18D . One change for purposes of this example is that the res_count is set to 1 and the res_max value is set to 1, rather than 4 in the previous example. This is done for this example to limit the tracking resources in the receiver 1501.
[0412] Reference is made to FIG. 9A In operation 27, the sender connection state machine 1502 provides packet 4 (PID = 4) to the receiver connection state machine 1506. Packet 4 is also lost, resulting in two lost packets, creating a range hole. In operation 28, the sender connection state machine 1502 provides a transmit message to the sender packet state machine 5, which advances to the packet not completed state. The sender connection state machine 1502 increments the tail PID value to 5. In operation 29, the receiver connection state machine 1506 provides an ACK response signaling packet to the sender connection state machine 1502 for packet 2. In operation 30, the ACK message with PID 2 is provided to the sender packet state machine 2. The sender packet state machine 2 advances to the packet delivered state and terminates in operation 31. After operation 30, the sender connection state machine 1502 increments the unacknowledged PID value to 3. In operation 32, the sender connection state machine 1502 provides packet 5 (PID = 5) to the receiver connection state machine 1506. Upon receipt of packet 5, the receiver connection state machine 1506 evaluates the res_depleted variable and determines that the tracking resources are depleted because the res_count value is 1 and a new range hole has been detected. The res_depleted variable being true causes the receiver connection state machine 1506 to transition from the ACTIVE superstate to the RECOVERY superstate (as shown in FIG. 18E FIG. 6B) without evaluating any do statements present in the ACTIVE state machine. This transition is indicated at operation 33, entering the RECOVERY state.
[0413] Referring to FIG. 9C , the receiver connection state machine 1506 is in the RECOVERY superstate. Based on the evaluation as shown in FIG. 9A , the inf_pid value is 4. The receiver connection state machine 1506 discards packet 5 because packet 5 falls into the infinite hole and will be retransmitted once the recovery process is complete. This is a simplification for illustration. The receiver connection state machine 1506 provides RX messages to the receiver packet state machines 3, 4, and 5 based on the do-first: RX(pid > epid) -> update(PktStateMachines[p > epid], RX) statement in the RECOVERY superstate machine of FIG. 7B . However, receiver packet state machines are not created for packets 4 and 5. Receiver packet state machine 4 is not created because FIG. 18F No condition applies for packet 4. The inf_pif value is 4, so the condition for the Pkt-Lost state is not met because the term p < inf_pid is not true. The condition for the Pkt-Rcvd state is not met because the term pid < inf_pid is not true. The condition for the Pkt-Dlver state is not met because p + pid == epid is not true. For the same reason, no receiver packet state machine 5 is created. Because packet 5 has been received but no state machine 5 is created, the packet is discarded.
[0414] In operation 34, the receiver connection state machine 1506 provides a RX message with a PID value of 5 and an epid value of 3 to create a receiver packet state machine 3 in the Pkt_Lost state. The receiver packet state machine 3 is created because the conditions RX(pid > p) and p < inf_pid are true. The receiver connection state machine 1506 decrements the res_count value to 0. In operation 35, the receiver packet state machine 3 provides a SNAK request message with a PID of 3 to the receiver connection state machine 1506. The receiver connection state machine 1506 has determined that a SNAK infinite hole signaling packet is needed based on the do: RX(pid) -> SNAK(inf_pid; inf_hole) statement, but has postponed sending the infinite hole indication, waiting for a SNAK for packet 3. The receiver connection state machine 1506 combines the requested SNAK for PID 3 with the infinite hole SNAK and provides a SNAK signaling packet with a PID value of 3 and an inf_pid value of 4 for the infinite hole indication in operation 36 to the sender connection state machine 1502. The sender connection state machine 1502 detects the infinite hole SNAK signaling packet and enters the RECOVERY superstate with an inf_pid value of 4 in operation 37. The sender connection state machine 1502 sets the enable_tx value to 0 to stop transmitting any previously untransmitted packets.
[0415] Reference is made to FIG. 8A , the sender connection state machine 1502 uses the do: SNAK(inf_pid) -> {update(PktStateMachines[p >= inf_pid], SNAK); set(tail_pid, inf_pid - 1)} statement in the RECOVERY state machine of FIG. 6B to provide the infinite hole message to the sender packet state machines 3, 4, and 5. In operation 38, the sender connection state machine 1502 provides a SNAK message with a PID of 3 to the sender packet state machine 3 to cause the state machine to update the PktStateMachines[p >= inf_pid] based on FIG. 18G The SNAK (pid=p) condition advances the state to Pkt-Lost. In operation 39, the sender connection state machine 1502 provides a SNAK message with inf_pid = 4 to the sender packet state machine 4 to return the state machine to the packet pending state based on the SNAK (inf_pid < p; inf_hole) condition. In operation 40, the sender packet state machine 3 provides a schedule_tx message with PID value = 3 to the sender connection state machine 1502 to cause packet 3 to be retransmitted. In operation 41, the sender connection state machine 1502 provides a SNAK message with inf_pid = 4 to the sender packet state machine 5 to return it to the packet pending state. In operation 42, the sender connection state machine 1502 provides a delivery indication to the sender packet state machine 3, which advances to the packet retransmit state. The sender connection state machine 1502 sets the tail_pid value to 3 and the inf_pid to 4. In operation 43, the sender connection state machine 1502 provides packet 3 (PID = 3) to the receiver connection state machine 1506. In operation 44, the receiver connection state machine 1506 provides an RX message with PID value = 3 and expected PID value = 3 to the receive packet state machine 3, which advances to the packet delivered state. The receiver connection state machine 1506 sets the epid value to 4 and the next_epid value to 5. Based on the received packet 3, the value of res_count is incremented to 1. In operation 45, the receiver packet state machine 3 provides an ACK request message for packet 3 to the receiver connection state machine 1506. In operation 46, the receiver packet state machine 3 terminates operation.
[0416] Reference is made to FIG. 18H In operation 47, the receiver connection state machine 1506 provides an ACK signaling message with PID value = 3 to the sender connection state machine 1502. The sender connection state machine 1502 determines that this is an ACK signaling message with PID value = inf_pid - 1 (ACK (inf_pid - 1)). This is an optional condition to return to the ACTIVE superstate, accompanied by a NAK signaling packet with PID value = inf_pid. In operation 48, the sender connection state machine 1502 exits the RECOVERY superstate and then sets enable_tx to 1 to allow new packet transmission. In operation 49, the sender connection state machine 1502 provides an ACK message with PID value = 3 to the sender packet state machine 3. The sender packet state machine 3 advances to the Pkt-Dlvr state and stops operation in operation 50. The unackd_pid value is incremented to 4.
[0417] At this point, the receiver connection state machine 1506 determines that res_count equals res_max and that the resource is again available. In operation 51, the receiver connection state machine 1506 exits the RECOVERY superstate and returns to the ACTIVE super set. Upon exiting the RECOVERY superstate, in operation 52, the receiver connection state machine 1506 provides a NAK signaling packet with a PID value of 4 (i.e., an inf_pid value) to inform the sender connection state machine that it has returned to the ACTIVE superstate. Since the sender connection state machine 1502 has entered the ACTIVE superstate, the NAK is essentially ignored. If the optional return to the ACTIVE superstate based on an ACK (inf_pid - 1) is not implemented, the receipt of the NAK (inf_pid) in this example (i.e., a NAK (PID = 4)) triggers a transition to the ACTIVE superstate, at which point the actions of the sender connection state machine 1502 described occur, causing the ACK message to the sender packet state machine 3 to be slightly late. The return to the ACTIVE superstate causes the sender packet state machines 4 and 5 to provide schedule_tx messages with PID values of 4 and 5 in operations 53 and 54, respectively. The schedule_tx() operations occur because the transition from the RECOVERY superstate to the ACTIVE superstate is an existing state of the state machine. Since the sender packet state machines 4 and 5 are in the Pkt-Pending state, reentering this state triggers the schedule_tx entry operation. In operation 55, the sender connection state machine 1502 provides a TX message to the sender packet state machine 4 to indicate that packet 4 is being transmitted. The sender packet state machine 4 advances to the Pkt-Outstanding state.
[0418] REFERENCE FIG. 14 In operation 56, packet 4 is provided to the receiver connection state machine 1506 and the sender connection state machine 1502 increments the tail_pid value to 4. In operation 57, the receiver connection state machine 1506 provides an RX message with a PID value of 4 and an expected PID value of 4 to create a receiver packet state machine 4 that starts in the packet delivered state. The receiver connection state machine 1506 increments the expected PID value to 5 and the next_epid value to 6. In operation 58, the sender connection state machine 1502 provides a TX message to the sender packet state machine 5 to cause the sender packet state machine 5 to advance to the Pkt-Outstanding state. In operation 59, the receiver packet state machine 4 provides an ACK request message for packet 4 to the receiver connection state machine 1506. In operation 60, the receiver packet state machine 4 terminates. In operation 61, the sender connection state machine 1502 provides a PID-5 packet to the receiver connection state machine 1506, then advances the tail_pid value to 5. In operation 62, the receiver connection state machine 1506 provides an RX message with a PID value of 5 and an epid value of 5, which causes the creation of a receiver packet state machine 5 in the Pkt-Deliver state. The receiver connection state machine 1506 increments the expected PID value to 6 and the next_epid value to 7. In operation 63, the receiver packet state machine 5 provides an ACK request message with a PID of 5 to the receiver connection state machine 1506 and terminates in operation 64. In operation 65, the receiver connection state machine 1506 provides an ACK response signaling packet for packet 5 to the sender connection state machine 1502, acknowledging packets 4 and 5 cumulatively. In operation 66, the sender connection state machine 1502 provides an ACK message with a PID of 5 to the sender packet state machine 4, which enters the packet delivered state. In operation 68, the sender packet state machine 4 terminates. In operation 68, the sender connection state machine 1502 forwards the ACK message with a PID of 5 to the sender packet state machine 5, which advances to the packet delivered state. The sender connection state machine 1502 increments the unackd_pid value to 6. In operation 69, the sender packet state machine 5 terminates. This completes the transmission of row (e), including the sender connection state machine 1502 and the receiver connection state machine 1506 entering the recovery state and returning to the active state. The transmission of row (e) is completed, including the sender connection state machine 1502 and the receiver connection state machine 1506 entering the recovery state and returning to the active state.
[0419] By using state machines with states and superstates (with some functionality changing between superstates), the operation in the two different modes is simplified. In the illustrated example, the use of the RECOVERY and ACTIVE superstates allows for accelerated recovery from resource deficient conditions without the need for additional reliability signaling or hardware.
[0420] While the above examples focus on SEND, RDMA READ, RDMA WRITE, and ATOMIC operations, similar inclusion of packet headers, reuse of existing header fields, state machine operations, etc. apply to other RDMA operations not specifically described, such that all RDMA operations will function under the described methods.
[0421] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples can be used in combination with each other. Many other examples will be readily ascertainable by those having ordinary skill in the art upon reviewing the above description. The scope of the application should, therefore, be determined not with reference to the above description, but instead with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the respective terms "comprising" and "wherein."< / description> < / description>
Claims
1. A remote direct memory access (RDMA) network interface controller (NIC) for connecting to a lossy Ethernet network to communicate with a remote RDMA NIC by exchanging packets with the remote RDMA NIC, the RDMA NIC comprising: a network interface for connecting to the network; an RDMA NIC processor coupled to the network interface; RDMA packet processing logic to control the receiving of packets from and the transmission of packets to the network interface; an RDMA NIC memory coupled to the RDMA NIC processor, the RDMA packet processing logic, and the network interface, the RDMA NIC memory comprising: a buffer memory for storing packets; and an RDMA NIC non-transitory memory for a program for execution by the program on the RDMA NIC processor from the RDMA NIC memory, the RDMA NIC non-transitory memory comprising an RDMA packet processing program for cooperating with the RDMA packet processing logic to perform packet transmission and reception. perform ACK and SACK signaling using iWarp reliability semantics; 2. The RDMA NIC of claim 1, wherein the RDMA packet processing logic and the RDMA packet processing program are configured to perform NAK signaling using reliability semantics of RoCE-RC; perform symmetric reliability for iWARP request and response traffic; perform one-sided single-phase transactions and two-sided single-phase transactions using iWarp and RoCE-RC operation semantics; and perform one-sided two-phase transactions using RoCE-RC operation semantics.
3. The RDMA NIC of claim 2, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a packet identifier in each request and response.
4. The RDMA NIC of claim 2, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to echo control information provided in a request message in a related response message.
5. The RDMA NIC of claim 2, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide target node RDMA addressing information in each one-sided single-phase transaction.
6. The RDMA NIC of claim 2, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a message number and a packet number within a message in each two-sided single-phase transaction, except for transactions carrying addressing information.
7. The RDMA NIC of claim 2, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a response message number and a packet number within a message in each one-sided two-phase transaction.
8. The RDMA NIC of claim 1, wherein the RDMA NIC is for connecting a computer and a network, the computer comprising: a computer processor; a memory controller coupled to the computer processor; a computer memory coupled to the computer processor and the memory controller; a computer peripheral interface coupled to the computer processor and the computer memory; and a computer non-transitory memory for a program for execution by the program from the computer memory on the computer processor, wherein the computer non-transitory memory includes a local program configured to provide a command indicative of a RDMA NIC function to be configured to the RDMA NIC, the RDMA NIC function including a requester, a responder, and a completer, the RDMA NIC further including: a RDMA NIC peripheral interface for connection to the computer, wherein the RDMA NIC processor and the RDMA NIC memory are coupled to the RDMA NIC peripheral interface, and wherein the RDMA NIC non-transitory memory further includes: a requester program for execution of requester functions; a responder program for execution of responder functions; a completer program for execution of completer functions; and a RDMA NIC function configuration program configured to interpret a received command indicative of a RDMA NIC function to be configured, and based on the received command indicative of a RDMA NIC function to be configured, cause instantiation of the requester program, the responder program, and the completer program into the RDMA NIC memory for execution by the RDMA NIC processor.
9. The RDMA NIC of claim 8, wherein the completer functions can include inbound request tracking, inbound response tracking, or both, wherein the command indicative of a RDMA NIC function to be configured further indicates which of the inbound request tracking and inbound response tracking are to be configured, and wherein the RDMA NIC function configuration program is further configured to, based on the received command indicative of a RDMA NIC function to be configured, cause activation of the inbound request tracking and inbound response tracking in the completer.
10. The RDMA NIC of claim 9, wherein the command indicative of a RDMA NIC function to be configured includes values for requester mode, send queues, receive queues, response queues, inbound request tracking, and inbound response tracking.
11. The RDMA NIC of claim 8, wherein the command indicative of a RDMA NIC function to be configured includes values for send queues, receive queues, and response queues.
12. The RDMA NIC of claim 8, wherein the RDMA NIC memory includes state machine memory for storing states of received and transmitted packets, and wherein the instantiation requestor function, the response function, or the completion function comprises: reserving memory in the RDMA NIC state machine memory for state machines associated with the requester function, responder function, or completer function, if not already reserved.
13. The RDMA NIC of claim 1, the RDMA NIC further for receiving a stream of sequentially numbered packets from the remote RDMA NIC, wherein the RDMA packet processing logic is further to provide reliability protocol elements, wherein further comprising an RDMA NIC reliability protocol procedure, and wherein the RDMA NIC reliability protocol is configured to: determine holes in the stream of sequentially numbered packets and post-hole received packets; provide a new hole indication packet to the network interface upon determining a new hole; and provide post-hole packet received packets to the network interface.
14. The RDMA NIC of claim 13, wherein the RDMA NIC reliability procedure is further configured to provide an acknowledgement of successful in-order receipt of at least one packet received prior to determining that a hole occurred and after all holes are filled to the network interface.
15. The RDMA NIC of claim 14, wherein the RDMA NIC reliability procedure is further configured to provide an acknowledgement of successful in-order receipt of at least one packet upon receiving a packet that fills the earliest hole.
16. The RDMA NIC of claim 15, wherein the new hole indication packet can indicate one or more individual holes or a range of holes.
17. The RDMA NIC of claim 13, wherein the RDMA NIC memory further comprises hole information regarding holes in the stream of received sequentially numbered packets, wherein the hole information is provided with an allocated amount of RDMA NIC memory, and wherein the RDMA NIC reliability procedure is further configured to store hole information in the RDMA NIC memory; determine when the allocated amount of RDMA NIC memory provided for hole information is exceeded; and provide a resource exhaustion packet to the network interface upon determining that the allocated amount of RDMA NIC memory provided for hole information has been exceeded.
18. The RDMA NIC of claim 17, wherein the RDMA NIC reliability procedure is further configured to provide a resource recovery packet to the network interface upon sufficient portions of the allocated memory provided for hole information being cleared.
19. The RDMA NIC of claim 1, wherein the RDMA packet processing logic is configured to detect a packet indicating that the remote RDMA NIC has exhausted out-of-order tracking resources and, in response to the detection, stop transmitting new packets to the network interface.
20. The RDMA NIC of claim 19, wherein the RDMA packet processing logic is further configured to detect a packet indicating that sufficient out-of-order tracking resources have been recovered and, in response to the detection, resume transmitting new packets to the network interface.
21. The RDMA NIC of claim 20, wherein the packet indicating that out-of-order tracking resources have been recovered indicates that all out-of-order tracking resources have been recovered.
22. The RDMA NIC of claim 19 or 20, wherein the RDMA packet processing logic is configured to detect a packet received indicating a new hole exists at the remote RDMA NIC and specifying the new hole, and wherein the RDMA packet processing logic is configured to retransmit the packet indicated by the new hole indication packet after receiving the out-of-order resource depletion packet.
23. The RDMA NIC of claim 1, the RDMA NIC providing request messages and response messages to the remote RDMA NIC, wherein the RDMA packet processing logic further constructs the RDMA packet header to include a value identifying individual packets in request packet traffic and response packet traffic, enabling request message packets to be interleaved with response message packets on a packet by packet basis without ambiguity and allowing packet reliability operations.
24. The RDMA NIC of claim 23, wherein the RDMA packet processing logic constructs the RDMA packet header to include the added header and repurposes a packet sequence number (PSN) field of a basic transport header (BTH) header, such that one of the added header and repurposed PSN field provides a message number that is incremented with each new message and remains the same for each packet in the message, and the other of the added header and repurposed PSN field provides an intra-message packet number that is incremented with each packet in the message.
25. The RDMA NIC of claim 24, wherein the RDMA NIC non-transitory memory includes an RDMA NIC reliability protocol program; and wherein the RDMA NIC reliability protocol program is configured to determine holes in a received packet stream by examining the added header and repurposed PSN field in the received packets.
26. The RDMA NIC of claim 23, wherein the RDMA packet logic constructs the RDMA packet header such that request messages use a first packet sequence number (PSN) sequence and response messages utilize a second, different PSN sequence, the first PSN sequence being incremented with each packet in a request message and the second PSN sequence being incremented with each packet in a response message.
27. The RDMA NIC of claim 26, wherein the RDMA NIC non-transitory memory includes an RDMA NIC reliability protocol program; and wherein the RDMA NIC reliability protocol program is configured to determine holes in a received packet stream by examining the first PSN sequence and the second PSN sequence.
28. A computer system, comprising: a computer, including: a computer processor; a memory controller coupled to the computer processor; a computer memory coupled to the computer processor and the memory controller; a computer peripheral interface coupled to the computer processor and the computer memory; and a network interface coupled to the computer processor. a computer non-transitory memory for a program for execution by the program from the computer memory on the computer processor, a remote direct memory access (RDMA) network interface controller (NIC) for connecting to an Ethernet network for communicating with a remote RDMA NIC by exchanging packets with the remote RDMA NIC, the RDMA NIC comprising: an RDMA NIC peripheral interface for connecting to the computer; a network interface for connecting to the network; an RDMA NIC processor coupled with the network interface; RDMA packet processing logic to control receiving packets from and transmitting packets to the network interface; an RDMA NIC memory coupled with the RDMA NIC processor, the RDMA packet processing logic, and the network interface, the RDMA NIC memory comprising: buffer memory for storing received and pending packets; and an RDMA NIC non-transitory memory for a program for execution by the program from the RDMA NIC memory on the RDMA NIC processor, the RDMA NIC non-transitory memory comprising an RDMA packet processing program for cooperating with the RDMA packet processing logic to perform packet transmission and reception.
29. The computer system of claim 28, wherein the RDMA packet processing logic and the RDMA packet processing program are configured to perform NAK signaling using the reliability semantics of RoCE-RC; to perform ACK and SACK signaling using the reliability semantics of iWarp; performing symmetric reliability for iWARP request and response traffic; performing one-sided single-phase transactions and two-sided single-phase transactions using RoCE-RC and iWarp operation semantics; and performing one-sided two-phase transactions using RoCE-RC operation semantics.
30. The computer system of claim 29, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a packet identifier in each request and response.
31. The computer system of claim 29, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to echo control information provided in a request message in a related response message.
32. The computer system of claim 29, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide target node RDMA addressing information in each one-sided single-phase transaction.
33. The computer system of claim 29, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a message number and a packet number within a message in each two-sided single-phase transaction except transactions carrying addressing information.
34. The computer system of claim 29, wherein the RDMA packet processing logic and the RDMA packet processing program are further configured to provide a response message number and a packet number within a message in each one-sided two-phase transaction.
35. The computer system of claim 28, wherein the computer non-transitory memory includes a local program configured to provide a command indicative of RDMA NIC functionality to be configured to the RDMA NIC, the RDMA NIC functionality including a requester, a responder, and a completer, and wherein the RDMA NIC non-transitory memory further comprises: a requester program for performing requester functionality; a responder program for performing responder functionality; a completer program for performing completer functionality; and a RDMA NIC functionality configuration program configured to interpret the received command indicative of RDMA NIC functionality to be configured and, based on the received command indicative of RDMA NIC functionality to be configured, cause instantiation of the requester program, the responder program, and the completer program into the RDMA NIC memory for execution by the RDMA NIC processor.
36. The computer system of claim 35, wherein the completer functionality can include inbound request tracking, inbound response tracking, or both, wherein the command indicative of RDMA NIC functionality to be configured further indicates which of the inbound request tracking and inbound response tracking are to be configured, and wherein the RDMA NIC functionality configuration program is further configured to cause activation of the inbound request tracking and inbound response tracking in the completer based on the received command indicative of RDMA NIC functionality to be configured.
37. The computer system of claim 36, wherein the command indicative of RDMA NIC functionality to be configured includes values for requester mode, send queue, receive queue, response queue, inbound request tracking, and inbound response tracking.
38. The computer system of claim 35, wherein the command indicative of RDMA NIC functionality to be configured includes values for send queue, receive queue, and response queue.
39. The computer system of claim 35, wherein the RDMA NIC memory includes state machine memory for storing states of received and transmitted packets, and wherein the instantiation requestor function, the response function, or the completion function comprises: reserving memory in the RDMA NIC state machine memory for state machines related to the requester functionality, responder functionality, or completer functionality, if not already reserved.
40. The computer system of claim 28, wherein the RDMA NIC is to receive a stream of sequentially numbered packets from the remote RDMA NIC, wherein the RDMA packet processing logic is further to provide reliability protocol elements, wherein the RDMA NIC non-transitory memory further includes a RDMA NIC reliability protocol program, and wherein the RDMA NIC reliability protocol is configured to determine holes in the stream of sequentially numbered packets and packets received after holes and, upon determining a new hole, provide a new hole indication packet to the network interface and provide packets received after holes to the network interface for packets received after holes.
41. The computer system of claim 40, wherein the RDMA NIC reliability protocol program is configured to determine holes in the stream of sequentially numbered packets and packets received after holes and, upon determining a new hole, provide a new hole indication packet to the network interface and provide packets received after holes to the network interface for packets received after holes.
42. The computer system of claim 41, wherein the RDMA NIC reliability protocol program is configured to determine holes in the stream of sequentially numbered packets and packets received after holes and, upon determining a new hole, provide a new hole indication packet to the network interface and provide packets received after holes to the network interface for packets received after holes.
43. The computer system of claim 42, wherein the RDMA NIC reliability protocol program is configured to determine holes in the stream of sequentially numbered packets and packets received after holes and, upon determining a new hole, provide a new hole indication packet to the network interface and provide packets received after holes to the network interface for packets received after holes.
44. The computer system of claim 43, wherein the RDMA NIC reliability protocol program is configured to determine holes in the stream of sequentially numbered packets and packets received after holes and, upon determining a new hole, provide a new hole indication packet to the network interface and provide packets received after holes to the network interface for packets received after holes.
41. The computer system of claim 40, wherein the RDMA NIC reliability program is further configured to provide the network interface with an acknowledgement of successful in-order receipt of at least one packet received prior to a determination that a hole has occurred and after all holes have been filled.
42. The computer system of claim 41, wherein the RDMA NIC reliability program is further configured to provide the acknowledgement of successful in-order receipt of at least one packet upon receiving a packet that at least fills an earliest hole.
43. The computer system of claim 42, wherein the new hole indication packet can indicate one or more individual holes or hole ranges.
44. The computer system of claim 40, wherein the RDMA NIC memory further comprises information about holes in a received stream of sequentially numbered packets, wherein an amount of RDMA NIC memory is provided for the hole information, wherein the RDMA NIC reliability program is further configured to store hole information in the RDMA NIC memory, determine when the amount of RDMA NIC memory provided for hole information has been exceeded, and provide a resource-exhausted packet to the network interface upon determining that the amount of RDMA NIC memory provided for hole information has been exceeded.
45. The computer system of claim 44, wherein the RDMA NIC reliability program is further configured to provide a resource-recovered packet to the network interface upon sufficient portions of the allocated memory for hole information being cleared.
46. The computer system of claim 28, wherein the RDMA packet processing logic is configured to detect receipt of a packet indicating that the remote RDMA NIC has exhausted out-of-order tracking resources, and in response to the detection, stop transmitting new packets to the network interface.
47. The computer system of claim 46, wherein the RDMA packet processing logic is further configured to detect a packet indicating that sufficient out-of-order tracking resources have been recovered, and in response to the detection, resume transmitting new packets to the network interface.
48. The computer system of claim 47, wherein the packet indicating that out-of-order tracking resources have been recovered indicates that all out-of-order tracking resources have been recovered.
49. The computer system of claim 46, wherein the RDMA packet processing logic is configured to detect a packet that is a new hole indication packet indicating that a new hole exists at the remote RDMA NIC and specifying the new hole, and wherein the RDMA packet processing logic is configured to retransmit packets indicated by the new hole indication packet after receiving the out-of-order resource-exhausted packet.
50. The computer system of claim 28, wherein the RDMA NIC communicates with the remote RDMA NIC by providing request messages and response messages to the remote RDMA NIC, and wherein the RDMA packet processing logic is configured to detect a packet that is a new hole indication packet indicating that a new hole exists at the remote RDMA NIC and specifying the new hole, and wherein the RDMA packet processing logic constructs the RDMA packet header to include values that identify individual packets in request packet traffic and response packet traffic, such that request message packets can be interleaved with response message packets on a packet by packet basis without ambiguity, and to allow packet reliability operations.
51. The computer system of claim 50, wherein the RDMA packet processing logic constructs the RDMA packet header to include the added header and to repurpose a packet sequence number (PSN) field of a basic transport header (BTH) header, such that one of the added header and repurposed PSN field provides a message number that is incremented with each new message and remains the same for each packet in the message, and the other of the added header and repurposed PSN field provides a packet within message number that is incremented with each packet in the message.
52. The computer system of claim 51, wherein the RDMA NIC non-transitory memory includes an RDMA NIC reliability protocol program; and wherein the RDMA NIC reliability protocol program is configured to determine holes in a received packet stream by examining the added header and repurposed PSN field in received packets.
53. The computer system of claim 50, wherein the RDMA packet processing logic constructs the RDMA packet header such that request messages use a first packet sequence number (PSN) sequence and response messages utilize a second, different PSN sequence, the first PSN sequence being incremented with each packet in a request message and the second PSN sequence being incremented with each packet in a response message.
54. The computer system of claim 53, wherein the RDMA NIC non-transitory memory includes an RDMA NIC reliability protocol program; and wherein the RDMA NIC reliability protocol program is configured to determine holes in a received packet stream by examining the first PSN sequence and the second PSN sequence.
55. A method of operating a remote direct memory access (RDMA) network interface controller (NIC) for connecting to a lossy Ethernet network to communicate with a remote RDMA NIC by exchanging packets with the remote RDMA NIC, the RDMA NIC comprising: a network interface for connecting to the network; an RDMA NIC processor coupled with the network interface; RDMA packet processing logic to control receiving packets from and transmitting packets to the network interface; an RDMA NIC memory coupled with the RDMA NIC processor, the RDMA packet processing logic, and the network interface, the RDMA NIC memory comprising: a buffer memory for storing packets; and a non-transitory memory for storing a RDMA NIC reliability protocol program. RDMA NIC non-transitory memory for a program for execution by the program from the RDMA NIC memory on the RDMA NIC processor, the RDMA NIC non-transitory memory including RDMA packet processing programs for cooperating with the RDMA packet processing logic to perform packet transmission and reception; the method comprising: performing RDMA operations.
56. The method of claim 55, further comprising: performing NAK signaling using reliability semantics of RoCE-RC; performing ACK and SACK signaling using reliability semantics of iWarp; performing symmetric reliability for request and response traffic of iWarp; performing one-sided one-phase transactions and two-sided one-phase transactions using operation semantics of RoCE-RC and iWarp; and performing one-sided two-phase transactions using operation semantics of RoCE-RC.
57. The method of claim 56, further comprising providing a packet identifier in each request and response.
58. The method of claim 56, further comprising echoing control information provided in a request message in a related response message.
59. The method of claim 56, further comprising providing target node RDMA addressing information in each one-sided one-phase transaction.
60. The method of claim 56, further comprising providing a message number and a packet number within a message in each two-sided one-phase transaction except transactions carrying addressing information.
61. The method of claim 56, further comprising providing a response message number and a packet number within a message in each one-sided two-phase transaction.
62. The method of claim 55, wherein the RDMA NIC is for connection to a computer, the computer comprising: a computer processor; a memory controller coupled to the computer processor; a computer memory coupled to the computer processor and the memory controller; a computer peripheral interface coupled to the computer processor and the computer memory; and computer non-transitory memory for a program for execution by the program from the computer memory on the computer processor, wherein the computer non-transitory memory includes a local program configured to provide commands indicative of RDMA NIC functions to be configured to the RDMA NIC, the RDMA NIC functions including a requester, a responder, and a completer, the RDMA NIC further comprising: an RDMA NIC peripheral interface for connection to the computer, wherein the RDMA NIC processor and the RDMA NIC memory are coupled with the RDMA NIC peripheral interface, and the method further comprising: configuring the RDMA NIC functions to the RDMA NIC. wherein the RDMA NIC non-transitory memory further comprises: a requester program for performing requester functionality; a responder program for performing responder functionality; a completer program for performing completer functionality; and a RDMA NIC functionality configuration program configured to interpret a received command indicating RDMA NIC functionality to be configured, the method further comprising: based on the received command indicating RDMA NIC functionality to be configured, causing instantiation of the requester program, the responder program, and the completer program into the RDMA NIC memory for execution by the RDMA NIC processor.
63. The method of claim 62, wherein the completer functionality can include inbound request tracking, inbound response tracking, or both, wherein the command indicating RDMA NIC functionality to be configured further indicates which of the inbound request tracking and inbound response tracking are to be configured, the method further comprising: based on the received command indicating RDMA NIC functionality to be configured, activating inbound request tracking and inbound response tracking in the completer.
64. The method of claim 62, wherein the command indicating RDMA NIC functionality to be configured includes values for requester mode, send queues, receive queues, response queues, inbound request tracking, and inbound response tracking.
65. The method of claim 62, wherein the command indicating RDMA NIC functionality to be configured includes values for send queues, receive queues, and response queues.
66. The method of claim 62, wherein the RDMA NIC memory includes state machine memory for storing state of received and transmitted packets, and wherein the instantiation requestor function, the response function, or the completion function comprises: reserving memory in the RDMA NIC state machine memory for state machines associated with the requester functionality, responder functionality, or completer functionality, if not already reserved.
67. The method of claim 55, wherein the RDMA NIC is further for receiving a stream of sequentially numbered packets from the remote RDMA NIC, wherein the RDMA packet processing logic provides reliability protocol elements, and wherein the RDMA NIC non-transitory memory further comprises a RDMA NIC reliability protocol program, the method comprising: determining holes in the stream of sequentially numbered packets and packets received after holes; upon determining a new hole, providing a new hole indication packet to the network interface; and providing packets received after a hole packet to the network interface for packets received after holes.
68. The method of claim 67, further comprising providing an acknowledgement of successful in-order receipt of at least one packet received prior to a determination that a hole occurred and after all holes have been filled to the network interface.
69. The method of claim 68, further comprising providing an acknowledgement of successful in-order receipt of at least one packet upon receiving a packet that fills at least an earliest hole.
70. The method of claim 69, wherein the new hole indication packet can indicate one or more individual holes or hole ranges.
71. The method of claim 67, wherein the RDMA NIC memory further comprises information about holes in a received stream of sequentially numbered packets, and wherein an allocated amount of RDMA NIC memory is provided for the hole information, the method further comprising: storing hole information in the RDMA NIC memory; determining when the allocated amount of RDMA NIC memory provided for hole information has been exceeded; and providing a resource exhaustion packet to the network interface upon determining that the allocated amount of RDMA NIC memory provided for hole information has been exceeded.
72. The method of claim 71, further comprising providing a resource recovery packet to the network interface upon sufficient portions of the allocated memory provided for hole information being cleared.
73. The method of claim 55, further comprising: detecting receipt of a packet indicating that the remote RDMA NIC has exhausted out-of-order tracking resources; and stopping transmission of new packets to the network interface in response to the detection.
74. The method of claim 73, further comprising: detecting a packet indicating that out-of-order tracking resources have been recovered, and resuming transmission of new packets to the network interface in response to the detection.
75. The method of claim 74, wherein the packet indicating that out-of-order tracking resources have been recovered indicates that all out-of-order tracking resources have been recovered.
76. The method of claim 73, further comprising: detecting a packet indicating receipt of a packet indicating a new hole at the remote RDMA NIC and specifying the new hole, and retransmitting packets indicated by the new hole indication packet after receipt of the out-of-order resource exhaustion packet.
77. The method of claim 55, wherein the RDMA NIC further communicates with the remote RDMA NIC by providing request messages and response messages to the remote RDMA NIC, the method further comprising: including in the RDMA packet header a value identifying individual packets in request packet traffic and response packet traffic such that request message packets can be interleaved with response message packets on a packet by packet basis without ambiguity and allowing packet reliability operations.
78. The method of claim 77, further comprising constructing the RDMA packet header to include an added header and reusing a packet sequence number (PSN) field of a basic transport header (BTH) header such that one of the added header and the reused PSN field provides a message number that is incremented with each new message and remains the same for each packet in the message, and the other of the added header and the reused PSN field provides a packet within message number that is incremented with each packet in the message.
79. The method of claim 78, wherein the RDMA NIC non-transitory memory comprises an RDMA NIC reliability protocol program; and The method further comprises: determining holes in the received packet stream by examining the added header and the re-used PSN field in the received packets.
80. The method of claim 77, further comprising constructing the RDMA packet header such that request messages use a first packet sequence number (PSN) sequence and response messages utilize a second, different PSN sequence, the first PSN sequence increasing with each packet in a request message and the second PSN sequence increasing with each packet in a response message.
81. The method of claim 80, wherein the RDMA NIC non-transitory memory comprises an RDMA NIC reliability protocol program, and The method further comprises: determining holes in the received packet stream by examining the first PSN sequence and the second PSN sequence.
82. The method of claim 77, wherein the RDMA NIC non-transitory memory comprises an RDMA NIC reliability protocol program, and determining holes in the received packet stream by examining the added header and the re-used PSN field in the received packets.