System and method for facilitating efficient message reconciliation in a network interface controller (NIC)
The NIC with a hardware list processing engine addresses scalability and efficiency issues by accelerating MPI list reconciliation through parallel processing and protocol support, improving network performance in diverse traffic scenarios.
Patent Information
- Application Number
- DE112020002754
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-23
- Filing Date
- 2020-03-23
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-03-23
AI Technical Summary
Existing network interface controllers face challenges in scalability, versatility, and efficiency in handling diverse network traffic types and increasing loads, particularly in applications like high-performance computing and IoT, necessitating improved performance metrics such as bandwidth and latency.
A network interface controller (NIC) equipped with a hardware list processing engine (LPE) that performs high-speed MPI list reconciliation, utilizing multiple processing devices and memory banks interconnected via a crossbar, and supports both eager and rendezvous protocols to accelerate list matching and reduce latency.
The NIC achieves efficient list matching by overlapping matching pipeline stages, using parallel processing, and separating endpoint interfaces, thereby enhancing network performance and reducing latency in handling diverse traffic types.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND area
[0001] This generally relates to the technical field of networking. More specifically, this disclosure relates to systems and methods for facilitating high-speed Message Passing Interface (MPI) list matching in a network interface controller (NIC). State of the art
[0002] As network-enabled devices and applications become increasingly ubiquitous, different types of traffic and ever-increasing network loads demand ever-greater performance from the underlying network architecture. For example, applications such as high-performance computing (HPC), media streaming, and the Internet of Things (IoT) can generate different types of traffic with distinct characteristics. As a result, network architects continue to face challenges such as scalability, versatility, and efficiency in addition to traditional network performance metrics such as bandwidth and latency.
[0003] US 2017 / 0054633 A1 relates to high performance computing and in particular to message passing in a network interface controller.
[0004] US 2019 / 0044875 A1 relates generally to the field of computer science and / or networking and, in particular, to the communication of a large message using multiple network interface controllers.
[0005] Against this background, it is the object of the present disclosure to provide a network interface controller and a method which at least partially solve the disadvantages of known network interface controllers. SUMMARY
[0006] A network interface controller and method having the features of independent claims 1 and 10, respectively, is disclosed, which is capable of performing an MPI list reconciliation. The network interface controller, hereinafter also referred to as NIC, may comprise a host interface, a network interface, and a hardware list processing engine (LPE). The host interface may connect the NIC to a host device. The NIC may be connected to a network via the network interface. During operation, the LPE may receive a reconciliation request and perform an MPI list reconciliation based on the received reconciliation request. BRIEF DESCRIPTION OF THE CHARACTERS Fig. shows an example network. Fig. shows an example NIC chip with a plurality of NICs. Fig. shows an example architecture of a NIC. Fig. shows an example architecture of a processing engine. Fig. shows an example operation pipeline of a matching engine. Fig. shows example queues for reconciliation requests. Fig. shows an example block diagram of a persistent list entry cache (PLEC). Fig. shows a flowchart for performing list matching in a NIC. Fig. shows an example computer system equipped with a NIC that facilitates MPI list matching.
[0007] In the figures, the same numbers refer to the same elements of the figure. DETAILED DESCRIPTION
[0008] Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the present disclosure. Therefore, the present invention is not limited to the embodiments shown. Overview
[0009] The present disclosure describes systems and methods that facilitate efficient list matching in a network interface controller (NIC). The NIC implements a hardware list processing engine coupled to a memory device. The list processing engine can perform high-speed list matching. The list processing engine (LPE) can perform atomic search and search-with-delete operators on the various lists defined by the Message Passing Interface (MPI) protocol and can route list operations to the appropriate matching devices. To increase speed, multiple processing devices can be used, and each processing device can include multiple memory banks interconnected via a crossbar. Furthermore, the LPE accelerates list matching by separating the endpoint network interfaces.The list matching hardware can reduce latency by overlapping the state of the matching attempt pipeline with the matching termination condition and can use a unified search pipeline for priority and unexpected lists, as well as for network search and host append commands. The LPE hardware can also use a unified processing pipeline to search persistent list entries related to an unordered network interface, as well as to search entries related to an ordered network interface. The NIC can also process MPI messages by efficiently using either the eager or rendezvous protocol.
[0010] One embodiment provides a NIC capable of performing MPI list matching. The NIC may include a host interface, a network interface, and a hardware LPE. The host interface may connect the NIC to a host device. The network interface may connect the NIC to a network. During operation, the LPE may receive a matching request and perform MPI list matching based on the received matching request.
[0011] In a variant of this embodiment, the reconciliation request may comprise a reconciliation request corresponding to a command received via the host interface or a reconciliation request corresponding to a message received via the network interface.
[0012] In another variation, the NIC may include a first set of reconciliation request queues for reconciliation requests corresponding to received commands and a second set of reconciliation request queues for reconciliation requests corresponding to received messages. The number of queues in the first or second set of reconciliation request queues corresponds to the number of physical endpoints supported by the NIC.
[0013] In another variant, the message is an MPI message.
[0014] In another variant, the message is based on an eager protocol or a rendezvous protocol connected to MPI.
[0015] In a variant of this embodiment, the hardware list processing engine may comprise a plurality of processing elements; and a respective processing element comprises a plurality of matching engines and a plurality of memory banks storing one or more lists, the memory banks being connected to the matching engines via a crossbar.
[0016] In a further variation, a corresponding matching engine may include a unified search pipeline for searching the one or more lists, and the one or more lists may include a priority list and an unexpected list.
[0017] In a further variant, a corresponding matching engine may comprise a single pipeline stage to perform in parallel a matching operation for a previous matching request and a calculation to determine a current read or write address.
[0018] In a variation of this embodiment, the hardware list processing engine may include a persistent list entry cache to store previously matched list entries and enable fast searching.
[0019] In a variant of this embodiment, the list processing engine may perform atomic search operations in a plurality of lists.
[0020] In this disclosure, the description in connection with Fig. on the network architecture, and the descriptions in connection with Fig. and following provide further details about the architecture and operations associated with a NIC that supports efficient MPI list matching.
[0021] Fig. shows an example network. In this example, a network 100 of switches, which may also be referred to as a "switch fabric," may include switches 102, 104, 106, 108, and 110. Each switch may have a unique address or ID within the switch fabric 100. Various types of devices and networks may be connected to a switch fabric. For example, a storage array 112 may be connected to the switch fabric 100 via switch 110; an InfiniBand (IB)-based HPC network 114 may be connected to the switch fabric 100 via switch 108; a number of end hosts, such as host 116, may be connected to the switch fabric 100 via switch 104; and an IP / Ethernet network 118 can be connected to the switch fabric 100 via the switch 102. In general, a switch can have edge ports and fabric ports. An edge port can be connected to a device located outside the fabric.A fabric port can be connected to another switch within the fabric via a fabric link. Typically, traffic can enter the switch fabric 100 through an ingress port of an edge switch and exit the switch fabric 100 through an egress port of another (or the same) edge switch. An ingress link can connect a NIC of an edge device (e.g., an HPC end host) to an ingress edge port of an edge switch. The switch fabric 100 can then transport the traffic to an egress edge switch, which can, in turn, forward the traffic to a destination edge device through another NIC. Exemplary NIC architecture
[0022] Fig. shows an exemplary NIC chip with a plurality of NICs. With respect to the example in Fig. A NIC chip 200 may be an application-specific integrated circuit (ASIC) designed for the host 116 to operate with the switch fabric 100. In this example, the chip 200 may provide two independent NICs 202 and 204. A corresponding NIC of the chip 200 may be equipped with a host interface (HI) (e.g., an interface for connecting to the host processor) and a high-speed network interface (HNI) for communicating with a network interface connected to the switch fabric 100 of Fig. coupled connection. For example, NIC 202 may include an HI 210 and an HNI 220, and NIC 204 may include an HI 211 and an HNI 221.
[0023] In some embodiments, HI 210 may be a Peripheral Component Interconnect (PCI) or a Peripheral Component Interconnect Express (PCIe) interface. HI 210 may be coupled to a host via a host connection 201, which may include N (e.g., N may be 16 for some chips) PCIe Gen 4 lanes capable of operating at signaling rates of up to 25 Gbps per lane. HNI 210 may enable a high-speed network connection 203 that can interface with an interconnect in the switch fabric 100 of Fig. can communicate. The HNI 210 can operate at total rates of either 100 Gbit / s or 200 Gbit / s using M (e.g., M can be 4 on some chips) full-duplex serial lanes. Each of the M lanes can operate at 25 Gbit / s or 50 Gbit / s based on non-return-to-zero (NRZ) or pulse amplitude modulation 4 (PAM4), respectively. The HNI 220 supports the Institute of Electrical and Electronics Engineers (IEEE) Ethernet-based 802.3 protocols, as well as an extended frame format that enables higher rates for small messages.
[0024] The NIC 202 may support one or more of the following: point-to-point messaging based on the Message Passing Interface (MPI), remote memory access (RMA) operations, offloading and progression of bulk data collection operations, and Ethernet packet processing. When the host issues an MPI message, the NIC 202 may correspond to the corresponding message type. In addition, the NIC 202 may implement both the Eager Protocol and the Rendezvous Protocol for MPI, thus offloading the corresponding operations from the host.
[0025] Additionally, the RMA operations supported by NIC 202 may include PUT, GET, and atomic memory operations (AMO). NIC 202 may provide reliable transport. For example, if NIC 202 is a source NIC, NIC 202 may provide a retry mechanism for idempotent operations. Furthermore, connection-based error detection and a retry mechanism may be used for ordered operations that may manipulate a target state. The hardware of NIC 202 may maintain the state required by the retry mechanism. In this way, NIC 202 may offload the host (e.g., software). The policy dictating the retry mechanism may be set by the host through the driver software, ensuring the flexibility of NIC 202.
[0026] In addition, the NIC 202 may facilitate triggered operations, a general mechanism for offloading and scheduling dependent sequences of operations, such as bulk data collections. The NIC 202 may support an application programming interface (API) (e.g., libfabric API) that facilitates fabric communication services provided by the switch fabric 100 of Fig. for applications on host 116. NIC 202 may also support a low-level network programming interface, such as Portals API. Furthermore, NIC 202 may provide efficient Ethernet packet processing, which may include efficient transmission when NIC 202 is a sender, flow control when NIC 202 is a destination, and checksum calculation. Furthermore, NIC 202 may support virtualization (e.g., with containers or virtual machines).
[0027] Fig. shows an example architecture of a NIC. In NIC 202, the port macro of HNI 220 may facilitate low-level Ethernet operations, such as Physical Coding Sublayer (PCS) and Media Access Control (MAC). Additionally, NIC 202 may provide support for Link Layer Retry (LLR). Incoming packets may be parsed by parser 228 and stored in buffer 229. Buffer 229 may be a PFC buffer that buffers a threshold (e.g., one microsecond) of delay bandwidth. HNI 220 may also include a control transmit unit 224 and a control receive unit 226 for managing outgoing and incoming packets, respectively.
[0028] The NIC 202 may include a command queue (CQ) unit 230. The CQ unit 230 may be responsible for fetching and issuing host-side commands. The CQ unit 230 may include command queues 232 and schedulers 234. The command queues 232 may include two independent sets of queues for initiator commands (PUT, GET, etc.) and target commands (Append, Search, etc.), respectively. The command queues 232 may be implemented as circular buffers. In some embodiments, the command queues 232 may be maintained in the host's main memory. Applications running on the host may write directly to the command queues 232. The schedulers 234 may include two separate schedulers for initiator commands and target commands, respectively. The initiator commands are sorted into flow queues 236 based on a hash function. One of the flow queues 236 can be assigned to a unique flow.In addition, the CQ unit 230 may include a triggered operations module (or logic block) 238 that is responsible for queuing and dispatching triggered instructions.
[0029] The outbound transfer engine (OXE) 240 may retrieve commands from the expiration queues 236 to process them for dispatch. OXE 240 may include an address translation request unit (ATRU) 244 that may send address translation requests to the address translation unit (ATU) 212. The ATU 212 may perform virtual-to-physical address translation on behalf of various engines, such as the OXE 240, the inbound transfer engine (IXE) 250, and the event engine (EE) 216. The ATU 212 may maintain a large translation cache 214. The ATU 212 may either perform the translation itself or utilize host-based address translation services (ATS). OXE 240 may also include a message chopping unit (MCU) 246, which can fragment a large message into packets of a size corresponding to a maximum transmission unit (MTU). MCU 246 may include a plurality of MCU modules.When an MCU module becomes available, the MCU module can retrieve the next command from an assigned expiration queue. The data received from the host can be written to data buffer 242. The MCU module can then send the packet header, the appropriate traffic class, and the packet size to traffic shaper 248. Traffic shaper 248 can determine which of the requests presented by MCU 246 can be forwarded to the network.
[0030] The selected packet may then be sent to the Packet and Connection Trace (PCT) 270. PCT 270 may store the packet in a queue 274. PCT 270 may also maintain status information for outgoing commands and update the status information when responses are returned. PCT 270 may also maintain packet status information (e.g., to match responses with requests), message status information (e.g., to track the progress of messages containing multiple packets), initiator completion status information, and retry status information (e.g., to obtain the information needed to retry a command if a request or response is lost). If a response does not return within a specified period of time, the corresponding command may be stored in the retry buffer 272.PCT 270 may facilitate connection management for initiator and target commands based on source tables 276 and target tables 278, respectively. For example, PCT 270 may update its source tables 276 to track the necessary state for reliable packet delivery and message completion notification. PCT 270 may forward outgoing packets to HNI 220, which stores the packets in the outgoing queue 222.
[0031] The NIC 202 may also include an IXE 250, which handles packet processing when the NIC 202 is a target or destination. IXE 250 may retrieve the incoming packets from HNI 220. Parser 256 may parse the incoming packets and forward the appropriate packet information to a List Processing Engine (LPE) 264 or a Message State Table (MST) 266 for matching. LPE 264 may match incoming messages against buffers. LPE 264 may determine the buffer and starting address to be used by each message. LPE 264 may also maintain a pool of list entries 262 used to represent buffers and unexpected messages. MST 266 may store matching results and the information required to generate destination-side completion events.MST 266 can be used by unrestricted operations, including multi-packet PUT instructions and single- or multi-packet GET instructions.
[0032] Parser 256 may then store the packets in packet memory 254. IXE 250 may retrieve the results of the matching for conflict checking. DMA Write and AMO module 252 may then perform memory updates generated by write and AMO operations. If a packet contains a command that generates target-side memory reads (e.g., a GET response), the packet may be forwarded to OXE 240. NIC 202 may also include an Event Engine (EE) 216, which may receive requests to generate event notifications from other modules or devices on NIC 202. An event notification may indicate that either a fill or a count event is generated. EE 216 may maintain event queues located in host processor memory to which it writes full events. EE 216 may forward count events to CQ unit 230. MPI list comparison
[0033] In MPI, send / receive operations are identified by an envelope, which can contain a set of parameters such as source, destination, message ID, and communicator. The envelope can be used to map a specific message to the corresponding user buffer. The entire list of buffers sent by a given process can be referred to as a matching list, and the process of searching for the corresponding buffer from the matching list to a specific buffer is called list matching or tag matching.
[0034] In some embodiments, the NIC may provide hardware acceleration of MPI list matching, and the list processing engine in the NIC may include a plurality (e.g., 2048) of physical endpoints. Each physical endpoint may contain four lists: priority, overflow, unexpected, and software request. The software request list may enable a smooth transition from hardware offload to software-managed lists. The priority, overflow, and request lists contain entries that contain matching criteria and memory descriptor information. The unexpected list contains header information of messages for which a list entry has not been created in advance. The NIC's LPE block may provide memory for a number (e.g.,64k) of list entries distributed among the matching entries (for the matching interface), the list entries (for the non-matching interface), and the unexpected list entries.
[0035] In some embodiments, the NIC's LPE block can be divided into multiple (e.g., four) processing engines, allowing the LPE to leverage process-level parallelism in applications or workloads. Each processing engine can access a subset of the list entries. For example, if the LPE block contains a total of 64k list entries and there are four processing engines, each processing engine can access 16k list entries. Software can be responsible for assigning physical endpoints to the processing engines to ensure load balancing.
[0036] The LPE can contain two interfaces for list matching: one interface that receives target-side commands from the CQ unit, and the other interface that receives message matching requests from an IXE. The IXE sends the first packet of each message to the LPE; the LPE searches the corresponding lists. If a matching entry is found, it can be uncoupled and returned to the IXE; otherwise, the header can be appended to the unexpected list. Each interface can be designated as matching or non-matching, depending on the physical endpoint's configuration. CQ command requests and IXE network requests are both referred to as matching requests.
[0037] In some embodiments, the interfaces for MPI can be initialized in the disabled state. Message matching of incoming traffic occurs only in the hardware offload state. Specifically, the processing engine can perform atomic search and search-with-delete operations on the priority, overflow, and unexpected lists. During the search, the processing engine can forward list operations to a correct matching unit.
[0038] Fig. shows the exemplary architecture of a processing engine. In this example, processing engine 300 may include a plurality of matching engines (e.g., matching engines 302 and 304) and four memory banks (e.g., memory banks 306, 308, 310, and 312).
[0039] In some embodiments, processing engine 300 may include up to eight matching engines. The memory banks may be connected to the matching engines via a crossbar to minimize bank contention and achieve high parallelism and utilization of the matching engines. Each matching engine may generate a memory address for one of the memory banks in processing engine 300. The multiple matching engines (e.g., matching engines 302 and 304) operate independently of each other. However, these multiple matching engines must mediate access to the memory banks.
[0040] Fig. shows an example matching engine operation pipeline. The matching engine pipeline 320 may include a number of stages: a setup stage 322, a read address and match stage 324, a read data stage 326, a correct read data stage 328, a mux match entry stage 330, a write address stage 332, and a write data stage 334.
[0041] In the setup stage 322, the match engine collects the match request information from the ready request queue (RRQ). In the read and match stage 324, the match engine initiates the read request in each memory bank. Each match engine may have logic that decides whether and to which memory bank a read or write request should be made. In some embodiments, each memory bank may have an arbiter used to select a matching engine and multiplex the address. If there are eight parallel match engines, the arbiter may be an 8:1 arbiter. In parallel with calculating the read address, the read and match stage 324 also checks whether there is a match with the previous match entry. If there is a match, it prepares the write update (calculating a new offset or deleting an entry).The address and data are then registered in the memory bank at the write address stage 332 and the data write stage 334. In the write address stage 332, the matching engine begins the write access, and in the data write stage 334, the matching engine completes the write operation.
[0042] In the read data stage 326, the read data is registered at the output of each memory bank. In the read data correction stage 328, the read data is corrected in the memory bank. In the mux match entry stage 330, a multiplexer in each match engine captures the match entry containing the new current address. A series of inner loops are executed, with each loop including stage 324 (read address and match), stage 326 (read data), stage 328 (correction of the read data), and stage 330 (mux match entry). In the case of four memory banks, the matching engine pipeline 320 may include four cycles. Each match engine contains storage space for the result of each operation. An arbiter selects a result from the multiple match engines to send to the output arbiter block.When the output arbiter block consumes a result, the matching engine that produced the result can retrieve another instruction from the RRQ.
[0043] The Fig. The pipeline shown can provide several advantages. First, the overlap between the matching attempt pipeline stage and the matching termination condition (e.g., the address read and match stage 324) can reduce latency in the matching engine. Second, pipeline 320 can provide a unified search pipeline for priority and unexpected list searches, as well as network searches and host append commands.
[0044] To increase concurrency and avoid endpoint and traffic class congestion, in some embodiments the network interface card may accelerate list matching by separating queues, with each endpoint network interface having its own queue. Specifically, the match request queues ensure that for matching interfaces, only one operation is processed per physical endpoint; for unmatched interfaces, concurrent access to certain persistent list entries may be permitted. Within a physical endpoint, command requests must be executed in the order in which they arrive, and network match requests must be executed in the order in which they arrive. However, the order of commands and network requests is not prescribed.The separate queues also ensure that requests from one physical endpoint cannot be blocked by requests from another physical endpoint. Likewise, requests in one traffic class cannot block requests in other traffic classes.
[0045] Fig. shows example queues for reconciliation requests. The reconciliation request queuing block 400 may include two sets of queues. The first set of queues includes CQ queues (MRQs) 402 for queuing CQ commands and IXE queues 404 for queuing IXE requests, with each queue indexed by the physical portal index. Each physical endpoint corresponds to a CQ queue for reconciliation requests and an IXE queue for reconciliation requests. For a NIC supporting 2048 physical endpoints, the CQ queues 402 may contain 2048 CQ queues for reconciliation requests and the IXE queues 404 may contain 2048 IXE queues for reconciliation requests.
[0046] One or more arbitrators 406 may be used to select between CQ queues 402 and IXE queues 404, and to select between the multiple queues within each type of queue. In some embodiments, a standard arbitration mechanism (e.g., round robin) may be used for arbitration.
[0047] When a reconciliation request is removed from one of these queues, a lookup table 408 is checked to determine the processing engine (PE) for the physical portal index. The lookup table 408 may be an array of flops containing the processing engine number for each physical endpoint and may be accessed in parallel. The reconciliation request is then placed in a corresponding processing engine / traffic class reconciliation request queue belonging to the second tier of queues (processing engine / traffic class (PE / TC) MRQs 410), unless it is an IXE request, which fits into the persistent list entry cache (PLEC) 412. A detailed description of the PLEC 412 follows. An arbitrator 414 may select among the PE / TC MRQs 410, and a multiplexer 416 may multiplex the output of arbitrator 414 and PLEC 412.
[0048] To further increase the speed of list matching, in some embodiments, the system may also use a unified processing pipeline to search persistent list entries related to an unordered network interface and entries related to an ordered network interface. In particular, the PLEC enables very fast searches with a unit delay.
[0049] Fig. shows an example block diagram of a persistent list entry cache (PLEC). The PLEC stores a number of entries (e.g., up to 256) that match the physical portal index. When a physical endpoint has an entry in the cache, the PLEC allows its queue for the physical endpoint's reconciliation request to be processed at full speed and without blocking.
[0050] When dequeueing an IXE match request queue (MRQ) for a physical endpoint that matches in the PLEC, the PLEC forwards the list entry (LE) to the memory block that stores match requests. When dequeuing the CQ MRQ or dequeuing the IXE MRQ that is then missing from the PLEC, a lock bit is set for the physical endpoint. The PLEC maintains a lock bit for each physical endpoint to ensure that matching requests and commands are processed atomically, while unmatched IXE requests to qualified persistent list entries are satisfied without blocking.
[0051] The PLEC intercepts IXE requests that match in its cache before enqueuing them in the processing engine / traffic class queue. When a persistent list entry is copied from the cache, no eviction from the processing engine / traffic class queue is initiated that cycle, allowing the persistent connection entry (LE) to proceed through the pipeline to the physical endpoint's memory. Specifically, when a PLEC hit occurs, a dequeue from the PE / TC MRQ is suppressed to create a bubble in the pipeline. The dequeue is suppressed when the PLEC memory (i.e., the LE cache) is read, so that the PLEC data is available when the bubble occurs. The LE from the PLEC and its match request ID can be forwarded to the physical endpoint's memory block.
[0052] The PLEC receives allocation and deallocation requests from the processing engines. An allocation request arrives when a processing engine matches a network request with a persistent LE on the priority list that has packet matching events disabled at a mismatched, non-space-checking physical endpoint. A physical endpoint allocation request that encounters an existing entry in the PLEC has no effect; otherwise, an entry is allocated. When the cache is full, an entry is evicted using round-robin selection. When a processing engine decouples a cacheable list entry, it sends a deallocation request to the PLEC. If the PLEC contains an entry with a matching physical endpoint, the corresponding entry is removed from the PLEC.
[0053] The LPE block on the NIC plays an important role in processing MPI messages. As mentioned earlier, MPI implements the "eager" protocol for processing small messages and the "rendezvous" protocol for processing large messages. Specifically, "eager" means that the data is sent along with the PUT (message) command. The system software sets an upper limit for eager messages. For messages whose size exceeds the eager message limit, MPI requires that the messages be sent using the rendezvous protocol.
[0054] In the software implementation of the eager protocol, data is delivered to a system buffer, from which the data must be copied to a user buffer. While this approach reduces synchronization, it is expensive in terms of memory capacity and memory bandwidth. In some embodiments, the NIC may provide a mechanism to write eager messages directly to the user buffer if the destination address can be determined quickly.
[0055] Specifically, when the LPE receives the first request packet (containing the MPI message envelope), it searches the physical endpoint's priority list for a matching buffer. Matching can be performed based on the source, a set of match bits contained in the message, and buffer-specific match and ignore bits. The match list entry contains information such as the starting address, length, translation context, and various attributes into which the PUT data (i.e., the eager message) will be written, allowing the direct memory access (DMA) dispatch logic to write data directly to the user buffer. If no match is found in the priority list, the LPE searches the overflow list for a description of the memory parameters into which it can write the PUT data and adds a list entry with the message description to the unexpected list.
[0056] In the software implementation of the Rendezvous protocol, the transfer of bulk data is delayed until the destination address is known. While this approach reduces system memory usage, it requires software intervention to ensure progress. In some embodiments, the Rendezvous protocol is offloaded to the network card, enabling strong progression.
[0057] More specifically, when transmitting large MPI messages, the initiator can send a small initial message containing the MPI envelope used for matching and a modest amount of eager data. After the matching operation is complete, the target executes a GET to transfer the bulk data to the user buffer. This can improve network performance because bulk data is delivered as unordered GET responses, which the network can forward adaptively on a packet-by-packet basis.
[0058] When a rendezvous request lands on the unexpected list, the GET is initiated when the user process sends the matching append. The start of the rendezvous GET is the same in both cases; it is triggered by the completion of the match on a rendezvous PUT request.
[0059] This is a valuable relief. MPI applications are expected to send non-blocking receive data early and then return to computation. Moving the rendezvous to the NIC ensures good overlap between computation and communication. The NIC performs the reconciliation and instantiates the bulk data movement asynchronously, achieving strong progression.
[0060] Fig. shows a flowchart for performing list matching in an NIC. During operation, the NIC may receive a match request (operation 502). The match request may be a command from the CQ to manipulate the lists or update the physical endpoint state, or a message match request from the IXE. The match request may be queued to an appropriate MRQ, depending on its type (operation 504). An arbitrator selects an MRQ to dequeue a match request (operation 506) and sends the dequeued match request to a lookup table, also known as a processing engine map, to determine a processing engine to process the match request (operation 508). The determination may be based on the physical portal index (i.e., the identification of the physical endpoint).
[0061] The match request is also sent to the PLEC (operation 510), which attempts to find a match (operation 512). If a match is found in the PLEC, the PLEC outputs the matching entry (operation 514). Otherwise, the match request is sent to a PE / TC MRQ (operation 516). An arbitrator selects a PE / TC MRQ to exit the queue (operation 518). In some embodiments, arbitration may occur in two steps. In the first step, a ready processing engine is selected using round robin. In the second step, a ready TC within that processing engine may be selected using weighted round robin arbitration, where each TC has a predetermined weighting factor.
[0062] The request taken from the PE / TC MRQ is sent to the corresponding processing engine, which in turn searches for the matching list entry (operation 520). The matching operations of the processing system are described in Fig. shown. Exemplary computer system
[0063] Fig. shows an exemplary computer system equipped with a NIC that facilitates MPI list matching. Computer system 650 includes a processor 652, a memory device 654, and a storage device 656. Storage device 654 may include a volatile memory device (e.g., a dual inline memory module (DIMM)). Computer system 650 may also be coupled to a keyboard 662, a pointing device 664, and a display device 666. Storage device 656 may store an operating system 670. An application 672 may operate with operating system 670.
[0064] Computer system 650 may be equipped with a host interface coupling a NIC 620, which enables efficient management of data requests. NIC 620 may provide one or more HNIs to computer system 650. NIC 620 may be coupled to a switch 602 via one of the HNIs. NIC 620 may include a list processing logic block 630, as described in connection with Fig. The list processing logic block 630 may include an MRQ logic block 632 that stores match requests to be processed, a PLEC logic block 634 that enables fast lookup, and a processing engine 636 for matching the incoming match request with a list entry stored in the memory bank.
[0065] In summary, the present disclosure describes a NIC that facilitates MPI list matching. The NIC may include a host interface, a network interface, and a hardware LPE. The host interface may connect the NIC to a host device. The network interface may connect the NIC to a network. The hardware LPE may perform high-speed list matching. A high degree of parallelism may be achieved by implementing multiple processing units (PEs) and multiple memory banks within a processing unit. Because each processing engine or TC is assigned its own queue, the system prevents one processing engine or TC from blocking the queues of other processing engines or TCs. In the hardware list processing engine, the match pipeline stage and the match termination condition overlap to reduce latency.The NIC also allows offloading of MPI message processing, including eager and rendezvous messages.
[0066] The methods and processes described above may be performed by hardware logic blocks, modules, logic blocks, or devices. The hardware logic blocks, modules, or devices may include, but are not limited to, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), dedicated or shared processors that execute code at a specific time, and other known or later developed programmable logic devices. When activated, the hardware logic blocks, modules, or devices execute the methods and processes contained therein.
[0067] The methods and processes described herein may also be embodied as code or data that can be stored in a memory device or computer-readable storage medium. When a processor reads and executes the stored code or data, the processor can perform these methods and processes.
[0068] The foregoing descriptions of embodiments of the present invention have been presented for purposes of illustration and description only. They are not intended to be exhaustive or to limit the present invention to the forms shown. Accordingly, many modifications and variations will be apparent to those skilled in the art. Furthermore, the present invention is not intended to be limited by the above disclosure. The scope of the present invention is defined by the appended claims.
Claims
[1] Network interface controller NIC (202, 204, 620), which includes: a host interface (210, 211) for coupling a host device; a network interface (220, 221) for coupling a network; and a hardware list processing engine LPE (264) for: Receiving a reconciliation request and Performing a Message Passing Interface (MPI) list reconciliation based on the received reconciliation request, wherein the LPE comprises a persistent list entry cache PLEC (412) to store previously matched list entries, and wherein performing the MPI list match comprises bypassing a processing pipeline comprising a lookup table (408) and a number of match request queues (410) in response to finding a matching entry in the PLEC for the received match request. [2] The network interface controller of claim 1, wherein the reconciliation request comprises: a reconciliation request corresponding to a command received via the host interface, or a reconciliation request corresponding to a message received over the network interface. [3] The network interface controller of claim 2, wherein the LPE further serves to: maintain a first set of reconciliation request queues (402) for reconciliation requests corresponding to received commands; and maintain a second set of reconciliation request queues (404) for reconciliation requests corresponding to received messages; wherein a number of queues in the first or second set of reconciliation request queues corresponds to a number of physical endpoints supported by the network interface controller. [4] The network interface controller of claim 2, wherein the message is an MPI message. [5] The network interface controller of claim 4, wherein the message is based on an eager protocol or a rendezvous protocol associated with MPI. [6] The network interface controller of claim 1, wherein the LPE further comprises a plurality of processing elements (300); and wherein a respective processing element comprises a plurality of matching engines (302, 304) and a plurality of memory banks (306, 308, 310, 312) storing one or more lists, the memory banks being connected to the matching engines using a crossbar. [7] The network interface controller of claim 6, wherein a respective matching engine comprises a unified search pipeline for searching the one or more lists, and wherein the one or more lists comprise a priority list and an unexpected list. [8] The network interface controller of claim 6, wherein a respective matching engine comprises a single pipeline stage to perform in parallel a matching operation for a previous matching request and a calculation to determine a current read or write address. [9] The network interface controller of claim 1, wherein the LPE is further operable to perform atomic search operations on a plurality of lists. [10] Procedure comprising: Receiving a reconciliation request by a network interface controller NIC (202, 204, 620), the NIC comprising a host interface (210, 211) for coupling a host device and a network interface (220, 211) for coupling a network; and Performing an MPI list reconciliation by a hardware list processing engine LPE (264) in the NIC based on the received reconciliation request, wherein the LPE comprises a persistent list entry cache PLEC (412) to store previously matched list entries, and wherein performing the MPI list match comprises bypassing a processing pipeline comprising a lookup table (408) and a number of match request queues (410) in response to finding a matching entry in the PLEC for the received match request. [11] The method of claim 10, wherein the reconciliation request comprises: a reconciliation request corresponding to a command received via the host interface, or a reconciliation request corresponding to a message received over the network interface. [12] The method of claim 11, further comprising: queuing reconciliation requests corresponding to received commands into a first set of reconciliation request queues (402) by the LPE; and queuing, by the LPE, reconciliation requests corresponding to received messages into a second set of reconciliation request queues (404); wherein a number of queues in the first or second set of reconciliation request queues corresponds to a number of physical endpoints supported by the NIC. [13] The method of claim 11, wherein the message is an MPI message. [14] The method of claim 13, wherein the message is based on an eager protocol or a rendezvous protocol coupled with MPI. [15] The method of claim 10, wherein performing the MPI list matching comprises: Selecting a processing element (300) from a plurality of processing elements within the hardware list processing engine to process the request; and Selecting a matching engine from a plurality of matching engines (302, 304) within a respective processing element to perform a matching operation, wherein a plurality of memory banks (306, 308, 310, 312) storing one or more lists are connected to the plurality of matching engines using a crossbar. [16] The method of claim 15, wherein a respective matching engine comprises a unified search pipeline for searching the one or more lists, and wherein the one or more lists comprise a priority list and an unexpected list. [17] The method of claim 15, wherein a respective matching engine performs the matching operation for a previous matching request in parallel with a calculation to determine a current read or write address. [18] The method of claim 10, wherein performing the MPI list matching comprises performing atomic search operations on a plurality of lists.
Citation Information
Patent Citations
Method and apparatus for managing applicaiton state in a network interface controller in a high performance computing system
US20170054633A1
Communication of a large message using multiple network interface controllers
US20190044875A1