Programmable Computer IO Device Interface
The programmable IO device interface addresses the inflexibility of existing network interfaces by enabling customizable host interfaces and hardware offloading, enhancing performance and efficiency in data communication.
Patent Information
- Application Number
- CN201980027381.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-22
- Filing Date
- 2019-02-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-11-13
AI Technical Summary
The prior art is difficult to provide flexible and efficient IO device interfaces, and cannot adapt to the rapid development and changes of network device structure and function sets, various protocols, operating systems and applications, resulting in insufficient performance and waste of resources.
The programmable IO device interface mechanism is adopted to achieve efficient interaction between the IO device and the host system through programmable device registers, memory-based data structures and DMA block pipeline design, and support flexible hardware shunt of custom host interfaces and network functions, improve performance and operate within the power budget.
It realizes low-latency interaction between IO devices and host systems, improves performance and reduces resource consumption, supports flexible network pipelines and custom functions, and adapts to rapidly changing network environments.
Smart Images

Figure CN112074808B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 634,090, filed Feb. 22, 2018, the disclosure of which is incorporated herein by reference. BACKGROUND OF THE INVENTION
[0003] A computing environment can include a host (e.g., a server), a computer (e.g., a virtual machine or a container) running one or more processes. The host and / or the processes can be configured to communicate with other processes or devices via a computing network. The host system interfaces with the computing network via an input / output (I / O) device (e.g., a network interface card (NIC)).
[0004] A computer system interfaces with an I / O device via a set of designated device registers and memory-based data structures. These registers and data structures are typically fixed for a given I / O device, allowing a specific device driver to run on the computer system and control the I / O device. In a data communication network, network interfaces are typically fixed-defined control structures, descriptors, registers, etc. Network data and control structures are memory-based and access storage using direct memory access (DMA) semantics. Network systems such as switches, routing devices receive messages or data packets on one of a set of input interfaces and forward them to one or more of a set of output interfaces. Users typically require such routing devices to operate as fast as possible to keep up with the high rate of incoming messages. One challenge associated with network systems involves providing flexible network interfaces to accommodate changes in network device architectures and function sets, various protocols, operating systems, applications, and device models in rapid evolution. SUMMARY OF THE INVENTION
[0005] There is a desire to provide a flexible and fully programmable I / O device interface mechanism so that an I / O device can be customized to be more suitable for a desired application or OS interface. There is a need to provide a programmable I / O device interface to work with a highly configurable network pipeline, a customizable host interface, and flexible hardware offload for storage, security, and network functions to improve performance and stay within a target power budget. The present invention meets this need and also provides related advantages.
[0006] The subject matter disclosed herein meets this need by providing a device interface that can be programmed in the form of device data structures and control registers and defines device behavior to coordinate with the device interface. The programmable I / O device interface may be able to work with a highly configurable network pipeline, a customizable host interface, and flexible hardware offloading for storage, security, and network functions, thereby improving performance and staying within the target power budget. An I / O device with the provided device interface may improve performance. The provided device interface mechanism may allow the I / O device interface to emulate existing host software drivers and effectively interact with a variety of different software drivers.
[0007] The performance of an I / O device can be improved by replacing the conventional fixed-function direct memory access (DMA) engine, control registers, and device state machine with a programmable pipeline of match, action, and DMA stages. For example, a stage in the pipeline can initiate DMA read and write operations to the host system, obtain memory-based descriptors, scatter-gather lists (SGLs), or custom data structures that describe I / O operations. The provided interface mechanism can include using a stack of fields mapped to data structures to describe host computer data structures (e.g., descriptors for describing how to make packets, different types of packets); and storing the internal DMA engine state in a programmable match table that can be updated by the hardware pipeline (e.g., a match processing unit (MPU)) as well as the host processor; defining device registers through separate programmable fields and supporting them by a hardware mechanism through an address remapping mechanism. The above interface mechanism enables the I / O device to directly interact with host data structures without the help of the host system, thereby achieving lower latency and deeper processing in the I / O device.
[0008] The I / O device interface can be a highly optimized ring-based I / O queue interface with an efficient software programming model to provide high performance through CPU and Peripheral Component Interconnect (PCIe) bus efficiency. The I / O device can be connected to the processor of the host system via the PCIe bus. The I / O device can be docked to the host system through one or more (e.g., one to eight) physical PCIe interfaces.
[0009] An I / O device can break down packet processing tasks into a series of table lookups or matches, along with processing actions. A Match Processing Unit (MPU) can be provided to perform table-based actions at each stage of the network pipeline. One or more MPUs can be used in conjunction with a table engine that is configured to extract a set of programmable fields and obtain table results. Once the table engine has completed obtaining the lookup results, the table engine can pass the table results and associated data packet header fields to the MPU for processing. The MPU can run a target program based on a domain-specific instruction set. The MPU can take the table lookup results and the data packet header as inputs and generate table update and data packet header rewrite operations as outputs. A predetermined number of such table engine and MPU pipeline stages can be combined to form a programmable pipeline capable of operating at a high packet processing rate. This can prevent the MPU from encountering data loss stalls and allow the MPU program to execute within a determined time, and then pipeline them together to maintain the target data packet processing rate. In some cases, a programmer or compiler may break down a data packet processing program into a set of related or independent table lookup and action processing stages (match + action), which are respectively mapped to the table engine and MPU stages. In some cases, if the required number of stages exceeds the implemented number of stages, the data packet can be redistributed for other processing.
[0010] Thus, in one aspect, a method for a programmable I / O device interface is disclosed herein. The method includes: providing a programmable device register, a memory-based data structure, a DMA block, and a pipeline of processing entities, and the pipeline of processing entities is configured to: (a) receive a packet including a header portion and a payload portion, where the header portion is used to generate a packet header vector; (b) generate table results by performing a packet matching operation using a table engine, where the table results are generated at least in part based on data stored in a programmable match table and the packet header vector; (c) receive, at a match processing unit, an address of a set of instructions associated with the programmable match table and the table results; (d) perform, by the match processing unit, one or more actions according to the loaded set of instructions until the instructions are completed, where the one or more actions include updating the memory-based data structure, inserting a DMA command, and / or initiating an event; (e) perform, by the DMA block, a DMA operation according to the inserted DMA command.
[0011] In some embodiments, the method further includes: providing the header portion to a subsequent circuit, where the subsequent circuit is configured to assemble the changed header portion into the corresponding payload portion.
[0012] In some embodiments, the programmable match table includes a DMA register table, a descriptor format, or a control register format. In some cases, the programmable match table is selected based on packet type information related to the packet type associated with the header portion. In some cases, the programmable match table is selected based on the ID of the match table selected in a previous stage.
[0013] In some embodiments, the table result includes a key related to the programmable match table and a match result of the match operation. In some embodiments, the memory unit of the match processing unit is configured to store multiple sets of instructions. In some cases, the multiple sets of instructions are associated with different actions. In some cases, each set of instructions in the multiple sets of instructions is stored in a contiguous region of the memory unit, and the contiguous region is identified by the address of each set of instructions.
[0014] In some embodiments, one or more actions further include updating the programmable match table. In some embodiments, the method further includes: locking the match table for exclusive access by the match processing unit while the match processing unit processes the match table. In some embodiments, the packets are processed in a non-stalling manner.
[0015] In a related but independent aspect, a device having a programmable IO device interface is provided. The device includes: (a) a first memory unit storing a plurality of programs thereon, where the plurality of programs are associated with a plurality of actions, the plurality of actions including updating a memory-based data structure, inserting a DMA command, or initiating an event; (b) a second memory unit for receiving and storing table results, where the table results are provided by a table engine configured to perform a packet matching operation on a packet header vector included in a header portion and data stored in a programmable match table; (c) circuitry for executing a program selected from the plurality of programs in response to the table results and an address received by the device, where the program is executed until completion and the program is associated with the programmable match table.
[0016] In some embodiments, the device is configured to provide the header portion to a subsequent circuit. In some cases, the subsequent circuit is configured to assemble the modified header portion into a corresponding payload portion.
[0017] In some embodiments, the programmable match table includes a DMA register table, a descriptor format, or a control register format. In some cases, the programmable match table is selected based on packet type information related to the packet type associated with the header portion. In some cases, the programmable match table is selected based on the ID of the match table selected in a previous stage.
[0018] In some embodiments, each of the plurality of programs includes a set of instructions stored in a contiguous region of the first memory unit, and the contiguous region is identified by the address. In some embodiments, the one or more actions include updating the programmable match table. In some embodiments, the event is not related to changing the header portion of the packet. In some embodiments, the memory-based data structure includes at least one of the following: an administrative token for initiating an event, an administrative command, a processing token.
[0019] A system including a plurality of devices, wherein the plurality of devices are coordinated to execute the set of instructions or the one or more actions concurrently or sequentially according to a configuration. In some embodiments, the configuration is determined by application instructions received from a main memory of a host device operatively coupled to the plurality of devices. In some embodiments, the plurality of devices are arranged to process the packet according to a pipeline of stages.
[0020] It should be understood that the different aspects of the present invention can be understood individually, jointly, or in combination with each other. The various aspects of the invention described herein can be applied to any particular application presented below or any other type of the data processing system disclosed herein. Any description of the data processing herein can be applied to and used for any other data processing scenario. In addition, any embodiment disclosed in the case of the data processing system or device is also applicable to the method disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The patent or application document contains at least one color drawing. The Patent Office will, upon request and after providing the necessary fees, provide a copy of the present patent or patent application publication with the color drawing. The novel features of the present invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and the drawings which illustrate the illustrative embodiments in which the principles of the invention are utilized, in the drawings:
[0022] Figure 1 A block diagram showing an exemplary computing system architecture according to an embodiment of the present invention;
[0023] Figure 2 An exemplary configuration of a plurality of MPUs for executing a program is shown;
[0024] Figure 3A is a diagram showing an example of the internal structure of a PCIe configuration register;
[0025] Figure 3B An example of a P4-defined descriptor ring is shown;
[0026] Figure 4Shows a block diagram of a matching processing unit (MPU) according to an embodiment of the present invention;
[0027] Figure 5 Shows a block diagram of an exemplary P4 ingress or egress pipeline (PIP pipeline) according to an embodiment of the present invention;
[0028] Figure 6 Illustrates an exemplary stage expansion pipeline for Ethernet packet transmission (i.e., Tx P4 pipeline);
[0029] Figure 7 Shows a block diagram of an exemplary Rx P4 pipeline according to an embodiment of the present invention;
[0030] Figure 8 Shows a block diagram of an exemplary Tx P4 pipeline according to an embodiment of the present invention; and
[0031] Figure 9 Shows an example of an extended transmission pipeline (i.e., TxDMA pipeline). Detailed Description
[0032] In certain embodiments described herein, network devices, systems, and methods for processing data (such as packets or tables) are disclosed, which have reduced data stalls.
[0033] Certain definitions
[0034] Unless otherwise defined, all technical terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains.
[0035] As used herein, the singular forms "a", "an", etc. include plural referents unless the context clearly dictates otherwise. Any reference to "or" herein is intended to include "and / or" unless otherwise stated.
[0036] References in this specification to "some embodiments" or "embodiments" mean that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in some embodiments" or "in embodiments" throughout this specification are not necessarily all referring to the same embodiment. Moreover, the specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0037] As used herein, the terms "component", "system", "interface", "unit", "block", "device", etc. refer to computer-related entities, hardware, software (e.g., in execution), and / or firmware. For example, a component can be a processor, a process running on a processor, an object, an executable program, a program, a storage device, and / or a computer. By way of illustration, a server and an application running on the server can be components. One or more components can reside within a process, and a component can be localized on one computer and / or distributed between two or more computers.
[0038] In addition, these components can execute from various computer-readable media having various data structures stored thereon. Components can communicate through local and / or remote processes, such as in accordance with signals having one or more data packets (e.g., data from one component interacts with another component in a local system, a distributed system, and / or across a network (e.g., the Internet, a local area network, a wide area network, etc.) via a signal).
[0039] As another example, a component can be a device having a specific function provided by a mechanical part operated by an electrical or electronic circuit; the electrical or electronic circuit can be operated by a software application or a firmware application executed by one or more processors; the one or more processors can be internal or external to the device and can execute at least a part of the software or firmware application. As another example, a component can be a device providing a specific function through an electronic component without the need for a mechanical part; the electronic component can include one or more processors for executing at least part of the software and / or firmware that gives the electronic component its function.
[0040] Furthermore, as used herein, the term "exemplary" means serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or superior to other aspects or designs. Instead, the use of the term "exemplary" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise stated or clear from the context, "X uses A or B" is intended to mean any natural inclusive permutation. That is, if X uses A; X uses B; or X uses both A and B simultaneously, then "X uses A or B" is satisfied in any of the above instances. Additionally, the articles "a" and "an" as used in this application and the appended claims should generally be construed to mean "one or more" unless otherwise stated or clearly indicated to the contrary from the context.
[0041] Embodiments of the present invention can be used in various applications. Some embodiments of the present invention can be used in combination with various devices and systems, such as, personal computers (PCs), desktop computers, mobile computers, laptop computers, notebook computers, tablet computers, server computers, handheld computers, handheld devices, personal digital assistant (PDA) devices, handheld PDA devices, wireless communication stations, wireless communication devices, wireless access points (APs), modems, networks, wireless networks, local area networks (LANs), wireless LANs (WLANs), metropolitan area networks (MANs), wireless MANs (WMANs), wide area networks (WANs), wireless WANs (WWANs), personal area networks (PANs), wireless PANs (WPANs), devices and / or networks operating according to existing IEEE 802.11, 802.11a, 802.11b, 802.11e, 802.11g, 802.11h, 802.11i, 802.11n, 802.16, 802.16d, 802.16 standards and / or future versions and / or derivatives and / or long term evolution (LTE) of the above standards, units and / or devices belonging to a part of the above networks, unidirectional and / or bi-directional wireless communication systems, cellular wireless telephone communication systems, cellular telephones, wireless telephones, personal communication system (PCS) devices, PDA devices including wireless communication devices, multiple input multiple output (MIMO) transceivers or devices, single input multiple output (SIMO) transceivers or devices, multiple input single output (MISO) transceivers or devices, etc.
[0042] The term "table" refers to various types of tables related to data or packet processing. For example, the table can be a match table used in the match + action phase, such as a forwarding table (e.g., a hash table for Ethernet address lookup, a longest prefix match table for IPv4 or IPv6, a wildcard lookup for access control lists (ACLs)). These tables can be stored in various memory locations, such as internal static random access memory (SRAM), network interface card (NIC) DRAM, or host memory.
[0043] The term "match+action" refers to a paradigm for network packet switching (such as those implemented by OpenFlow switches or P4 pipelines, which use match tables, action tables, statistics memories, meter memories, state memories, and ternary indirect memories). The term "P4" refers to a high-level language for writing protocol-independent packet processors. P4 is a declarative language for expressing how the pipelines of network forwarding elements (such as switches, network interface cards, routers, or network function devices) process packets. It is based on an abstract forwarding model that consists of a parser and a set of match+action table resources divided between ingress and egress. The parser identifies the headers present in each incoming packet. Each match+action table performs a lookup on a subset of the header fields and applies the operation corresponding to the first match in each table.
[0044] Although, for illustrative purposes, some portions of the present disclosure relate to wired and / or wireline communication systems or methods, embodiments of the present invention are not limited thereto. For example, one or more wired communication systems may utilize one or more wireless communication components, one or more wireless communication methods or protocols, and the like.
[0045] Although, for illustrative purposes, some portions discussed herein may relate to fast or high-speed interconnect infrastructures, fast or high-speed interconnect components or adapters with OS bypass functionality, high-speed interconnect cards or network interface cards (NICs) with OS bypass functionality, or fast or high-speed interconnect infrastructures or fabrics, embodiments of the present invention are not limited thereto and may be used in combination with other infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs, which may or may not be fast or high-speed or have an operating system bypass functionality. For example, some embodiments of the present invention may be used in combination with InfiniBand (IB) infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs; with Ethernet infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs; with Gigabit Ethernet (GEth) infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs; with infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs with an operating system; with an operating system of an infrastructure, fabric, component, adapter, host channel adapter, card, or NIC that allows user-mode applications to directly access this hardware and bypass calls to the operating system (i.e., has an OS bypass functionality); with infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs; with connectionless and / or stateless infrastructures, fabrics, components, adapters, host channel adapters, cards, or NICs; and / or with other suitable hardware.
[0046] Computer systems employ a variety of peripheral components or I / O devices. An example of a computer system host processor connected to an I / O device via a component bus defined by the Peripheral Component Interconnect Express (PCIe), a high-speed serial computer expansion bus standard. A device driver (also known as a driver) is hardware-specific software that can control the operation of a hardware device connected to a computing system.
[0047] In computing, virtualization technologies are used to allow multiple operating systems to share processor resources simultaneously. One such virtualization technology is Single Root I / O Virtualization (SR-IOV), which is described in the PCI-SIG Single Root I / O Virtualization and Sharing Specification. A physical I / O device can allow multiple virtual machines to use the device simultaneously via SR-IOV. In SR-IOV, a physical device may have a Physical Function (PF) that allows input / output operations and device configuration, and one or more Virtual Functions (VFs) that allow data input / output. According to SR-IOV, a Peripheral Component Interconnect Express (PCIe) device appears as multiple separate physical PCIe devices. For example, an SR-IOV network interface card (NIC) with a single port can have up to 256 virtual functions, each representing its own NIC port.
[0048] In one aspect, a programmable device interface is provided. The device interface can be a highly optimized ring-based I / O queue interface with an efficient software programming model to provide high performance with CPU and PCIe bus efficiency. Figure 1 A block diagram of an exemplary computing system architecture 100 in accordance with an embodiment of the present invention is shown. A hypervisor 121 on a host computing system 120 can interact with a physical I / O device 110 using a PF 115 and one or more VFs 113. As shown, the computing system 110 can include a management device 117 configured to manage interface devices. The management device 117 can communicate with a processing entity 111 (e.g., an ARM core) and a management entity 119 (e.g., a management virtual machine system). It should be noted that the computing system shown is only an example mechanism and does not imply any limitation on the scope of the present invention. The programmable I / O interface and method provided can be applied to any operating system-level virtualization (e.g., containers and Docker systems) or machine-level virtualization or computing systems without virtualization features.
[0049] The hypervisor 121 typically provides the host with operating system functions (e.g., process creation and control, file system process threads, etc.) as well as CPU scheduling and memory management. In some cases, the host computing system 120 may include programs that implement a machine emulator and a virtualizer. The machine emulator and virtualizer can assist in virtualizing various computer I / O devices in the virtual machine, such as virtualized hard disks, optical disk drives, and NICs. Virtio is a virtualization standard for implementing virtual I / O devices in a virtual machine and can be regarded as an abstraction of a set of general-purpose emulated devices in the hypervisor.
[0050] When using a device emulator, the provided programmable I / O device interface mechanism allows native hardware speeds. The programmable I / O device interface allows the host system to interface with the I / O device using existing device drivers without reconfiguration or change. In some cases, VF devices, PF devices, and management devices may have similar driver interfaces so that such devices can be supported by a single driver. In some cases, such devices may be referred to as Ethernet devices.
[0051] The I / O device 110 can provide various services and / or functions to the operating system operating as a host on the computing system 110. For example, the I / O device can provide network connection functions, coprocessor functions (e.g., graphics processing, encryption / decryption, database processing, etc.) to the computing system. The I / O device 110 can interface with other components in the computing system 100 via, for example, a PCIe bus.
[0052] As described above, the SR-IOV specification allows a single root function (e.g., a single Ethernet port) to appear as multiple physical devices in a virtual machine. A physical I / O device with SR-IOV capabilities can be configured to appear as multiple functions in the PCI configuration space. The SR-IOV specification supports physical functions and virtual functions.
[0053] A physical function is a complete PCIe device that can be discovered, managed, and configured as a normal PCI device. The physical function configures and manages SR-IOV functionality by allocating virtual functions. An I / O device can expose one or more physical functions (PFs) 115 to a host computing system 120 or a hypervisor 121. A PF 115 can be a full-featured PCIe device that includes all configuration resources and capabilities for the I / O device. In some cases, a PF can be a PCIe function that includes SR-IOV extension capabilities, which assist in the configuration or management of the I / O device. A PF device is essentially the basic controller of an Ethernet device. A PF device can be configured with up to 256 VFs. In some cases, a PF may include extended operations such as allocating, configuring, and releasing VFs, discovering the hardware capabilities of VFs such as receive side scaling (RSS), discovering the hardware resources of VFs such as queue and interrupt number resources, configuring the hardware resources and characteristics of VFs, saving and restoring the hardware state, etc. In some cases, a PF device can be configured as a boot device that can present an Option ROM base address register (BAR).
[0054] The I / O device can also provide one or more virtual functions (VFs) 113. A VF can be a lightweight PCIe function that contains the resources necessary for data movement but can have a minimized set of configuration resources. In some cases, a VF may include a lightweight PCIe function that supports SR-IOV. To use an SR-IOV device in a virtual system, the hardware can be configured to create multiple VFs. These VFs can be made available to the hypervisor for allocation to virtual machines. The VFs can be manipulated (e.g., created, configured, monitored, or destroyed) by, for example, an SR-IOV physical function device. In some cases, each of the multiple VFs is configured with one or more base address registers (BARs) to map NIC resources to the host system. A VF can map one or more logical interfaces (LIFs) or ports that are used in the I / O device for forwarding and transaction identification. A LIF may belong to only one VF. Within a physical device, all virtual functions may have the same BAR resource layout and are stacked sequentially in the host PCIe address space. The I / O device PCIe interface logic can be programmed to map control registers and NIC memory regions with programmable access permissions (e.g., read, write, execute) to the VF BAR.
[0055] The IO device may include a management device 117 for managing the IO device. The management device 117 may not have direct access to the network uplink port. The management device may communicate with the processing entity 111. For example, traffic on the management device may be directed to an internal receive queue for processing by the management software on the processing entity 111. In some cases, the management device may be used to pass to the management entity 119 through a hypervisor, such as managing virtual machines. For example, a different device ID may be assigned to the management device 117 from the PF device 115, so that when the PF device does not require the management device, the device driver in the hypervisor can be released for the PF device.
[0056] Figure 2 Another exemplary IO device system 200 with the described programmable device interface according to some embodiments of the present invention is shown. The system 200 serves as an example for implementing P4 and extended P4 pipelines and various other functions to provide improved network performance. In some cases, the device interface may improve network performance in the following ways: not requiring reading PCIe bus registers in the packet sending or receiving path; providing a single posted (non-blocking) PCIe bus register write operation for packet transmission; supporting Message-Signaled Interrupts (MSI) and Message-Signaled Interrupt Extended (MSI-X) modes, with driver-configurable interrupt moderation for high-performance interrupt handling; supporting I / O queues with outstanding requests per queue (e.g., up to 64k); transmitting TCP Segmentation Offload (TSO) with improved send sizes; providing Transmission Control Protocol (TCP) / User Datagram Protocol (UDP) checksum offload; supporting a variable number of receive queues to support industry-standard Receive-Side Scaling (RSS); supporting SR-IOV with up to 255 virtual functions.
[0057] The IO device system 200 may be the same IO device as described in Figure 1 and is implemented as a rack-mounted device, and includes one or more Application-Specific Integrated Circuits (ASICs) and / or boards with components mounted thereon. As Figure 2 shown, the system 200 may include four Advanced RISC Machines (ARM) processors with coherent L1 and L2 caches, a shared local memory system, flash non-volatile memory, a DMA engine, and other IO devices for operation and debugging. The ARM processors may observe and control all NIC resources through address mapping. As described later herein, the ARM processors may implement the P4 pipeline and the extended P4 pipeline.
[0058] The system may include a host interface and a network interface. The host interface may be configured to provide a communication link with one or more hosts (e.g., host servers). The host interface block may also observe regions of the address space through PCIe BAR mapping to expose NIC functionality to the host system. In an example, the address mapping may be initially created according to the principle of the ARM-limited ARM memory mapping, which provides SOC addressing guidelines for a 34-bit memory mapping.
[0059] The network interface may support a network connection or uplink to a compute network, which may be, for example, a local area network, a wide area network, and various other networks described elsewhere herein. The physical link may be controlled by a management agent (e.g., management entity 119) through a device driver. For example, the physical link may be configured via a "virtual link" associated with a device logical interface (LIF).
[0060] All memory transactions in system 200, including host memory, high bandwidth memory (HBM), and registers, may be connected based on IP from an external system via an on-chip coherent network (NOC). The NOC may provide a cache-coherent interconnect between NOC masters (including P4 pipelines, extended P4 pipelines, DMA, PCIe, and ARM). The interconnect may use a programmable hash algorithm to distribute HBM memory transactions across multiple (e.g., 16) HBM interfaces. All traffic to the HBM may be stored in the NOC cache (e.g., 1MB cache). The NOC cache may be coherent with the ARM cache. Since HBM is not efficient when processing small writes, the NOC cache may be used to aggregate HBM write transactions that may be smaller than a cache line (e.g., 64-byte size). The NOC cache may have high bandwidth because it is ahead of the 1.6Tb / s HBM and thus supports operations up to 3.2Tb / s.
[0061] The system may include an internal HBM memory system for running Linux, storing large data structures (e.g., flow tables and other analytics), and providing buffer resources for advanced functions (including TCP termination and proxy, deep packet inspection, storage offloading, and connected FPGA functionality). The memory system may include an HBM module, which may support a 4GB capacity or an 8GB capacity depending on the package and HBM.
[0062] As described above, the system may include a PCIe host interface. The PCIe host interface may support a bandwidth of, for example, 100 Gb / s per PCIe connection (e.g., dual PCIe Gen4x8 or single PCIe Gen3x16). A mechanism or scheme for mapping the resources available on an IO device to a memory-mapped control region associated with a virtual IO device may be implemented by using a pool of configurable PCIe base address registers (BARs) coupled with a resource mapping table. The IO resources provided by the IO device may be mapped to host addresses within the PCIe standard framework so that the same device drivers used to communicate with a physical PCIe device may be used to communicate with the corresponding virtual PCIe device.
[0063] The IO device interface may include programmable registers. These registers may include, for example, PCIe base address registers (BARs), which may include a first memory BAR containing device resources (e.g., device command registers, doorbell registers, interrupt control registers, interrupt status registers, MSI-X interrupt table, MSI-X interrupt pending bit array, etc.), a second BAR containing a device doorbell page, and a third BAR mapping a controller memory buffer.
[0064] The device command registers are a set of registers for submitting management commands to hardware or firmware. For example, the device command registers may specify a single 64-byte command and a single 16-byte completion response. The register interface may allow a single outstanding command at a time. The device command doorbell is a dedicated doorbell for indicating that a command in the device command registers is ready.
[0065] The second BAR may contain a doorbell page. The general form of the second BAR may contain multiple LIFs, each LIF with multiple doorbell pages. A network device (i.e., an IO device) may have at least one LIF with at least one doorbell page. Any combination of single / multiple LIFs with single / multiple doorbell pages is possible, and the driver may be prepared to recognize and operate on different combinations. In an example, the doorbell pages may be presented in 4k strides by default to match common system page sizes. The stride between doorbell pages may be adjusted in the virtual function device 113 to match the system page size configuration setting in the SR-IOV capability header in the parent physical function device 115. This page size partitioning allows each process to have protected independent direct access to a set of doorbell registers by allowing each process to map and access a doorbell page dedicated to its use. Each page may provide the doorbell resources required to operate the data path queues of a LIF while protecting access to these resources from another process.
[0066] The doorbell register can be written by software to adjust the producer index of the queue. Adjusting the producer index is the mechanism for transferring the ownership of queue entries in the queue descriptor ring to the hardware. Some doorbell types, such as management queues, Ethernet transmit queues, and RDMA transmit queues, may cause the hardware queue arrangement to further process the available descriptors in the queue. After updating the producer index, other queue types (e.g., completion queues and receive queues) may not require further action by the hardware queue.
[0067] The interrupt status register can contain bits for each interrupt resource of the device. The register can set a bit indicating that the corresponding interrupt resource has asserted its interrupt. For example, bit 0 in the interrupt status indicates that interrupt resource 0 has been asserted, and bit 1 indicates that interrupt resource 1 has been asserted.
[0068] The controller memory buffer can be a region of general-purpose memory resident on the IO device. A user or kernel driver can map into this controller memory BAR and establish descriptor rings, descriptors, and / or payload data in this region. A bit can be added to the descriptor to select whether to interpret the descriptor address field as a host memory address or an offset relative to the start of the device controller memory window. If the extended P4 program is a host address, a specified bit of the address (e.g., bit 63) can be set, or the bit can be cleared and the device controller memory base address added to the offset when building a TxDMA operation for the DMA level.
[0069] The MSI-X resources can be mapped through the first BAR, and the format can be described by the PCIe base specification. The MSI-X interrupt table is a region of control registers that allows the operating system to program the MSI-X interrupt vectors on behalf of the driver.
[0070] The MSI-X interrupt pending bit array (PBA) is an array of bits for each MSI-X interrupt supported by the device.
[0071] The IO device interface can support a programmable DMA register table, descriptor format, and control register format, allowing for a dedicated VF interface and user-defined behavior. The IO device PCIe interface logic can be programmed to map control registers and NIC memory regions with programmable access permissions (e.g., read, write, execute) to the VF BAR.
[0072] Figure 3AIt is a diagram showing an example 300 of the internal arrangement of PCIe configuration registers. As an example of an address, the Device ID specifies a vendor-specific device number, the Vendor ID specifies the manufacturer's number (both at offset 00h), and the Class Code (offset 08h) specifies the device attributes. Addresses with offsets 10h - 24h and 30h are used for base address registers. The configuration software contained in SR-PCIM can identify a device by looking up register values. When allocating an address space for an I / O device, the configuration software in SR-PCIM uses the base address register to write the base address. Device identification and related processes occur during the PCIe configuration cycle. Such a configuration cycle occurs during system startup and possibly after a hot-plug operation.
[0073] The transmit ring may include a ring buffer. Figure 3B An example of a descriptor ring 301 defined by P4 is shown. In some cases, the memory structure of an I / O device interface may be defined in descriptor format. In the example, a receive queue (RxQ) descriptor may have the following fields:
[0074]
[0075]
[0076] In the case of a transmit queue, the transmit queue descriptor may have a single DMA address field for the first data buffer segment to be sent. If there is only one segment, the single DMA address is sufficient to send the entire packet. In the case of more than one segment, a scatter / gather list may be used to describe the DMA address and the lengths of subsequent segments.
[0077] As described above, the provided I / O device interface extends the P4 programmable pipeline mechanism to the host driver. For example, a P4-programmed DMA interface can be presented directly to the host virtual functions and processing entities (e.g., an ARM CPU) of a network device or offload engine interface. The I / O device interface can support up to 2048 or more PCIe virtual functions for direct container mapping with multiple transmit and receive queues. Combining the programmable I / O device interface with P4 pipeline features allows offloading the host virtual switch / NIC to a programmable network device, thereby increasing bandwidth and reducing latency.
[0078] Matching Processing Unit (MPU)
[0079] In one aspect of the present invention, a match processing unit (MPU) is provided to process data structures. The data structures can include various types, such as data packets, management tokens, management commands from a host, processing tokens, descriptor rings, etc. The MPU can be configured to perform various operations according to the type of data being processed or different uses. For example, these operations can include table-based actions for processing packets, table maintenance operations (such as writing a timestamp to a table or collecting table data for export), management operations (such as creating a new queue or memory mapping, collecting statistics), and various other operations that may result in writing any type of changed data to the host memory (such as initiating bulk data processing).
[0080] In some embodiments, the MPU can process the data structure to update a memory-based data structure or initiate an event. The event may involve changing a packet, for example, changing the PHV field of a packet as described elsewhere herein. Or, the event may not be related to changing or updating a packet. For example, the event can be a management operation, such as creating a new queue or memory mapping, collecting statistics, initiating bulk data processing that may result in writing any type of changed data to the host memory, or performing calculations on a descriptor ring, scatter-gather list (SGL).
[0081] Figure 4 A block diagram of a match processing unit (MPU) 400 according to an embodiment of the present invention is shown. In some embodiments, the MPU unit 400 can include multiple functional units, a memory, and at least one register file. For example, the MPU unit can include an instruction fetch unit 401, a register file unit 407, a communication interface 405, an arithmetic logic unit (ALU) 409, and various other functional units.
[0082] In an illustrative example, the MPU unit 400 may include a write port or communication interface 405 that allows memory read / write operations. For example, the communication interface may support writing to or reading from an external memory (e.g., high bandwidth memory (HBM) of a host device) or an internal static random access memory (SRAM) in packets. The communication interface 405 may employ any suitable protocol, such as the Advanced Microcontroller Bus Architecture (AMBA) Advanced eXtensible Interface (AXI) protocol. AXI is a bus protocol for high-speed / high-end on-chip bus protocols and has channels associated with read, write, address, and write response that are separate, operate independently, and have transaction attributes such as multiple outstanding addresses or write data crossing. The features of the AXI interface 405 may include support for unaligned data transfer using byte strobes, burst-based transactions that only issue the starting address, separation of the address / control and data phases, issue of multiple outstanding addresses with out-of-order responses, and ease of adding a register phase to provide timing convergence. For example, when the MPU executes a table write instruction, the MPU may keep track of which bytes are written (i.e., dirty bytes) and which remain unchanged. When the table entry is flushed back to memory, the dirty byte vector may be provided to the AXI as a write strobe, allowing multiple writes to safely update a single table data structure as long as they do not write to the same byte. In some cases, the dirty bytes in the table do not need to be contiguous, and the MPU may only write back the table when at least one bit is set in the dirty vector. Although the packet data is transmitted in accordance with the AXI protocol in the packet data communication on-chip interconnect system of this exemplary embodiment according to this specification, it may also be applied to a packet data communication on-chip interconnect system that operates with other protocols that support locking operations, such as the Advanced High-Performance Bus (AHB) protocol or the Advanced Peripheral Bus (APB) protocol in addition to the AXI protocol.
[0083] The MPU 400 may include an instruction fetch unit 401 configured to fetch an instruction set from a memory external to the MPU based on an input table result or at least a portion of the table result. The instruction fetch unit may support branches and / or linear code paths based on the table result or a portion of the table result provided by the table engine. In some cases, the table result may include table data, key data, and / or a starting address of a set of instructions / programs. Details regarding the table engine will be described subsequently herein. In some embodiments, the instruction fetch unit 401 may include an instruction cache 403 for storing one or more programs. In some cases, one or more programs may be loaded into the instruction cache 403 upon receipt of the starting address of the program provided by the table engine. In some cases, a set of instructions or a program may be stored in a contiguous region of a memory unit, and the contiguous region may be identified by an address. In some cases, one or more programs may be fetched and loaded from an external memory via a communication interface 405. This provides flexibility, allowing the same processing unit to execute different programs associated with different types of data. In an example, when injecting a management packet header vector (PHV) into a pipeline to perform, for example, a management table direct memory access (DMA) operation or an entry aging function (i.e., adding a timestamp), one of the management MPU programs may be loaded into the instruction cache to perform the management function. The instruction cache 403 may be implemented using various types of memory, such as one or more SRAMs.
[0084] One or more programs may be any programs, such as P4 programs related to reading tables, constructing headers, DMA to / from memory regions in HBM or host devices, and other actions. One or more programs may be executed at any stage of the pipeline, as described elsewhere herein.
[0085] The MPU 400 may include a register file unit 407 to segment data between the memory of the MPU and the functional units or between the memory external to the MPU and the functional units of the MPU. The functional units may include, for example, an ALU, a gage, a counter, an adder, a shifter, an edge detector, a zero detector, a condition code register, a status register, etc. In some cases, the register file unit 407 may include a plurality of general-purpose registers (e.g., R0, R1, … Rn), which may initially be loaded with metadata values and then be used to store temporary variables during program execution until the program is completed. For example, the register file unit 407 may be used to store SRAM addresses, ternary content addressable memory (TCAM) search values, ALU operands, comparison sources, or action results. The register file unit of a stage may also provide a data / program environment to the register file of a subsequent stage and make the data / program environment available to the execution data path of the next stage (i.e., the source registers of the adder, shifter, etc. of the next stage). In one embodiment, each register of the register file is 64 bits and may initially be loaded with special metadata values, such as hash values from a table, lookups, packet sizes, PHV timestamps, programmable table constants, etc., respectively.
[0086] In some embodiments, the register file unit 407 may further include a comparator flag unit (e.g., C0, C1, … Cn) configured to store comparator flags. The comparator flags may be set by a calculation result generated by the ALU, which is then compared with a constant value in the encoded instruction to determine a conditional branch instruction. In an embodiment, the MPU may include eight one-bit comparator flags. However, it should be noted that the MPU may include any number of comparator flag units, and each unit may have any suitable length.
[0087] The MPU 400 may include one or more functional units, such as an ALU 409. The ALU may support arithmetic and logical operations on values stored in the register file unit 107. Then the results of ALU operations (e.g., add, subtract, and, or, exclusive or, not, nand, shift, and compare) may be written back to the register file. For example, the functional units of the MPU may update or change fields anywhere in the PHV, write to memory (e.g., table flush), or perform operations unrelated to PHV updates. For example, the ALU may be configured to perform calculations on a descriptor ring, a scatter-gather list (SGL), and control data structures loaded from the host memory into the general-purpose registers.
[0088] The MPU 400 may include various other functional units, such as meters, counters, action insertion units, etc. For example, the ALU may be configured to support P4-compatible meters. A meter is a type of action executable on table matching for measuring the data flow rate. A meter may include several bands, typically two or three bands, each with a defined maximum data rate and an optional burst size. Using the leaky bucket analogy, a meter band is a bucket filled with packet data rate and drained at a constant allowed data rate. If the set of data rates exceeding the quota is greater than the burst size, an overflow occurs. Overflowing one band triggers activity in the next band, which may allow a higher data rate. In some cases, the fields of a packet can be considered the result of overflowing the base band. This information can then be used to direct the packet to different queues, where it may be more likely to experience delays or losses in case of congestion. Counters can be implemented by MPU instructions. The MPU may include one or more types of counters for different purposes. For example, the MPU may include performance counters that count MPU stalls. The action insertion unit may be configured to push register file results back to the PHV for header field changes.
[0089] The MPU may be able to lock tables. In some cases, the table being processed by the MPU may be locked or marked as "locked" in the table engine. For example, when the MPU loads a table into its register file, the table address may be reported back to the table engine, causing future reads of the same table address to stall until the MPU releases the table lock. For example, the MPU may release the lock when an explicit table flush instruction is executed, when the MPU program ends, or when the MPU address changes. In some cases, the MPU may lock more than one table address, for example, one for previous table write-back and another address lock for the current MPU program.
[0090] MPU pipelining
[0091] A single MPU may be configured to execute the instructions of a program until the program is completed. Optionally or additionally, multiple MPUs may be configured to execute a program. In some embodiments, table results may be distributed to multiple MPUs. According to the MPU distribution mask configured for the table, table results may be distributed to multiple MPUs. This provides the advantage of preventing data stalls or a reduction in million packets per second (MPPS) when the program is too long. For example, if the PHV needs to read four tables in one stage, then each MPU program may be limited to only eight instructions to maintain 100 MPPS if operating at 800 mhz (in which case multiple MPUs may be expected).
[0092] To achieve the desired performance, any number of MPUs can be used to execute a program. For example, at least two, three, four, five, six, seven, eight, nine, or ten MPUs can be used to execute a program. Each MPU can execute at least a portion of the program or a subset of the instruction set. Multiple MPUs can perform the execution simultaneously or sequentially. Each MPU may or may not execute the same number of instructions. The configuration can be determined based on the length of the program (i.e., the number of instructions, cycles) and / or the number of available MPUs. In some cases, the configuration can be determined by application instructions received from the main memory of a host device operatively coupled to the multiple MPUs.
[0093] P4 pipeline
[0094] In one aspect, a flexible, high-performance match-action pipeline that can execute a wide range of P4 programs is provided. The P4 pipeline can be programmed to provide various features including, but not limited to, routing, bridging, tunneling, forwarding, network ACL, L4 firewall, flow-based rate limiting, VLAN tagging policies, membership, isolation, multicast and group control, label push / pop operations, L4 load balancing, L4 flow tables for analysis and flow-specific processing, DDOS attack detection, suppression, and telemetry data collection at any data packet field or flow state and various other states. Figure 5 A block diagram of an exemplary P4 ingress or egress pipeline (PIP pipeline) 500 in accordance with an embodiment of the present invention is shown.
[0095] In some embodiments, the provided invention may support a match + action pipeline. A programmer or compiler may break a packet handler into a set of dependent or independent table lookups and action processing stages (i.e., match + action), which are respectively mapped to a table engine and an MPU stage. The match + action pipeline may include multiple stages. For example, a packet entering the pipeline may first be parsed by a parser (e.g., parser 507) according to a packet header stack specified by a P4 program. This parsed representation of the packet may be referred to as a parsed header vector. The parsed header vector may then pass through the stages of the ingress match + action pipeline (e.g., stage 501-1, stage 501-2, stage 501-3, stage 501-4, stage 501-5, stage 501-6), where each stage is configured to match one or more table parsed header vector fields to a table and subsequently update the packet header vector (PHV) and / or table entries according to actions specified by the P4 program. In some cases, if the number of required stages exceeds the number of implemented stages, the packet may be recycled for additional processing. In some cases, the packet payload may move in a separate first-in-first-out (FIFO) queue until it is reassembled with its PHV in a deparser (e.g., deparser 509). The deparser may rewrite the original packet according to the changed (e.g., added, deleted, or updated) PHV fields. In some cases, a packet processed by the ingress pipeline may be located in a packet buffer for scheduling and possible replication. In some cases, once the packet is scheduled and leaves the packet buffer, it may be parsed again to create an egress parsed header vector. The egress parsed header vector may pass through a series of match + action pipeline stages in a manner similar to the ingress match + action pipeline, after which a final deparsing operation may be performed before the packet is sent to its destination interface or recycled for additional processing.
[0096] In some embodiments, the ingress pipeline and the egress pipeline may be implemented using the same physical block or processing unit pipeline. In some embodiments, the PIP pipeline 500 may include at least one parser 507 and at least one deparser 509. The PIP pipeline 500 may include multiple parsers and / or multiple deparsers. The parser and / or deparser may be a P4-compatible programmable parser or deparser. In some cases, the parser may be configured to extract packet header fields according to a P4 header definition and place the packet header fields into a packet header vector (PHV). The parser may select any field in the packet and align the information of the selected field to create a packet header vector. In some cases, after passing through the pipeline of the match + action stage, the deparse block may be configured to rewrite the original packet according to the updated PHV.
[0097] The packet header vector (PHV) generated by the parser can have any size or length. For example, the PHV can be at least 512 bits, 256 bits, 128 bits, 64 bits, 32 bits, 8 bits, or 4 bits. In some cases, when a long PHV (e.g., 6Kb) is needed to contain all relevant header fields and metadata, a single PHV can be time-division multiplexed (TDM) across several cycles. This TDM functionality provides the following benefits: allowing the invention to support variable-length PHVs, including very long PHVs, to enable complex features. The PHV length can vary as the packet passes through the match + action stage.
[0098] The PIP pipeline can include multiple match + action stages. After the parser 307 generates the PHV, the PHV can pass through the ingress match + action stage. In some embodiments, multiple stage units 501-1, 501-2, 501-3, 501-4, 501-5, 501-6 can be used to implement the PIP pipeline, and each stage unit can include a table engine 505 and multiple MPUs 503. The MPU 503 can be the same as the MPU described in Figure 1 In an illustrative example, four MPUs are used in one stage unit. However, any other number of MPUs, such as at least one, two, three, four, five, six, seven, eight, nine, or ten MPUs, can be used or grouped by the table engine.
[0099] The table engine 505 can be configured to support table matching for each stage. For example, the table engine 505 can be configured to hash, look up, and / or compare keys with table entries. The table engine 505 can be configured to control the table matching process by controlling the address and size of the table, the PHV fields used as lookup keys, and the MPU instruction vector that defines the P4 program associated with the table. The table results generated by the table engine can be distributed to multiple MPUs 303.
[0100] The table engine 505 can be configured to control table selection. In some cases, at the entry stage, the PHV can be examined to select which tables are enabled for the arriving PHV. The table selection criteria can be determined based on the information contained in the PHV. In some cases, the matching table can be selected based on the packet type information related to the packet type associated with the PHV. For example, the table selection criteria can be based on the packet type or protocol (e.g., Internet Protocol Version 4 (IPv4), Internet Protocol Version 6 (IPv6), and Multiprotocol Label Switching (MPLS)) or the next table ID determined by a previous stage or the previous stage. In some cases, the incoming PHV can be analyzed by table selection logic, and then the table selection logic generates a table selection key and uses a TCAM to compare the results to select a valid table. The table selection key can be used to drive table hash generation, table data comparison, and drive associated data to the MPU.
[0101] In some embodiments, the table engine 505 may include a hash generation unit. The hash generation unit may be configured to generate a hash result from a PHV input, and the hash result may be used to perform a DMA read from a DRAM or SRAM array. In an example, the input to the hash generation unit may be masked based on which bits in the table selection key contribute to the hash entropy. In some cases, the table engine may use the same mask for comparison with the returned SRAM read data. In some cases, the hash result may be scaled based on the size of the table, and then a table base offset may be added to create a memory index. The memory index may be sent to the DRAM or SRAM array and a read may be performed.
[0102] In some cases, the table engine 505 may include a TCAM control unit. The TCAM control unit may be configured to allocate memory to store multiple TCAM search tables. In an example, the PHV table selection key may be directed to a TCAM search phase prior to an SRAM lookup. The TCAM search tables may be configured to be up to 1024 bits wide and of a depth allowed by the TCAM resources. In some cases, multiple TCAM tables may be carved out of the shared quadrant TCAM resources. The TCAM control unit may be configured to allocate the TCAM to individual phases to prevent TCAM resource conflicts or to allocate the TCAM to multiple search tables within a phase. The TCAM search index result may be forwarded to the table engine for an SRAM lookup.
[0103] The PIP pipeline 500 may include multiple stage units 501-1, 501-2, 501-3, 501-4, 501-5, 501-6. The PIP pipeline may include any number of stage units, such as at least two, three, four, five, six, seven, eight, nine, ten stage units that may be used in the PIP pipeline. In the illustrated embodiment, six match + action stage units 501-1, 501-2, 501-3, 501-4, 501-5, 501-6 are grouped. This group of stage units may share a common set of SRAM 511 and TCAM 513. The SRAM 511 and TCAM 513 may be components of the PIP pipeline. This arrangement may allow the six stage units to divide the match table resources in any suitable ratio, which provides convenience to the compiler and simplifies the compiler's resource mapping task. Each PIP pipeline may use any suitable amount of SRAM resources and any amount of TCAM resources. For example, the illustrated PIP pipeline may be coupled to ten SRAM resources and four or eight TCAM resources. In some cases, the TCAM may be vertically or horizontally fused for a wider or deeper search.
[0104] Extended P4 pipeline
[0105] In one aspect, the provided invention can support an extended P4 programmable pipeline to allow direct docking with a host driver. The extended P4 programmable pipeline implements the IO device interface as described above. For example, a P4-programmed DMA interface can be directly coupled to a host virtual function (VF) as well as an Advanced RISC Machine (ARM) CPU or offload engine interface. The extended P4 pipeline can handle the required DMA operations and loops. The extended P4 pipeline can include features including but not limited to stateless NIC offloading such as TCP segmentation offloading (TSO) and receive-side scaling (RSS); memory swap table-style transaction services in the extended P4 pipeline; fine-grained load balancing decisions, which can be extended to a single data structure for performance-critical applications such as DPDK or key-value matching; initiation of TCP flow termination and proxy services; support for RDMA (RoCE) over converged Ethernet and similar remote direct memory access (RDMA) protocols; custom descriptors and SGL formats can be specified in P4 to match the data structures of performance-critical applications; new device and VF behavior can be modeled using a P4 program coupled with host driver development and various other features.
[0106] Data can be transferred between the packet domain in the P4 pipeline to / from the memory transaction domain in the host and NIC memory systems. This packet-to-memory transaction conversion can be performed by the extended P4 pipeline, which includes DMA write (TxDMA) and / or DMA read (RxDMA) operations. In this specification, the extended P4 pipeline including TxDMA can also be referred to as Tx P4 or TxDMA, and the extended P4 pipeline including RxDMA can also be referred to as Rx P4. The extended P4 pipeline can include the same match + action stages in the P4 pipeline, as well as a payload DMA stage at the end of the pipeline. Packets can be segmented or reassembled into a data buffer or memory region (e.g., RDMA register memory) according to the extended P4 program. The payload DMA stage can be a P4 extension that enables the programmable P4 network pipeline to extend to the host memory system and driver interface. This P4 extension allows for customized data structures and application interactions to adapt to the needs of an application or container.
[0107] The match tables used in the extended P4 pipeline can be programmable tables. The stages of the extended P4 pipeline can include multiple programmable tables, which can be present in SRAM, NIC DRAM, or host memory. For example, the host memory structure can include descriptor rings, SGLs, and control data structures, which can be read into the register file unit of the MPU for calculation. The MPU can add PHV commands to control DMA operations in and out of the host and NIC memories and insert DMA commands into the PHV for execution by the payload DMA stage. The extended P4 program can include, for example, completion queue events, interrupts, timer settings, and control register writes, as well as various other programs.
[0108] Figure 6 Illustrates an exemplary stage extended pipeline for Ethernet packet transmission (i.e., Tx P4 pipeline) 600. In the example, the table engine of stage 0 can extract the queue status (e.g., Q status) table for processing by the MPU of stage 0. In some cases, the queue status can also include an instruction offset address based on the queue pair type to accelerate MPU processing. Other separate Tx P4 programs can be written for Ethernet Tx queues, RDMA command queues, or any new type of transmission DMA behavior customized for a specific application. The number of supported Tx queue pairs can be determined based on the hardware scheduler resources assigned to each queue pair. As described above, the PHV can pass through each stage, in which the associated stage unit can execute a match + action program. The MPU of the final stage (e.g., stage 5) can insert DMA commands into the PHV for execution by the payload DMA stage (e.g., PDMA).
[0109] Figure 7 and Figure 8 Shows exemplary Rx P4 pipeline 700 and Rx P4 pipeline 800 according to embodiments of the present invention. The Rx P4 stage and / or the Tx P4 stage can be substantially similar to the P4 pipeline stages described elsewhere herein, with some different features. In some cases, the extended P4 stage can not use TCAM resources and can use fewer SRAM resources than the P4 stage. In some cases, the extended P4 pipeline can include a different number of stages than the P4 pipeline by having a payload DMA stage at the end of the pipeline. In some cases, the extended P4 pipeline can have a local PHV recycle data path that can not use packet buffers.
[0110] Referring to the Rx P4 pipeline (i.e., RxDMA P4 pipeline) as shown in Figure 7 the Rx P4 pipeline can include a plurality of stage units 701-1, 701-2,... 701-n, each stage unit can have the same asFigure 5 physical blocks that are the same as the stage units described in
[0111] In some embodiments, the Rx P4 pipeline 700 may include a PHV allocator block 703 configured to generate an RxDMA PHV. For example, metadata fields of the PHV required for RxDMA (e.g., a logical interface (LIF) ID) may be passed through the packet buffer as a consecutive field block before the packet from the P4 network pipeline. Before entering the first stage of the RxDMA P4 pipeline, the PHV allocator block 703 may extract the pre - positioned metadata and locate it in the RxDMA PHV. The PHV allocator block 703 may maintain a count of the number of PHVs currently in the RxDMA pipeline, as well as a count of the number of packet payload bytes in the pipeline. In some cases, when the PHV count or the total packet byte count exceeds a high watermark, the PHV allocator block 703 may stop accepting new packets from the packet buffer. This provides the benefit of ensuring that packets recycled from the payload DMA block 705 have priority to be processed and leave the pipeline.
[0112] The Rx P4 pipeline may include a packet DMA block 705 configured to control the order between related events. The packet DMA block may also be referred to as a payload DMA block. Packet data may be sent to the packet DMA block 705 in a first - in - first - out manner to wait for DMA commands created in the Rx P4 pipeline. The packet DMA block at the end of the Rx P4 pipeline may execute packet DMA write commands, DMA completion queue (CQ) write commands, interrupt assertion writes, and doorbell writes in the order in which the DMA commands are placed in the PHV.
[0113] Referring to the Tx P4 pipeline 800 shown in Figure 8 the Tx P4 pipeline may include a plurality of stage units 801 - 1, 801 - 2, …… 801 - k, and each stage unit may have the same physical blocks as the stage units described in Figure 7 The number of stage units in the Tx P4 pipeline may or may not be the same as the number of stage units in the above - mentioned Rx P4 pipeline. In an example, the Tx P4 pipeline may be used to transmit packets from a host or NIC memory. The Tx queue scheduler may select the next queue to be serviced and submit the LIF, QID to the start of the Tx P4 pipeline.
[0114] The Tx P4 pipeline may include an empty PHV block 803 configured to generate addresses to be read by the table engine in stage 0. The empty PHV block 803 may also insert information into the inherent fields of the PHV, such as LIF or LIF type. The empty PHV block 803 may also insert recycled PHVs, as well as software-generated PHVs, into the pipeline from the final stage of the Tx P4 pipeline. The Tx P4 pipeline may include a packet DMA block 805, similar to the packet DMA block as described in Figure 7 . In some embodiments, the DMA commands generated in the Tx P4 pipeline may be arranged in contiguous space such that the commands can be executed in order as long as the first and last commands are indicated.
[0115] In some embodiments, the Tx DMA pipeline, Rx DMA pipeline, and P4 pipeline may be able to insert software-generated PHVs before the first stage of their respective pipelines. The software may use the generated PHVs to start an MPU program, perform table changes, or initiate DMA commands from an extended P4 pipeline.
[0116] In one aspect, a system may be provided that includes a Tx DMA pipeline, an Rx DMA pipeline, a P4 pipeline, and other components. The system may support host interface features based on an extended P4 pipeline (e.g., DMA operations and loops), provide improved network performance (e.g., increased MMPS with reduced data stalls), fault detection and isolation, P4-based network features (e.g., routing, bridging, tunneling, forwarding, network ACL, L4 firewall, flow-based rate limiting, VLAN tagging policies, membership, isolation, multicast and group control, label push / pop operations, L4 load balancing, L4 flow tables for analysis and flow-specific processing, DDOS additional detection, suppression, telemetry data collected on any data packet field or flow state), security features, and various other features.
[0117] Figure 9An example of an extended transmit pipeline (i.e., TxDMA pipeline) 900 is shown. The provided IO device interface can support TxDMA pipeline operations to send packets from a host system or NIC memory. The IO device interface can have improved performance, such as 16 reads per stage (i.e., 500 ns host latency), 32 MPPS or higher for direct packet processing, and reduced latency. The provided IO device interface can include other features, such as flexible rate limit control. For example, a logical interface (LIF) can be rate-limited independently by a Tx scheduler. Each LIF can have a dedicated qstate array that has a programmable base address and a programmable number of queues in HBM. The device interface can support up to 2K or more LIFs. A LIF can contain eight or more queue types, each queue type having a programmable number of queues. For example, the qstate array may have a separate base address in HBM. This allows each LIF to independently scale the number and type of queue pairs it owns up to the maximum array size determined at virtual function creation. There may be one or more qstate types (e.g., eight qstate types) in the qstate array. Each type can have independent entry size control as well as an independent number of entries. This enables each LIF to flexibly mix different types of queues.
[0118] As described above, the IO device interface can be a highly optimized ring-based I / O queue interface. An extended P4 program can be used to rate-limit to a finer granularity than a scheduler-based rate limiter. The extended P4 program can use the time and bucket data stored in qstate to rate-limit individual queue pairs. The extended P4 rate limit can be applied in the presence of scheduler (per LIF) rate limits, thus allowing application of per-VM or per-interface rates, while per-queue limits can cause congestion or other fine-grained rate control. For example, the extended P4 rate limit program can set the XOFF state for the main data transfer ring while leaving other rings for timer, management, or congestion message transmission. In an exemplary process for implementing the rate of each queue limited in an extended P4 pipeline, the program can first determine whether the current queue is below or above its target rate. If it is determined to be above its target rate, the program can discard the current scheduler token and set timer resources to reschedule the queue at some time in the future calculated based on the current token level and rate. Next, the program can disable the scheduler bit for the queue, but only for the rings in the COS of the flow control. The queue may not be rescheduled until a new doorbell event occurs for that ring, even if work on the main data transfer ring still needs to be completed.
[0119] As Figure 9The illustrated TxDMA pipeline 900 may include a pipeline of stage units 901-1, 901-2, … 901-8 and a payload DMA block 902. Each of the plurality of stage units 901-1, 901-2, ... 901-8 may have a physical block of the same stage unit as described in Figure 8 In an example, the TxDMA pipeline may be used to send packets from a host or NIC memory. The Tx queue scheduler may select the next service queue and submit the LIF, QID to the beginning of the TxDMA pipeline.
[0120] In some cases, a single queue may include up to eight or more rings to allow multiple event signaling options. A queue under direct host control may contain at least one host-accessible ring, typically the primary descriptor ring required for the queue type. Host access to the ring p_index may be through a doorbell mapping. The illustrated example shows multiple stages for implementing a Virtio program. The table engine (TE) of stage 0 may fetch a queue status (e.g., Q status) table for processing by the MPU of stage 0. In some cases, the queue status may also contain an instruction offset address based on the queue pair type to speed up MPU processing. Separate TxDMA programs may be written for Ethernet Tx queues, RDMA command queues, or any new type of transmission DMA behavior customized for a particular application. The number of supported Tx queue pairs may be determined based on the hardware scheduler resources allocated to each queue pair.
[0121] It should be noted that various embodiments may be used in conjunction with one or more types of wireless or wired communication signals and / or systems, e.g., radio frequency (RF), infrared (IR), frequency division multiplexing (FDM), orthogonal FDM (OFDM), time division multiplexing (TDM), time division multiple access (TDMA), extended TDMA (E-TDMA), general packet radio service (GPRS), extended GPRS, code division multiple access (CDMA), wideband CDMA (WCDMA), CDMA 2000, multi-carrier modulation (MDM), discrete multi-tone (DMT), ZigBee TM etc. Embodiments of the present invention may be used in a variety of other devices, systems, and / or networks.
[0122] While the preferred embodiments of the subject matter have been shown and described herein, it will be readily apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the subject matter described herein may be employed in practicing the invention.
Claims
1. A method for a programmable IO device interface, comprising: Provide a pipeline of programmable device registers, memory-based data structures, DMA blocks, and processing entities, wherein the programmable IO device interfaces with a hypervisor in a host system using at least one virtual function and at least one physical function, and wherein the pipeline of the processing entities is configured to: a) Receive a packet including a header part and a payload part, wherein the header part is used to generate a packet header vector; b) Generate a table result by using a table engine and performing a packet matching operation, wherein the table result is generated based at least in part on data stored in a programmable match table and the packet header vector; c) Receive, at a match processing unit, an address of a set of instructions associated with the programmable match table and the table result; d) Execute, by the match processing unit, one or more actions according to the loaded set of instructions until the instructions are completed, wherein the one or more actions include inserting a DMA command into the packet header vector; and e) Perform a DMA operation by the DMA block according to the inserted DMA command.
2. The method according to claim 1, further comprising providing the header part to a subsequent circuit, wherein the subsequent circuit is configured to assemble the header part into a corresponding payload part.
3. The method according to claim 1, wherein the programmable match table includes a DMA register table, a descriptor format, or a control register format.
4. The method according to claim 3, wherein the programmable match table is selected based on packet type information related to a packet type, the packet type being associated with the header part.
5. The method according to claim 3, wherein the programmable match table is selected based on an ID of a match table selected by the table engine.
6. The method according to claim 1, wherein the table result includes a key related to the programmable match table and a match result of the matching operation.
7. The method according to claim 1, wherein a memory unit of the match processing unit is configured to store multiple sets of instructions.
8. The method according to claim 7, wherein the multiple sets of instructions are associated with different actions.
9. The method according to claim 7, wherein the multiple sets of instructions are stored in a continuous area of the memory unit, and the continuous area is identified by an address of the multiple sets of instructions.
10. The method according to claim 1, wherein the one or more actions further include updating the programmable match table.
11. The method according to claim 1, further comprising locking the match table for exclusive access by the match processing unit when the match processing unit processes the match table.
12. The method according to claim 1, wherein multiple packets are processed in a non-stopping manner.
13. A device having a programmable IO device interface, comprising: a) A first memory unit storing multiple programs thereon, wherein the multiple programs are associated with multiple actions, the multiple actions including inserting a DMA command into a packet header vector; b) A second memory unit for receiving and storing a table result, wherein the table result is provided by a table engine configured to perform packet matching operations on (i) the packet header vector included in the header part of a packet and (ii) data stored in a programmable match table; And c) A circuit for executing a program selected from the plurality of programs in response to the table result and an address received by the device, wherein the program execution continues until completion and the program is associated with the programmable match table; Wherein the programmable IO device interfaces with a hypervisor in a host system using at least one virtual function and at least one physical function.
14. The device according to claim 13, wherein the device is configured to provide the header part to a subsequent circuit, and wherein the header part is modified by the circuit.
15. The device according to claim 14, wherein the subsequent circuit is configured to assemble the modified header part into a corresponding payload part.
16. The device according to claim 13, wherein the programmable match table includes a DMA register table, a descriptor format, or a control register format.
17. The device according to claim 16, wherein the programmable match table is selected based on packet type information related to the packet type associated with the header part.
18. The device according to claim 16, wherein the programmable match table is selected based on the ID of the match table selected by the table engine.
19. The device according to claim 13, wherein each of the plurality of programs includes a set of instructions stored in a contiguous region of the first memory unit, and the contiguous region is identified by the address.
20. The device according to claim 13, wherein the plurality of actions includes updating the programmable match table.
21. The device according to claim 13, wherein the circuit is further configured to lock the match table for exclusive access by the device while the device processes the programmable match table.
22. The device according to claim 13, wherein the plurality of actions further includes initiating an event that is not related to modifying the header part of the packet.
23. The device according to claim 13, wherein the plurality of actions further includes updating a memory-based data structure that includes at least one of the following: an administrative token for initiating an event, an administrative command, and one or more processing tokens.
24. A system comprising a plurality of devices according to claim 19, wherein the plurality of devices are coordinated to execute the set of instructions or the plurality of actions concurrently or sequentially according to a configuration.
25. The system according to claim 24, wherein the configuration is determined by application instructions received from a main memory of a host device operably coupled to the plurality of devices.
26. The system according to claim 24, wherein the plurality of devices are arranged to process a plurality of packets according to a pipeline of stages.
Citation Information
Patent Citations
Flexible header alteration in network devices
CN111490969A
Network system including match processing unit for table-based actions
CN111684769A