Dynamic data path at the edge gateway

By using data path pipelines at the gateway, dynamically determine the processing stage sequence, and combining centralized and distributed routing stages, the efficiency reduction problem of existing gateways under high traffic conditions is solved, and efficient traffic processing and resource utilization is achieved.

CN112769695BActive Publication Date: 2025-07-01VMWARE INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202110039636.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-11-02
Filing Date
2016-10-29
Publication Date
2025-07-01
Estimated Expiration
2036-10-29

AI Technical Summary

Technical Problem

When existing gateways process large amounts of network traffic, it is difficult for existing gateways to dynamically adjust data pathways, resulting in reduced efficiency and waste of resources.

Method used

By using data path pipelines at the edge of the network, dynamically determine the processing phase sequence, combining centralized and distributed routing phases, performing service provision phases such as NAT and firewalls, utilizing DP configuration databases to store configuration data for dynamic routing and service processing.

Benefits of technology

It realizes efficient processing of the gateway under high traffic conditions, dynamically adjusts the data path pipeline to meet different traffic needs, and improves resource utilization and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112769695B_ABST
    Figure CN112769695B_ABST
Patent Text Reader

Abstract

The present disclosure relates to dynamic data paths at an edge gateway. A novel design of a gateway is provided that processes traffic in and out of a network by using a data path pipeline. The data path pipeline includes multiple stages for performing various data plane packet processing operations at the network edge. The processing stages include a centralized routing stage and a distributed routing stage. The processing stages may include service provisioning stages such as NAT and firewall. The gateway caches the results of previous packet operations and reapplies the results to subsequent packets that meet certain criteria. For packets that do not have an applicable or valid result from a previous packet processing operation, the gateway data path daemon performs the pipelined packet processing stages and records a set of data from each stage of the pipeline and synthesizes this data into a cache entry for subsequent packets.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application date of October 29, 2016, the application number of 201680069286.6, and the invention name of "Dynamic Data Path at the Edge Gateway". Technical Field

[0002] The present invention generally relates to a dynamic data path at an edge gateway. Background Art

[0003] A gateway is a network point that serves as an entrance to another network. In a network provided by a data center, the computing resources allocated as gateway nodes facilitate and regulate the traffic between the data center network and the external physical network. A gateway is typically associated with a router that knows where to direct a given packet of data arriving at the gateway and a switch that supplies the actual path for the given packet to enter and exit the gateway.

[0004] A gateway is also a computing node that provides various network traffic services such as a firewall, network address translation (NAT), and security protocols (such as SSL-based HTTP). As data centers become larger and provide more computing and networking resources, gateways must also handle more traffic. In other words, the gateway and its associated routers and switches must perform more switching, routing, and service tasks at a faster speed. Summary of the Invention

[0005] Some embodiments provide a gateway that processes traffic entering and exiting a network by using a data path pipeline. The data path pipeline includes multiple stages for performing various data plane packet processing operations at the network edge. In some embodiments, the processing stages include a centralized routing stage and a distributed routing stage. In some embodiments, the processing stages include a service provisioning stage such as NAT and a firewall.

[0006] In some embodiments, the sequence of stages to be executed as part of the data path pipeline is dynamically determined based on the content of the received packet. In some embodiments, each stage of the data path pipeline corresponds to a packet processing logic entity (such as a logical router or a logical switch) in a logical network, and the next stage identified by the packet processing at that stage corresponds to the next hop of the packet in the logical network, where the next hop is another logical entity.

[0007] In some embodiments, the packet processing operations for each logical entity are based on configuration data stored in the data path configuration database of that logical entity. Such configuration data also defines the criteria or rules for identifying the next hop for a packet. In some embodiments, such next-hop identification rules are stored in the DP configuration database as a routing table or forwarding table associated with that stage. In some embodiments, such next-hop identification rules allow the data path daemon to determine the identity of the next hop by examining the contents of the packet and / or by recording the logical port through which the packet enters the logical entity.

[0008] In some embodiments, each packet processing stage is implemented as a function call to a data path daemon thread. In some embodiments, the functions called to implement the various stages of the data path are part of the programming of the data path daemon operating at the core, but the functions called perform different operations based on different configuration data for different network identities. In other words, the programming of the core provides functions that can be called by the data path daemon to perform the functions of various logical routers, logical switches, and service providing entities. The function call uses the contents of the packet as an input argument. In some embodiments, the function call also uses the identity of the logical port through which the packet enters the corresponding logical entity as an input argument. In some embodiments, the function call also identifies the egress port, which is used to identify the ingress port for the next function call for the next pipeline stage. In some embodiments, each logical port of each logical entity is associated with a Universally Unique Identifier (UUID) such that the logical port can be uniquely identified by the gateway. The UUID of the logical port also allows the data path daemon to identify the logical entity to which the logical port belongs, which in turn allows the data path daemon to identify the configuration data of the identified logical entity and execute the corresponding pipeline stage.

[0009] In some embodiments, some of the logical entities / constructs / components of the logical network are distributed among multiple physical machines in the data center, and some logical entities / entities are not distributed but are centralized or concentrated on one physical machine. In some embodiments, such a centralized router acts as a centralized point for routing packets between the logical network and an external router. In some embodiments, when processing an incoming packet, the data path daemon will execute both distributed logical entities and centralized logical entities as part of its pipeline stages. In some embodiments, a service router is a centralized logical router. Each service router has only one instance running on one gateway machine. Thus, the data path daemon running on the gateway machine will call the centralized or gateway machine-concentrated service router as one of its data path pipeline stages.

[0010] In some embodiments, a data center supports multiple logical networks for multiple different tenants. Different tenant logical networks share the same set of gateway machines, and each gateway machine provides packet switching, forwarding, and routing operations for all connected tenant logical networks. In some embodiments, a data path daemon is capable of performing a packet processing phase for packets going to and coming from different logical networks belonging to different tenants. In some of these embodiments, a DP configuration database provides configuration data (i.e., routing tables, forwarding tables, etc.) and service specifications that enable tenant-specific packet forwarding operations at the gateway.

[0011] In some embodiments, in addition to performing L3 routing and L2 routing pipeline phases, a gateway data path daemon also performs a service provisioning phase for L4 to L7 processing. These services support end-to-end communication between a source application and a destination application and are used whenever a message is passed from or to a user. The data path daemon applies these services to packets at the vantage point of an edge gateway without changing the applications running at the source or destination of the packet. In some embodiments, the data path may include service phases for traffic filtering services (such as a firewall), address mapping services (such as NAT), and encryption and security services (such as IPSec and HTTPS).

[0012] In some embodiments, some or all of these service provisioning phases are performed when the data path daemon is executing a service router pipeline phase. Additionally, in some embodiments, the data path daemon may perform different service provisioning pipeline phases for different packets. In some embodiments, the data path daemon performs different service provisioning phases based on the L4 flow to which a packet belongs and the state of the flow. In some embodiments, the data path daemon performs different service provisioning phases based on the tenant to which a packet belongs.

[0013] In some embodiments, instead of performing pipeline phases on all packets, the gateway caches the results of previous packet operations and reapplies the results to subsequent packets that meet certain criteria (i.e., cache hits). For packets that do not have an applicable or valid result from a previous packet processing operation, i.e., cache misses, the gateway data path daemon performs pipeline packet processing phases. In some embodiments, when the data path daemon performs pipeline phases to process a packet, it records a set of data from each phase of the pipeline and synthesizes this data into a cache entry for subsequent packets. When the data path pipeline is being executed, some or all of the execution phases emit data or instructions that will be used by the synthesizer to synthesize cache entries. In some embodiments, the cache entry synthesis instructions or data emitted by the pipeline phases include a cache enable field, a bitmask field, and an action field.

[0014] The synthesizer collects all cache entry synthesis instructions from all pipeline stages and synthesizes entries in the flow cache from all received instructions, unless one or more pipeline stages specify that cache entries should not be generated. The synthesized cache entries specify the final action for packets that meet certain criteria (i.e., belong to certain L4 flows). When generating cache entries, in some embodiments the synthesizer also includes a timestamp that specifies the time at which the cache entry was created. This timestamp will be used to determine whether the cache entry is valid for subsequent packets.

[0015] Some embodiments dynamically update the DP configuration database even while the data path daemon is actively accessing the DP configuration database. To ensure that the data path daemon does not use incomplete (and thus corrupted) configuration data for its pipeline stages when the DP configuration database is updated, some embodiments maintain two copies of the DP configuration database. One copy of the database serves as a staging area for new updates from the network controller / manager, such that the data path daemon can safely use the other copy of the database. Once the update is complete, the roles of the two database copies are atomically reversed. In some embodiments, the network controller waits for the data path daemon to complete its current run-to-completion packet processing pipeline stage before switching.

[0016] The foregoing invention content is intended to serve as a brief introduction to some embodiments of the present invention. It is not meant to be an introduction or overview of all inventive subject matter disclosed in this document. The following detailed description and the accompanying drawings that refer to the detailed description will further describe the embodiments described in the invention content and other embodiments. Therefore, to understand all embodiments described in this document, a comprehensive review of the invention content, detailed description, and the accompanying drawings is required. Additionally, the claimed subject matter is not limited by the illustrative details in the invention content, detailed description, and the accompanying drawings, but rather is defined by the appended claims, because the claimed subject matter can be embodied in other specific forms without departing from the spirit of the subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The novel features of the invention are set forth in the appended claims. However, for purposes of explanation, several embodiments of the invention are set forth in the following figures.

[0018] Figure 1 The conceptual diagram illustrates a data center where traffic to and from an external network passes through a gateway.

[0019] Figure 2Illustrates in more detail the gateway machine that implements the data path pipeline.

[0020] Figure 3 Illustrates the dynamic identification of processing phases by the data path daemon.

[0021] Figure 4 Conceptually illustrates the data path daemon that executes each phase of the data path pipeline as a function call.

[0022] Figure 5 Illustrates an example DP configuration database that provides configuration data for each data path pipeline phase.

[0023] Figure 6 Conceptually illustrates the process executed by the processing core when the data path pipeline is executed using the DP configuration database.

[0024] Figure 7 Illustrates a logical network with distributed and centralized logical entities.

[0025] Figure 8 Illustrates the gateway data path daemon that executes the pipeline phases for incoming packets to the logical network from an external network to the data center.

[0026] Figure 9 Illustrates the gateway data path daemon that executes the pipeline phases for outgoing packets from the logical network of the data center to an external network.

[0027] Figure 10 Illustrates the logical view of the entire network of the data center.

[0028] Figure 11 Illustrates the data path daemon that performs gateway packet processing for different tenants at the gateway machine.

[0029] Figure 12 Shows the data path daemon that processes packets by calling the pipeline phases corresponding to various logical entities.

[0030] Figure 13 Illustrates the gateway data path daemon that executes the service provisioning pipeline phases for different tenants.

[0031] Figure 14 Conceptually illustrates the process for providing services within a logical router.

[0032] Figure 15a -b Illustrates the data path daemon that maintains a cache to accelerate packet processing.

[0033] Figure 16 Illustrates the synthesis of cache entries for the data path cache.

[0034] Figure 17 Illustrates an example synthesis of an aggregated cache entry and an exact match cache entry.

[0035] Figure 18 Illustrates an example of a data path stage of an action that specifies an action that overrides all other actions.

[0036] Figure 19 Illustrates checking entries in a data path cache to determine whether there is a cache miss or a cache hit.

[0037] Figure 20 A conceptual diagram illustrates a process for operating a data path cache.

[0038] Figure 21 Illustrates a gateway with a DP configuration database that supports updates in an atomic manner.

[0039] Figure 22a -b Illustrates an atomic update of the data path configuration database 2110.

[0040] Figure 23 A conceptual diagram illustrates a process for controlling read and write pointers of a DP configuration database.

[0041] Figure 24 Illustrates the architecture of a gateway machine according to some embodiments of the present invention.

[0042] Figure 25a A conceptual diagram illustrates an RTC thread that uses IPC to communicate with a service process to provide a service.

[0043] Figure 25b A conceptual diagram illustrates an RTC thread that uses the Linux kernel to communicate with a service process to provide a service.

[0044] Figure 26 Illustrates a computing device that serves as a host machine for virtualization software that runs some embodiments of the present invention.

[0045] Figure 27 A conceptual diagram illustrates an electronic system that implements some embodiments of the present invention. Detailed Description

[0046] In the following description, many details are set forth for purposes of explanation. However, one of ordinary skill in the art will recognize that the present invention may be practiced without the use of these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the description of the present invention with unnecessary detail.

[0047] Some embodiments provide gateways that process traffic into and out of a network by using a data path pipeline. The data path pipeline includes multiple stages for performing various data plane packet processing operations at the network edge. In some embodiments, the processing stages include a centralized routing stage and a distributed routing stage. In some embodiments, the processing stages include service provisioning stages, such as NAT and firewall.

[0048] Figure 1 The conceptual diagram shows a data center 100 where traffic to and from an external network 190 passes through gateways 111 - 112. Each gateway operates a data path pipeline (141 and 142 respectively) for processing packets passing through the gateway.

[0049] The data center includes various other computing and networking resources 121 - 129 interconnected by a provider network. These resources communicate with each other via the provider network and with the external network 190 via network traffic over a physical communication medium, which can include wired communication such as Ethernet or wireless communication such as WiFi. Packets from the computing and networking resources 121 - 129 can reach the external network 190 via one of the gateways 111 - 112, and packets from the external network 190 can reach the computing and networking resources 121 - 129 via one of the gateways 111 - 112. Thus, the gateways of the network are considered to be at the edge of the network and are therefore also referred to as edge devices.

[0050] In some embodiments, some of these resources are provided by computing devices that serve as host machines 121 - 129. Some of these host machines operate virtualization software that allows these host machines to host various virtual machines (VMs). The host machines running virtualization software will be described in more detail below by reference Figure 26 to. In some embodiments, the gateway itself is a host machine, and the data path pipeline (141 or 142) of the gateway is provided by one of the VMs running on its virtualization software. Some of these resources run as "bare metal", i.e., without virtualization software. In some embodiments, the gateway is a bare metal computing device that operates its data path pipeline directly on its own operating system without virtualization software.

[0051] In some embodiments, packet traffic within a data center is steered by using an overlay logical network such as Virtual eXtensible LAN (VXLAN), Generic Network Virtualization Encapsulation (GENEVE), and Network Virtualization using Generic Routing Encapsulation (NVGRE). VXLAN. In some of these embodiments, each of the host machines and the gateway machines is a VXLAN endpoint (referred to as a VTEP) that uses overlay encapsulation to transmit packets. In some embodiments, the external physical network is steered by VLAN, and the gateway relays traffic between the data center and the external network by converting VXLAN packets into VLAN packets and vice versa.

[0052] In some embodiments, the compute and networking resources of the data center implement one or more logical networks, each of which can access gateways 111 - 112 for traffic to and from the external network 190. In some embodiments, each logical network has its own set of logical switches and logical routers for steering the network traffic of the logical network. Some or all of these logical routers and switches are provided by software operating in the host machines (either as virtualization software or as a program executing on a bare metal host machine). In some embodiments, some of the logical routers and switches operate as stages within their respective data path pipelines 141 - 142 in gateways 111 - 112. In some embodiments, the data center includes a network manager 180 for provisioning / creating the logical networks in the data center 100 and a network controller 170 (or a cluster of controllers) for controlling the various logical routers and switches of the various logical networks, including those operating in gateways 111 - 112. The logical routers and switches are described in U.S. Patent Application 14 / 814,473, titled "Logical Router with Multiple Routing Components," filed on Jun. 30, 2015, which is incorporated herein by reference.

[0053] The control plane of some embodiments configures and manages one or more logical networks for one or more tenants of a host system (e.g., a data center). In some embodiments, the logical networks of the host system use a set of logical forwarding elements (e.g., logical L2 and L3 switches) to logically connect a set of endpoints (e.g., virtual machines, physical servers, containers, etc.) to a set of physical machines. In some embodiments, different subsets of endpoints reside on different host machines that execute managed forwarding elements (MFEs). The MFE implements the logical forwarding elements of the logical network to which the local endpoints are logically connected. In various embodiments, these MFEs can be flow-based forwarding elements (e.g., Open vSwitch) or code-based forwarding elements (e.g., ESX), or a combination of both. These different types of forwarding elements implement the various logical forwarding elements differently, but in each case, they execute a pipeline for each logical forwarding element that may need to process packets.

[0054] Figure 2 A gateway machine that implements a data path pipeline is illustrated in more detail. As shown, gateway 111 includes processing cores 211 - 214 and a network interface controller (NIC) 220. NIC 220 receives data packets from the network communication medium to which gateway 111 is connected and provides the received packets to cores 211 - 214 for processing.

[0055] Each processing core is operating one or more processing threads. Specifically, core 211 is operating data path pipeline 141 as a processing thread, which is referred to as data path daemon 241. As shown, data path daemon 241 receives packet 270 and processes packet 270 through a series of stages 221 - 229 to produce processed packet 275. In some embodiments, each core executes only one thread at a time, and each thread processes one packet at a time. In other words, each packet processing thread is a run-to-completion (RTC) thread that does not start processing another packet until it has completed processing the current packet (i.e., 270) through all of its stages 221 - 229.

[0056] The operation of the data path daemon 241 is defined or specified by the Data Path Configuration Database Store (DP Configuration Database) 230. The configuration data stored in the DP Configuration Database 230 specifies what functions or operations each stage of the pipeline should perform for each incoming packet. For some stages corresponding to logical routers or switches, the DP Configuration Database provides the content specifying the next hop for the routing table or forwarding table in some embodiments. For some stages corresponding to network services such as firewalls, the DP Configuration Database 230 provides service rules. In some embodiments, the network controller 170 (or network manager 180) loads and updates the content of the DP Configuration Database 230.

[0057] Several more detailed embodiments of the present invention are described below. Part I discusses the dynamic pipeline stages for packet processing at the gateway. Part II describes the caches for the gateway data path pipeline. Part III describes the update of the configuration database for the gateway data path pipeline. Part IV describes the software architecture of the gateway implementing the data path pipeline. Part V describes the computing device operating the virtualization software. Finally, Part VI describes the electronic system implementing some embodiments of the present invention.

[0058] I. Dynamic Pipeline Stages

[0059] In some embodiments, the sequence of stages to be executed as part of the data path pipeline is dynamically determined based on the content of the received packet. In Figure 2 this context, this means that the content of the packet 270 dynamically determines which processing stages are to be executed as part of the data path pipeline 141. In some embodiments, when processing a packet at a particular stage, the core 211 determines or identifies the next stage to be used for processing the packet. In some embodiments, each stage of the data path pipeline corresponds to a packet processing logical entity such as a logical router or a logical switch, and the next stage identified by the packet processing at that stage corresponds to the next hop of the packet in the logical network, where the next hop is another packet processing logical entity. (For simplicity, packet forwarding logical entities will be referred to as logical entities throughout this document).

[0060] In some embodiments, the pipeline stage corresponding to a logical router or logical switch is a full functional model of the logical router or switch, i.e., it specifies all of its logical ports, its routing / forwarding tables, the services it provides, its security policies, its encapsulation protocols, etc. In some embodiments, all of these characteristics of the logical router are specified by a computer-executable code package and can be executed as a pipeline stage through function calls. It performs forwarding either through the destination MAC (L2 switching) or through the destination IP (L3 routing). Such a pipeline stage can thus be distinguished from the flow tables under OpenFlow or Open vSwitch, where the flow tables under OpenFlow or Open vSwitch perform flow forwarding based on a set of flow entries, each entry describing a match condition and a corresponding action.

[0061] In some embodiments, the packet processing operations (i.e., the pipeline stage) for each logical entity are based on configuration data stored in the DP configuration database for that logical entity. This configuration data also defines the criteria or rules for identifying the next hop for a packet. In some embodiments, such next-hop identification rules are stored in the DP configuration database as a routing table or forwarding table associated with that stage. In some embodiments, such next-hop identification rules allow the data path daemon to determine the identity of the next hop by examining the contents of the packet (e.g., its source and destination addresses) and / or by recording the logical port through which the packet enters the logical entity. In other words, the DP configuration database can be considered to store the logical relationships between the various hops of the logical network, and the data path daemon processes each packet by traversing the logical network based on these logical relationships and the contents of the packet.

[0062] Figure 3 Illustrates the data path daemon's dynamic identification of processing stages. As shown, the core 211 is operating the data path daemon 241 as a processing thread. The data path daemon 241 is processing packet 371 according to the logical network 300, the configuration data of which is stored in the DP configuration database 230.

[0063] As shown in the figure, the logical network 300 includes service routers 311 and 312 (SR1 and SR2), logical switches 321, 322, and 323 (LS1, LS2, and TLS), and a distributed router 313 (DR). Each of the service routers SR1 and SR2 has an uplink (U1 and U2) for connecting to the external network 190. The logical switch TLS is a transit logical switch that provides L2 switching for packets from routers SR1, SR2, and DR1, and the routers SR1, SR2, and DR1 are assigned logical ports with MAC addresses "MAC1", "MAC2", and "MAC3", respectively. On the other hand, the distributed router DR1 provides L3 routing between the L2 network segments defined by the logical switches LS1, LS2, and TLS.

[0064] The figure illustrates how the data path daemon 241 processes two different packets 371 and 372 according to the configuration data stored in the DP configuration database 230. These two different packets cause the data path daemon 241 to traverse the logical network 300 differently and execute different corresponding pipeline stages.

[0065] Packet 371 is a packet from the external network 190 destined for the VM 381 (VM1) behind the L2 segment of the logical switch LS1. When the processing core 211 receives packet 371, it executes the packet processing stage 351 corresponding to the service router SR1. The operation of stage 351 is defined by the configuration in the DP configuration database. The service router SR1 logically forwards the packet to the logical switch TLS, which causes the data path daemon 241 to identify the next packet processing stage 352 corresponding to the logical switch TLS.

[0066] The processing stage 352 is configured by the DP configuration database 230 to perform L2 switching operations as the logical switch TLS, and the logical switch TLS forwards packet 371 from its "MAC1" port to its "MAC3" port. The MAC3 port corresponds to the distributed router DR1, and the data path daemon 241 correspondingly identifies the next packet processing stage 353 corresponding to DR1.

[0067] The processing stage 353 is configured by the DP configuration database 230 to perform L3 routing operations as the distributed logical router DR1, which operates according to the routing table provided by the DP configuration database 230. Based on the routing table and the destination IP address of packet 371, the logical router DR1 routes packet 371 from the L2 segment defined by the logical switch TLS to the L2 segment defined by the logical switch LS1. Correspondingly, the data path daemon 241 identifies the next packet processing stage 354 corresponding to LS1.

[0068] Processing stage 352 is configured by the DP configuration database 230 to perform L2 switching operations as logical switch LS1, and logical switch LS1 forwards packet 371 to virtual machine VM1(381) according to the destination MAC address of the packet.

[0069] Packet 372 goes to VM 382 attached to the L2 segment defined by logical switch LS2. Packet 372 causes data path daemon 241 to identify packet processing stage 361 to perform service router SR1, then identify packet processing stage 362 to perform logical switch TLS, and then identify packet processing stage 363 to perform distributed router DR. When processing packet 372, packet processing stage 363 routes the packet from the L2 side defined by logical switch TLS to the L2 segment defined by logical switch LS2. Correspondingly, data path daemon 241 identifies the next packet processing stage 364 corresponding to logical switch LS2, and logical switch LS2 forwards packet 372 to virtual machine VM2(382) according to the destination MAC address of the packet.

[0070] In Figure 3 the example, although data path daemon 241 operates according to the same DP configuration database 230, two different packets 371 and 372 cause data path daemon to traverse logical network 300 differently and perform different packet processing stages (for packet 371, SR1-TLS-DR1-LS1, for packet 372, SR1-TLS-DR1-LS2).

[0071] In some embodiments, each packet processing stage is implemented as a function call to a data path daemon thread. In some embodiments, a function (also known as a subroutine or procedure) is a sequence of program instructions encapsulated as an element to perform a specific task. In some embodiments, the functions called to implement the various stages of the data path are part of the programming of the data path daemon operating at the core, but the functions called perform different operations based on different configuration data for different network identities. In other words, the programming of the core provides functions that can be called by the data path daemon to perform the functions of various logical routers, logical switches, and service providing entities.

[0072] The function call uses the content of the packet as input arguments. In some embodiments, the function call also uses the identity of the logical port through which the packet enters the corresponding logical entity as input arguments. In some embodiments, the function call also identifies the exit port, and the exit port is used to identify the entry port for the next function call for the next pipeline stage.

[0073] Figure 4The conceptual diagram illustrates a data path daemon 400 for each stage of a data path pipeline that is executed as a function call. As shown, the data path daemon is processing packet 470 according to a logical network. For each stage, the data path daemon executes a function call that executes a set of instructions corresponding to the operations of a packet processing stage (e.g., logical router, logical switch, service router, etc.). The function call operates on packet 470 based on a set of configuration data (i.e., rule table, routing table, forwarding table, etc.) for that function / stage. The result of the function call is used by data path daemon 400 to identify the next hop and execute the next function call as the next stage. Each function call uses the packet as an input argument, along with other information such as the identity of a logical entity or the logical port to which the packet is forwarded.

[0074] For some embodiments, Figure 5 The figure illustrates an example DP configuration database 500 that provides configuration data for each data path pipeline stage, and a connection map that allows the data path daemon to identify the next pipeline stage between different logical entities. In some embodiments, such a connection map specifies the connection destination for each exit port of each logical entity, whether it is an entry port of another logical entity or a connection leaving the gateway (such as an uplink to an external network). Thus, such a connection map effectively provides the logical topology of the network for some embodiments.

[0075] In some embodiments, each logical port of each logical entity is associated with a Universally Unique Identifier (UUID) such that the logical port can be uniquely identified by the gateway. The UUID of the logical port also allows the data path daemon to identify the logical entity to which the logical port belongs, which in turn allows the data path daemon to identify the configuration data of the identified logical entity and execute the corresponding pipeline stage.

[0076] As shown, the DP configuration database 500 specifies several logical entities 501 - 503 (Logical Entities A, B, C) and their corresponding logical ports. Each logical port is associated with a UUID. For example, logical entity 501 has logical ports with UUIDs: UUID1, UUID2, UUID3, and UUID4, while logical entity 502 has logical ports with UUIDs: UUID5, UUID6, and UUID7. The DP configuration database also specifies the connections for each logical port. For example, logical port UUID4 is connected to logical port UUID5, logical port UUID6 is connected to logical port UUID8, logical port UUID7 is connected to logical port UUID9, etc.

[0077] The DP configuration database 500 also includes configuration data for each logical entity, which includes the ingress and egress ports of the logical entity and its routing table or forwarding table, services provided or enabled on the logical entity, and so on. In some embodiments, such configuration data includes other data to be consumed by the data path during packet processing, such as the MAC-to-VTEP mapping table.

[0078] Figure 6 The conceptual diagram illustrates a process 600 performed by a processing core when executing a data path pipeline using the DP configuration database. The process begins when it receives (at 610) an incoming packet. The process then identifies (620) an initial connection for the received packet. In some embodiments, such identification is based on the content of the packet, such as the header fields of the packet. In some embodiments, such connection is associated with a UUID, and thus it will have a connection mapping according to the DP configuration database.

[0079] The process then determines (at 630) whether the identified connection is an ingress port of a logical entity or an egress port of a gateway. In some embodiments, the process checks the connection mapping provided by the DP configuration database to identify the destination of the connection. If the connection is an ingress port to another logical entity within the gateway, the process proceeds to 640. If the connection is an egress port of the gateway, the process proceeds to 690.

[0080] At 640, the process identifies the logical entity based on the connection. At this stage, the process has determined that the connection connects to the ingress port of a logical entity. By using the DP configuration database, the process is able to identify the logical entity to which the connection is the ingress port. For example, in Figure 5 According to the DP configuration database 500, the logical port UUID5 is the ingress port of the logical entity B (501).

[0081] The process then performs (at 650) the operations of the logical entity. These operations are specified by the configuration data associated with the identified logical entity (e.g., routing table, forwarding table, service rules, etc.). The results of these operations are also based on the content of the packet, such as the source and destination addresses of the packet. The operation corresponds to a function call in some embodiments, which executes a series of instructions by referring to the configuration data in the DP configuration database.

[0082] The process then identifies (at 660) the exit port of the logical entity. For a logical entity that is a logical switch or a logical router, in some embodiments, such exit port identification corresponds to looking up a forwarding table or a routing table to perform routing based on the identity of the ingress port or the content of the packet. In some embodiments, the identification of the exit port (and thus the next processing stage or next hop) is based on some of the following parameters in the packet: (1) the source and destination MAC addresses of the packet for L2 switching / forwarding; (2) the source and destination IP addresses of the packet for L3 routing; (3) the source and destination ports of the packet for L4 transport / connection / flow processing; (4) the identity of the logical network or tenant to which the packet belongs. Configuration data from the DP configuration database provides rules that specify the actions to be taken at each stage based on these packet parameters.

[0083] Next, the process identifies (at 670) the connection of the exit port based on the connection mapping provided by the DP configuration database. For example, in Figure 5 , according to the DP configuration database 500, the logical port UUID4 is connected to the logical port UUID5. The process then returns to 630 to determine whether the exit port is the egress port connected to the gateway or is connected to another logical entity.

[0084] At 690, the process ends the data path pipeline and sends the packet to its destination physical router or host machine. The process then ends. In some embodiments, the gateway communicates with other host machines that are VTEPs of an overlay network (e.g., VXLAN) in the data center, and the next hop is behind another VTEP in the data center. In such a case, the process encapsulates the packet according to the overlay network and sends the encapsulated packet to the destination VTEP. If the next hop is a physical router in an external network (which is typically a VLAN), the gateway will remove the overlay encapsulation and bridge the packet to the physical router. In some embodiments, the DP configuration provides a mapping that maps the destination MAC address of the packet to the corresponding VTEP address.

[0085] a. Centralized and Distributed Pipeline Stages

[0086] In some embodiments, some logical entities / entities / components of a logical network distributed among multiple physical machines in a data center, i.e., each of these host machines has a copy or instance of the distributed logical entity. Packets that need to be processed by the distributed logical entity can be processed by any machine running an instance of the distributed logical entity. On the other hand, some of the logical entities / entities are not distributed but centralized or concentrated on one physical machine, i.e., the logical entity has only one physical instance. In some embodiments, such a centralized router acts as a centralized point for routing packets between the logical network and an external router. Packets that need to be processed by the centralized logical entity must be forwarded to the machine operating the centralized logical entity. The distributed logical router and the centralized logical router are described in U.S. Patent Application No. 14 / 814,473, titled "Logical Router with Multiple Routing Components", filed on June 30, 2015.

[0087] In some embodiments, when processing incoming packets, the data path daemon will execute both the distributed logical entity and the centralized logical entity as part of its pipeline stages. In some embodiments, the service router is a centralized logical router. Each service router has only one instance running on one gateway machine. Therefore, the data path daemon running on the gateway machine will invoke the centralized or gateway machine-concentrated service router as one of its data path pipeline stages.

[0088] In some embodiments, the configuration data (DP configuration database) that controls the operation of the service router stage includes the definition of any services that the logical router should provide, whether the logical router will be configured in an active-active mode or an active-standby mode, how many uplinks are configured for the logical router, the IP and MAC addresses of the uplinks, the L2 and L3 connections of the uplinks, the subnets of any southbound interfaces of the logical router, and any static routes and other data for the routing information base (RIB) of the logical router.

[0089] Figure 7 Illustrated is a logical network with both distributed and centralized logical entities. Specifically, the figure illustrates the logical view and the physical view of the logical network 300. The logical view of the logical network 300 shows the logical relationships and connections between the various logical entities of the network. The physical view of the logical network 300 shows the physical instantiation of the various logical entities in the physical host machines and the physical gateways of the data center.

[0090] According to the logical view, the logical network 300 includes service routers 311 and 312 (SR1 and SR2), logical switches 321, 322, and 323 (LS1, LS2, and TLS), and a distributed router 313 (DR). Among these logical entities, service routers SR1 and SR2 are centralized logical entities, while LS1, LS2, TLS, and DR are distributed logical entities.

[0091] Some embodiments provide a distributed logical router implementation that enables first-hop routing in a distributed manner (instead of centralizing all routing functions at the gateway). In a physical implementation, the logical router of some embodiments includes a single distributed routing component (also referred to as a distributed router or DR) and one or more service routing components (also referred to as service routers or SRs). In some embodiments, the DR spans managed forwarding elements (MFEs) that are directly coupled to virtual machines (VMs) or other data computing nodes that are directly or indirectly logically connected to the logical router. The DR of some embodiments also spans the gateways to which the logical router is bound and one or more physical routers capable of performing routing operations. The DR of some embodiments is responsible for first-hop distributed routing between logical switches and / or other logical routers that are logically connected to the logical router. Service routers (SRs) only span the edge nodes of the logical network and are responsible for delivering services that are not implemented in a distributed manner (e.g., some stateful services).

[0092] The physical view of the network shows the physical instantiation of these centralized and distributed logical entities in the actual physical machines of the data center 100. As shown, the data center 100 includes gateways 111 - 112 and host machines 121 - 123 interconnected by physical connections. Instances of the distributed elements TLS, LS1, LS2, and DR are distributed across the gateways 111 - 112 and the host machines 121 - 123. In some embodiments, different physical instances of the distributed elements operate according to the same set of forwarding tables. However, the centralized element SR1 is only active in gateway 111, and the centralized element SR2 is only active in gateway 112. In other words, only the data path daemon of gateway 111 executes SR1 as a pipeline stage and only the data path daemon of gateway 112 executes SR2 as a pipeline stage.

[0093] Figure 8Illustrated is a gateway data path daemon that executes pipeline stages for incoming packets (also known as southbound traffic) to the logical network of a data center from an external network. As shown, packets received from the external network via uplink U1 are processed by gateway 111, and the data path daemon 141 of gateway 111 executes pipeline stages corresponding to SR1, TLS, DR, and LS-A (or LS-B, depending on the destination L2 segment of the packet). The processed packets are then sent to one of the host machines for forwarding to the destination VM. On the other hand, packets received from the external network via uplink U2 are processed by gateway 112, and the data path daemon 142 of gateway 112 executes pipeline stages corresponding to SR2, TLS, DR, and LS-A (or LS-B, depending on the destination address of the packet). The processed packets are then sent to one of the host machines for forwarding to the destination VM. In some embodiments, the logical switching stage (LS-A or LS-B) of the data path daemon identifies the host machine of the destination VM.

[0094] Both gateways 111 and 112 execute pipeline stages corresponding to the distributed transit logical switch TLS, the distributed router DR, and the logical switches LS-A and LS-B because these are distributed logical network constructs. However, only the data path daemon of gateway 111 executes the pipeline stage for service router SR1 because SR1 is a centralized router located at gateway 111, and only the data path daemon of gateway 112 executes the pipeline stage for service router SR2 because SR2 is a centralized router located at gateway 112.

[0095] Figure 9 Illustrated is a gateway data path daemon that executes pipeline stages for outgoing packets (also known as northbound traffic) from the logical network of a data center to an external network. As shown, packets originating from a VM operating in a host machine undergo pipeline stages corresponding to LS-A (or LS-B, depending on the L2 segment of the source VM), DR, and TLS. The TLS stage of the host machine identifies the next-hop router, which is either SR1 in gateway 111 or SR2 in gateway 112. In some embodiments, the selection of the service router is based on the destination address of the packet and routing decisions made previously in the pipeline.

[0096] For packets sent to gateway 111, the data path daemon 141 of gateway 111 executes pipeline stages corresponding to TLS and SR1 before relaying the packets to the external network via uplink U1. For packets sent to gateway 112, the data path daemon 142 of gateway 112 executes pipeline stages corresponding to TLS and SR2 before relaying the packets to the external network via uplink U2.

[0097] b. Data path pipelines for different tenants

[0098] In some embodiments, the data center supports multiple logical networks for multiple different tenants. Different tenant logical networks share the same set of gateway machines, and each of these gateway machines provides packet switching, forwarding, and routing operations for all connected tenant logical networks. In some embodiments, the data path daemon is capable of performing a packet processing stage on packets going to and coming from different logical networks belonging to different tenants. In some of these embodiments, the DP configuration database provides configuration data (i.e., routing tables, forwarding tables, etc.) and service specifications that enable tenant-specific packet forwarding operations at the gateway.

[0099] Different tenant logical networks have different packet processing logic entities, e.g., different logical routers and logical switches. Figure 10 FIG. 1000 illustrates a logical view of the entire network of the data center. The entire network includes various logical entities belonging to different tenant logical networks and a set of logical entities provided by the data center. These logical entities provided by the data center are shared by all tenants to access the external network through the gateway and use the services provided by the gateway.

[0100] According to FIG. 1000, the entire network of the data center includes a provider logical router (PLR) 1090 and tenant logical routers (TLR) 1010, 1020, and 1030 (TLR1, TLR2, and TLR3). The PLR 1090 is connected to TLR1 1010 through a transit logical router 1019 (TLS1), to TLR2 1020 through a transit logical router 1029 (TLS2), and to TLR3 1030 through a transit logical router 1039 (TLS3). TLR1 is used to perform L3 traffic for tenant 1, TLR2 is used to perform L3 traffic for tenant 2, and TLR3 is used to perform L3 traffic for tenant 3. The logical networks of different tenants are connected together by the PLR 1090. The PLR 1090 serves as an intermediate logical router between each tenant logical network and the external physical network.

[0101] In some embodiments, the logical router is part of a two-tier logical network structure. The two-tier structure of some embodiments includes (1) a single logical router (referred to as a provider logical router (PLR) and managed by, for example, the owner of the data center) for connecting the logical network to the network external to the data center; and (2) multiple logical routers (each referred to as a tenant logical router (TLR) and managed by, for example, different tenants of the data center) that are connected to the PLR and do not communicate separately with the external network. In some embodiments, the control plane defines the transit logical exchange between the distributed components of the PLR and the service components of the TLR.

[0102] For the PLR logical router, some embodiments use the active-active mode as much as possible and use the active-standby mode only when a stateful service (e.g., NAT, firewall, load balancer, etc.) is configured for the PLR. In the active-standby mode, only one service routing component is active, i.e., fully operational at a given time, and only this active routing component sends messages to attract traffic. All other service routing components are in standby mode. In some embodiments, the active service component and the standby service component use the same IP address but different MAC addresses to communicate with the distributed components. However, only the active component answers the Address Resolution Protocol (ARP) requests from the distributed component. In addition, only the active service component advertises routes to the external network to attract traffic.

[0103] For the TLR logical router, when a stateful service is configured for the TLR, some embodiments either do not use service components or use two service components in the active-standby mode. The TLR operates internally in the same manner as the PLR in the active-standby mode, i.e., having an active component and a standby component that share the same network layer address, but only the active component responds to ARP requests. To connect to the PLR, some embodiments assign the same network layer address to each of the two service components of the TLR (although different from the IP address used to connect to its own distributed component).

[0104] The above-described logical router is a distributed logical router implemented by a single distributed routing component and a set of service routing components. Some embodiments provide other types of logical router implementations in a physical network such as a centralized logical router (e.g., a data center network). In the centralized logical router, the L3 logical routing function is performed only in the gateway machine, and the control plane of some embodiments does not define any distributed routing components but only defines a plurality of service routing components, where each service routing component is implemented in a separate gateway machine.

[0105] Different types of logical routers having multiple routing components (e.g., distributed logical routers, multi-layer logical routers, etc.) and implementing different types of logical routers on edge nodes and managed forwarding elements operating on host machines in a data center are described in more detail in U.S. Patent Application 14 / 814,473, filed on July 30, 2015, which is incorporated herein by reference.

[0106] The PLR includes service routers 1001 - 1003 (SR1, SR2, and SR3) that provide access to the physical network and edge services. The PLR also includes a distributed router 1005 (PLR - DR) for routing packets to and from different tenant logical networks. The PLR distributed router 1005 is connected to service routers SR1, SR2, and SR3 via a transit logical router (PLR - TLS) 1099.

[0107] Each TLR serves as an L3 hub for a tenant logical network. Each TLR includes a distributed router (DR) for connecting different L2 segments defined by different logical switches. Specifically, TLR1 includes a TLR1 - DR (1015) for connecting logical switches LS - A and LS - B (1011 and 1012), TLR2 includes a TLR2 - DR (1025) for connecting logical switches LS - C and LS - D (1021 and 1022), and TLR3 includes a TLR3 - DR (1035) for connecting logical switches LS - E and LS - F (1031 and 1032).

[0108] In some embodiments, the DP configuration database stores routing tables, forwarding tables, rule tables, etc. for different logical entities as configuration data. The DP configuration database provides a mapping between connection identifiers (ingress port and egress port) and network logical entity identifiers. The data path daemon, in turn, executes data path pipelines for different tenants through function calls and following the connection mappings between different logical entities, where some of the logical entities correspond to various tenant - specific logical network constructs (e.g., TLR - LS or TLR - DR for different tenants). The data path daemon provides common network services to all tenants by executing pipeline stages corresponding to various provider logical entities (e.g., SR and PLR - DR).

[0109] Figure 11 Illustrated is a data path daemon 1105 that performs gateway packet processing for different tenants at a gateway machine 1100. The data path daemon 1105 is a processing thread operating on a processor core 1110 of the gateway machine 1100. It operates independently of the DP configuration database 1130. The DP configuration database 1130 stores configuration data (such as routing tables, forwarding tables, and rule tables) of various logical entities of the data center as shown in the network diagram 1000.

[0110] As shown, the DP configuration database 1130 includes configuration data for each logical entity / entities of the network (logical routers and logical switches 1001 - 1099), which includes tenant - specific entities (e.g., TLR) and provider entities (e.g., PLR) shared by all tenants. Figure 12Illustrated is the packet data path daemon 1105 that processes packets by invoking pipeline stages corresponding to various logical entities.

[0111] Packets 1211 - 1216 are southbound packets that enter the data center from an external network via the uplink of gateway 1100. Packets 1211 - 1216 go to VMs belonging to different tenants: Packets 1211 and 1212 go to the logical network of tenant 1, packets 1213 and 1214 go to the logical network of tenant 2, and packets 1215 and 1216 go to the logical network of tenant 3. Since packets 1211 - 1216 come from an external network, they are unencapsulated VLAN packets.

[0112] Packets for different tenants have different destination IP or MAC addresses, and the packet data path daemon accordingly identifies and executes different pipeline stages (e.g., function calls to different network logical entities) corresponding to different tenants. The packet data path daemon initially invokes the PLR stages PLR - SR1, PLR - TLS, PLR - DR, which route the packets to their corresponding TLSs based on the destination addresses of the packets. These TLSs in turn switch the packets to their corresponding tenant - specific TLRs.

[0113] For example, packet 1211 is a tenant 1 packet going to a VM behind logical switch LS - A. The packet data path daemon 1105 thus executes the pipeline stages corresponding to the following logical entities: PLR - SR1, PLR - TLS, PLR - DR, TLS1, TLR1 - DR, and LS - A. Packet 1214 is a tenant 2 packet going to a VM behind logical switch LS - D. The packet data path daemon 1105 accordingly executes the pipeline stages PLR - SR1, PLR - TLS, PLR - DR, TLS2, TLR2 - DR, and LS - D. Packet 1215 is a tenant 3 packet going to a VM behind logical switch LS - E. The packet data path daemon 1105 accordingly executes the pipeline stages PLR - SR1, PLR - TLS, PLR - DR, TLS3, TLR3 - DR, and LS - E.

[0114] Among these logical entities, PLR-SR1, PLR-TLS, and PLR-DR are provider constructs common to all tenants. TLS1, TLR1-DR, and LS-A are tenant-specific constructs for Tenant 1. TLS2, TLR2-DR, and LS-D are tenant-specific constructs for Tenant 2. TLS3, TLR2-DR, and LS-E are tenant-specific constructs for Tenant 3. Each of these phases has corresponding configuration data provided by the DP configuration database for routing packets, identifying the next hop, providing services, etc. In some embodiments, tenant-specific logical network constructs use tenant-specific forwarding tables, routing tables, rule tables, and other tenant-specific configuration data.

[0115] Since the destinations of packets 1211 - 1216 are VMs elsewhere in the data center, the gateway tunnels these packets to their corresponding destination host machines via an encapsulation overlay network. Specifically, packets 1211 - 1216 are encapsulated according to their corresponding tenant logical network and sent as encapsulated packets 1221 - 1226.

[0116] Packets 1231 - 1236 are northbound packets exiting the data center to the external network via gateway 1100. These packets 1231 - 1236 are encapsulated under the provider overlay because they have been routed to PLR-TLS at their corresponding source host machines. They are tunneled to the gateway via the provider overlay encapsulation tunnel, and gateway 1100 invokes PLR-TLS and PLR-SR1 to provide the necessary services before sending them as VLAN packets to the external network via the uplink.

[0117] Although not illustrated, in some embodiments, packets of different tenants are encapsulated differently for different overlay networks, and the data path daemon uses tenant-specific information in the encapsulation to identify and execute different pipeline phases corresponding to different tenants.

[0118] c. Service provisioning pipeline phase

[0119] In some embodiments, in addition to performing the L3 routing and L2 routing pipeline phases, the gateway data path daemon also performs a service provisioning phase for L4 to L7 processing. These services support end-to-end communication between the source application and the destination application and are used whenever a message is passed from or to the user. The data path daemon applies these services to the packets at the vantage point of the edge gateway without changing the applications running at the source or destination of the packet. In some embodiments, the data path can include service phases for traffic filtering services (such as firewalls), address mapping services (such as NAT), and encryption and security services (such as IPSec and HTTPS).

[0120] In some embodiments, some or all of these service provisioning phases are performed while the data path daemon is executing the service router pipeline stage. Additionally, in some embodiments, the data path daemon may execute different service provisioning pipeline stages for different packets. In some embodiments, the data path daemon executes different service provisioning stages based on the L4 flow to which the packet belongs and the state of the flow. In some embodiments, the data path daemon executes different service provisioning stages based on the tenant to which the packet belongs.

[0121] Figure 13 Illustrated is a gateway data path daemon that executes service provisioning pipeline stages for different tenants. The data path daemon 1305 is a processing thread operating on the processor core 1310 of the gateway machine 1300. It operates independent of the DP configuration database 1330, which provides configuration data and connection mappings for various data path pipeline stages. Some of these pipeline stages are service provisioning stages of services such as firewall, NAT, and HTTPS. The data path daemon decides which service stage to execute based on the configuration data of the logical router, the result of L3 routing, and / or the content of the packet, which may indicate which L4 flow and / or which tenant the packet belongs to.

[0122] The data path daemon executes these service provisioning stages after centralized routing (1321) and before the transit logical switch (1322), distributed router (1323), and logical switch (1324) stages of the pipeline. In some embodiments, these service provisioning stages are considered part of the service router (SR) pipeline stage. In some embodiments, some of the service provisioning stages are for providing stateful services and are thus centralized or centralized logical entities operating at one gateway machine. In some embodiments, the L4 service stage provides a stateful service by maintaining the state of each L4 connection.

[0123] This figure illustrates the data daemon 1305 executing different packet service provisioning stages for different packets 1371 - 1373. These packets may belong to different tenants or different L4 flows, or to the same L4 flow in different states. As shown, when processing packet 1371, the data path daemon executes service stages 1311, 1312, and 1314 that provide firewall, NAT, IPSec, and HTTPS services respectively. When processing packet 1372, the data path daemon only executes the firewall service stage (1311). When processing packet 1373, the data path executes the NAT and HTTPS service stages (1312 and 1314).

[0124] Figure 14The conceptual diagram illustrates process 1400 for providing services within a logical router (e.g., a service router). In some embodiments, the core executing the data path daemon executes process 1400 when it executes a function call for a pipeline stage corresponding to the service router. The process begins when it receives (at 1400) a packet and the identity of the logical port that is the ingress port. As discussed above with reference to Figure 5 and Figure 6 , in some embodiments, the DP configuration database provides the necessary mapping that allows the data path daemon to identify the corresponding logical entity when provided with the identity of the logical port. The process then accesses (at 1420) the configuration data of the identified logical entity. As mentioned, the configuration data of a logical entity such as a service router can include its routing table and the specifications of the services to be provided by the service router.

[0125] The process then performs (at 1430) the routing of the packet (since this is the service router stage). In some embodiments, this routing is based on the source or destination address of the packet or the identity of the ingress port. The process then identifies (at 1440) the network service according to the configuration data of the logical entity. In some embodiments, the service router can belong to different tenant logical networks, which can have different policies and require different services. Thus, the DP configuration database will specify different services for different tenant logical routers, and the service routers of these different tenant logical routers will perform different services.

[0126] The process then performs (at 1450) the operations specified by the identified service (e.g., NAT, firewall, HTTPS, etc.). In some embodiments, these operations are also based on the current content of the packet (e.g., the destination IP address), which may have been changed by a previous service executed by the process.

[0127] At 1460, the process determines whether the DP configuration database has specified another service for the service router. If so, the process returns to 1440 to perform another service. Otherwise, the process proceeds to 1470 to identify the egress port and output the packet to the next hop. Process 1400 then ends.

[0128] II. Cache for Data Path Pipeline

[0129] In some embodiments, instead of performing the pipelining stage on all packets, the gateway caches the results of previous packet operations and reapplies the results to subsequent packets that meet certain criteria (i.e., cache hit). For packets that do not have applicable or valid results from previous packet processing operations, i.e., cache miss, the gateway data path daemon performs the pipelined packet processing stage. In some embodiments, when the data path daemon performs the pipelining stage to process a packet, it records a set of data from each stage of the pipeline and synthesizes this data into a cache entry for subsequent packets.

[0130] In some embodiments, each cache entry corresponds to an L4 flow / connection (e.g., a five-tuple with the same source IP, destination IP, source port, destination port, and transport protocol). In other words, the data path daemon determines whether a packet has an applicable cache entry by identifying the packet flow. Thus, in some of these embodiments, the data path cache is also referred to as a flow cache.

[0131] Figure 15a -b illustrates a data path daemon that maintains a cache to accelerate packet processing. As shown, the data path daemon 1510 running on the core of the processor is processing packet 1570 from NIC 1590. The data path daemon 1510 is a processor thread that can process packet 1570 by executing the stages of the data path pipeline 1520 as described in Part I or by applying an entry from the data path cache 1530. The data path daemon uses the configuration data stored in the data path configuration database 1540 to configure and execute its pipelining stages.

[0132] Figure 15a Illustrated is packet processing when the incoming packet 1570 is a cache hit. As shown, the data path daemon is able to find a valid matching entry (e.g., with the same flow identifier) for the incoming packet 1570 in the data path cache 1530. The daemon 1510 then uses the matching entry in the cache to directly specify the action to be taken regarding the packet, e.g., specifying the next hop for the packet, resolving the IP address, rejecting the packet, translating the IP address in the packet header, encrypting / decrypting the packet, etc. No pipeline stages are executed (i.e., the data path daemon does not execute any pipeline stages).

[0133] Figure 15bThe figure illustrates packet processing when the incoming packet 1570 is a cache miss. As shown, the packet 1570 does not have a valid matching entry in the data path cache 1530 for the packet 1570. The data path daemon 1510 thus executes the stages of the data path pipeline (i.e., as described in Part I above, by making function calls and applying the configuration data of the logical entities). Since the data path executes the stages of the data path pipeline, each stage of the data path pipeline generates a set of information for synthesizing cache entries in the data path cache. This new cache entry (or updated cache entry) will apply to subsequent packets belonging to the same packet category as the packet 1570 (e.g., belonging to the same L4 flow).

[0134] Figure 16 The figure illustrates the synthesis of cache entries in the data path cache. As shown, the packet 1570 has caused a cache miss and the data path daemon 1510 is executing the stages of the data path pipeline 1520. When the data path pipeline is being executed, some or all of the executed stages emit data or instructions that will be used by the synthesizer 1610 to synthesize the cache entry 1620. In some embodiments, the cache entry synthesis instructions or data emitted by the pipeline stages include the following: a cache enable field 1631, a bitmask field 1632, and an action field 1633.

[0135] The cache enable field 1631 specifies whether to create a cache entry. In some embodiments, the pipeline stage may determine that the result of packet processing should not be used as a cache entry for future packets, i.e., only the packet 1570 should be processed in this way and future packets should not reuse the processing result of the packet 1570. In some embodiments, even if all other pipeline stages indicate that enabling the creation of a cache entry is okay, one pipeline stage that specifies that a cache entry should not be created will prevent the synthesizer 1610 from creating a cache entry.

[0136] The bitmask field 1632 defines which part of the packet header the pipeline stage actually uses to determine the action to take regarding the packet. Some embodiments apply the bitmask only to fields in the inner headers (IP header and MAC header) rather than the outer headers (i.e., headers of overlay encapsulations such as VXLAN). In some embodiments where the cache entry is flow-based, the bitmask field 1632 is used to create a cache entry applicable to multiple flows, i.e., by setting certain bit fields in the inner header to "don't care".

[0137] The action field 1633 specifies what action the pipeline stage has taken regarding the packet.

[0138] The synthesizer 1620 collects all cache entry synthesis instructions from all pipeline stages and synthesizes the entries 1620 in the flow cache from all the received instructions (unless one or more pipeline stages specify that cache entries should not be generated). The synthesized cache entries specify the final actions for packets that meet certain criteria (i.e., belong to certain L4 flows). When generating the cache entries 1620, the synthesizer also includes a timestamp specifying the time at which the cache entry was created in some embodiments. This timestamp will be used to determine whether the cache entry is valid for subsequent packets.

[0139] In some embodiments, the synthesizer 1610 creates aggregate cache entries that apply to multiple L4 flows or "jumbo flows". In some embodiments, these are entries whose match criteria have certain bits or fields masked (i.e., are considered "don't cares"). In some embodiments, the synthesizer creates jumbo flow entries based on the bitmask fields 1633 received from the executed pipeline stages. The synthesizer 1610 also creates exact match entries whose match criteria are fully specified to apply to only one flow or "microflow".

[0140] Figure 17 An example synthesis of an aggregate cache entry and an exact match cache entry is illustrated. The data path 1520 processes the packet 1700 through its stages 1521 - 1523, and each stage generates a set of cache synthesis instructions for the cache entry synthesizer 1610. The cache entry synthesizer 1610 in turn creates an aggregate cache entry 1751 and an exact match cache entry 1752 for the data path cache 1530.

[0141] As shown, the exact match entry 1752 fully specifies all fields as its match criteria. These fields exactly match the fields of the packet 1700 (e.g., the 5 - tuple flow identifier in its header). On the other hand, the aggregate entry 1751 only specifies some of its fields in its match criteria while masking some other fields. The packet 1700 will match these match criteria, but other packets that may have different values in these corresponding fields will also likely match these match criteria. In some embodiments, which fields / bits are masked in the match criteria of the aggregate entry are determined by the bitmask fields (e.g., 1632) generated by the respective data path stages.

[0142] Each cache entry also specifies the final actions to be taken on each packet that matches the cache entry. In some embodiments, these actions include all actions that affect the packet when it is output. In Figure 17In the example of, stages 1521, 1522, and 1523 respectively specify Action (1), Action (2), and Action (3). Action (1) and (3) affect the packet, but do not affect Action (2), so only Action (1) and (3) become part of the composite cache entries 1751 and 1752. For example, an action to update a register to indicate a packet processing stage does not affect the output packet and is thus not included in the flow entries of the cache, while an action to modify a header value (e.g., modify the MAC address as part of an L3 routing operation) is included. If the first action modifies the MAC address from a first value to a second value, and a subsequent action modifies the MAC address from the second value to a third value, then some embodiments specify directly modifying the MAC address to the third value in the flow entry of the cache.

[0143] In some embodiments, an action specified by one stage will override all other stages. Figure 18 Illustrates an example of a data path stage that specifies an action that overrides all other actions. The figure illustrates two example packets 1871 and 1872 processed by the data path 1520.

[0144] As shown, when the data path 1520 processes packets 1871 and 1872, each of its stages specifies certain actions. For packet 1871, the specified action includes "reject packet" by stage 1523. This action will override all other actions, and the cache entry created by this packet 1871 will only perform the action "reject packet". For packet 1872, stage 1522 specifies disabling the cache (cache enable = 0). As described above, in some embodiments, each stage of the data path can (e.g., via its cache enable bit) specify not to create a cache entry for a given packet, regardless of what other stages in the data path have specified. Thus, the cache entry synthesizer 1610 (not shown) does not create a cache entry for packet 1872.

[0145] Figure 19 Illustrates checking the entries of the data path cache to determine if there is a cache miss or a cache hit. As shown, the data path daemon (at the match function 1910) compares certain fields of the packet 1570 (e.g., the flow identification field) with the entries in the cache 1530 to find a cache entry applicable to the packet 1570. If the match function 1910 cannot find a matching cache entry, the data path daemon will proceed as a cache miss.

[0146] As shown, each entry is also associated with a timestamp, thereby marking the time when the cache entry (created by synthesizer 1610) was created and stored into the data cache. Some embodiments (by comparison function 1920) compare the timestamp of the matching cache entry with the timestamp of the DP configuration database 1540 to determine whether the cache entry is still valid. (In some embodiments, this timestamp records the time when the data in the database was last updated by the network controller or manager. The update of the DP configuration database will be described in Section III below). Specifically, if the DP configuration database 1540 has not been changed since the cache entry was created, i.e., the timestamp of the DP configuration database is before the timestamp of the matching entry, then the cache entry is still valid, and the data path daemon will proceed as a cache hit. Conversely, if the DP configuration database 1540 has been changed since the cache entry was created, i.e., the timestamp of the DP configuration database is after the timestamp of the matching entry, then the cache entry is considered no longer valid, and the data path daemon will proceed as a cache miss.

[0147] Figure 20 The conceptual diagram illustrates process 2000 for operating the data path cache. In some embodiments, a processor core operating as a data path daemon thread executes process 2000. Process 2000 begins when it receives (at 2010) a packet either from an external physical network or from a data center. The process then determines (at 2020) whether the packet has a matching entry in the data path cache. If so, the process proceeds to 2025. If the packet does not have a matching entry in the data path cache, the process proceeds to 2030.

[0148] At 2025, the process determines whether the matching cache entry is still valid, e.g., whether its timestamp indicates that the cache entry was made after the most recent update to the DP configuration database. The determination of cache entry validity is described by reference to the above Figure 19 If the matching cache entry is valid, the process proceeds to 2060. Otherwise, the process proceeds to 2030.

[0149] At 2030, the process indicates that the packet has caused a cache miss and starts the data path pipeline by executing its stages. The process then synthesizes (at 2040) a cache entry based on the data or instructions generated by the stages of the data path pipeline. The synthesis of the cache entry is described by reference to the above Figure 16 The process then stores (at 2050) the synthesized cache entry and associates the entry with the current timestamp. Process 2000 then ends.

[0150] At 2060, the process indicates a cache hit and performs an action based on the matching cache entry. Process 2000 then ends.

[0151] III. Data Path Configuration Update

[0152] As described above, the pipeline stages of the data path daemon use the configuration data in the DP configuration database as the forwarding table, routing table, rules table, etc. Since these tables contain real-time information about what actions should be taken regarding packets at the gateway, some embodiments dynamically update the DP configuration database even while the data path daemon is actively accessing the DP configuration database. To ensure that the data path daemon does not use incomplete (and thus corrupted) configuration data in its pipeline stages when the DP configuration database is being updated, some embodiments maintain two copies of the DP configuration database. One copy of the database is used as a staging area for new updates from the network controller / manager, so that the data path daemon can safely use the other copy of the database. Once the update is complete, the roles of the two database copies are atomically reversed. In some embodiments, the network controller waits for the data path daemon to complete its current run through the packet processing pipeline stage before switching.

[0153] Figure 21 Illustrated is a gateway 2100 with a DP configuration database 2110 that supports updates atomically. As shown, the gateway 2100 has a set of processor cores 2121 - 2123, each core operating a data path daemon that uses the configuration data in the DP configuration database 2110 to perform pipeline stages corresponding to logical entities. The network controller / manager 2190 dynamically updates the configuration data stored in the DP configuration database while the kernel is actively using the database.

[0154] As shown, the DP configuration database 2110 has two copies: an odd copy 2111 and an even copy 2112 ("DP Configuration Odd" and "DP Configuration Even"). Each copy of the database stores the complete configuration data for operating the data path daemon at cores 2121 - 2123. In some embodiments, the two different copies are stored in two different physical storage devices. In some embodiments, the two copies are stored in different locations of the same storage device.

[0155] When updating the DP configuration database 2110, the network controller 2190 uses the write pointer 2195 to select whether to write to the odd or even copy. When the pipeline stage is executed, cores 2121 - 2123 respectively use the read pointers 2131 - 2133 to select whether to read the odd or even copy. The network controller 2190 selects and updates one copy of the DP configuration database, while cores 2121 - 2123 each select and use the other copy of the DP configuration database.

[0156] Figure 22a -Figure b illustrates an atomic update of the data path configuration database 2110. The figure illustrates the update process in six stages 2201 - 2206.

[0157] At the first stage 2201, all read pointers 2131 - 2133 point to the odd copy 2111, and the write pointer 2195 points to the even copy 2112. In addition, all cores 2121 - 2123 are respectively reading configuration data from the odd copy 2111 and executing the packet processing pipeline for processing packets 2271 - 2273. The network controller 2190 is writing to the even copy 2112. The data in the even copy 2112 is thus incomplete or corrupted, but the data path daemons in cores 2121 and 2123 are isolated from this because they are operating off the odd copy 2111.

[0158] At the second stage 2202, the network controller has completed the update to the DP configuration database, that is, it has completed writing to the even copy 2112. The updated database is associated with a timestamp 2252 (as described in section II above for determining cache entry validity). Meanwhile, all cores are still in the middle of their respective run - to - completion pipelines.

[0159] At the third stage 2203, core 2121 has completed its previous run - to - completion pipeline for packet 2271. The read pointer 2131 then switches to the even copy 2112 before core 2121 starts processing another packet. In other words, the data path daemon of core 2121 will use the updated configuration data for its next packet. By using the odd copy of the database, the other two cores 2122 and 2123 are still in their current run - to - completion pipelines.

[0160] At stage 2204, core 2122 has also completed its packet processing pipeline for packet 2272, and its corresponding read pointer has switched to the even copy 2112. Meanwhile, core 2121 has started processing another packet 2274 by using the updated new data in the even copy 2112. Core 2123 is still processing packet 2273 by using the old configuration at the odd copy 2111.

[0161] At stage 2205, core 2123 has also completed its packet processing pipeline for packet 2273 and its corresponding read pointer 2133 has switched to the even copy 2112. Core 2122 has started processing packet 2275. At this time, no data path daemon is processing packets by using the old configuration data in the odd copy 2111. Since there is already a newer updated version of the DP configuration database in the even copy 2112, the old data in the odd copy 2111 is no longer useful. Therefore, some embodiments reset the odd copy 2111 of the DP configuration database to indicate that the data therein is no longer valid, and the network controller is free to write new configuration data to it.

[0162] At the sixth stage, the write pointer 2195 has switched to the odd copy of the database, while cores 2121 - 2123 are processing packets 2274 - 2276 by using the configuration data stored in the even copy 2112. This allows updates to the DP configuration database without affecting the operation of any data path daemon.

[0163] Figure 23 The conceptual diagram illustrates processes 2301 and 2302 for controlling the read and write pointers of the DP configuration database. In some embodiments, the gateway machine executes both processes 2301 and 2302.

[0164] Process 2301 is used to control the write pointer for writing to the DP configuration database by the network controller. Process 2301 starts when the gateway machine receives (at 2310) updated data for the DP configuration database from the network controller / manager. Then the process determines (at 2320) whether the copy of the database pointed to by the write pointer is currently being read by any core running a data path daemon. If so, the process returns to 2320 and waits until the copy of the database is no longer being used by any data path daemon. Otherwise, the process proceeds to 2330. Some embodiments make this determination by checking the read pointers used by the data path daemons: when none of the read pointers currently point to the copy of the database pointed to by the write pointer, the copy of the database pointed to by the write pointer is not being used (and thus it is safe to write to it).

[0165] At 2330, the process updates (at 2330) the configuration data stored in the copy of the database pointed to by the write pointer (e.g., by adding, deleting, modifying table entries, etc.). Once the update is complete, the process flips (at 2340) the write pointer to point to another copy of the database (if it is an odd copy, it flips to point to an even copy, and vice versa). Process 2301 then ends.

[0166] Process 2302 is used to control the read pointer used by the core / data path daemon to read from the DP configuration database. Process 2300 starts when the data path daemon receives (at 2350) a packet to process and starts the data path pipeline.

[0167] The process then reads (at 2360) and applies the configuration data stored in the copy of the database pointed to by the read pointer of the data path daemon. The process also processes (at 2370) the packet through all stages of the pipeline (running to completion). The operation of Process 2301 ensures that the applied configuration data will not be corrupted due to any updates to the data path configuration database.

[0168] Next, the process determines (at 2375) whether other copies of the database have updated configuration data. If another copy of the database does not have a newer version of the configuration data, Process 2302 ends. If another copy of the database does have a newer version of the configuration data, the process flips (at 2380) the read pointer to point to another copy of the database. Process 2302 then ends.

[0169] IV. Software Architecture

[0170] Figure 24 Illustrated is the architecture of a gateway machine 2400 according to some embodiments of the present invention. The memory of the gateway machine is partitioned into user space and kernel space. The kernel space is reserved for running the privileged operating system kernel, kernel extensions, and most device drivers. The user space is the memory area where application software and some drivers execute.

[0171] As shown, the packet processing thread 2410 (i.e., the data path daemon) is operating in the user space to handle L2 switching, L3 routing, and services such as firewalls, NAT, and HTTPS. Other service tasks such as ARP (Address Resolution Protocol) learning, BFD (Bidirectional Forwarding Detection) are considered to run slower and are thus processed by a separate process 2420 in the user space. These slower tasks are not processed by the data path daemon and are not part of the data path pipeline. The packet processing thread 2410 depends on a set of DPDK libraries 2430 for receiving packets from the NIC (by Developed )。In some embodiments, the NIC operation relies on a user - space NIC driver that uses a polling mode to receive packets.

[0172] In kernel space, the operating system kernel 2440 (e.g., Linux) operates the TCP / IP stack and processes the BGP stack (Border Gateway Protocol) for exchanging routing information with an external network. Some embodiments use KNI (Kernel NIC Interface) to allow user - space applications to access the kernel - space stack.

[0173] As described above, the gateway machine in some embodiments is implemented by using a processor with multiple cores, and each data - path daemon executes all of its pipeline stages in an RTC (Run - to - Completion) thread at one core. In some embodiments, the data - path daemon can insert service - pipeline stages executed by a service process executed by another thread at another core.

[0174] In some embodiments, these service processes communicate with the RTC thread using some form of inter - process communication (IPC) (such as shared memory or sockets). The RTC thread receives packets from the NIC, performs normal L2 / L3 forwarding, and classifies the packets to determine whether the packets need service. When a packet needs service, the packet is sent via the IPC channel to the corresponding service process. The IPC service process dequeues and processes the packet. After processing the packet, the service process passes it back to the RTC thread, which continues to process the packet (and may send the packet to another service process for other services). Effectively, the RTC thread is used to provide basic forwarding and direct packets between service processes. Figure 25a The conceptual diagram illustrates an RTC thread that uses IPC to communicate with service processes to provide services.

[0175] In some other embodiments, the service process runs inside a container and does not use IPC to communicate with the RTC thread and is actually unaware of the RTC thread. The process opens standard TCP / UDP sockets to send and receive packets from the Linux kernel. Instead of using IPC to communicate between the service process and the RTC thread, a tun / tap device or a KNI device is created inside the container. The routing table of the container is correctly populated so that packets sent by the service process can be routed using the correct tun / tap / KNI device.

[0176] When the RTC thread determines that a packet needs to be served, it sends the packet to the Linux kernel. After receiving the packet, the Linux kernel processes it as if the packet was received from the NIC. Eventually, the packet is delivered to the service process. After the service process finishes processing the packet, it sends the packet to the socket. The packet will be routed by the Linux kernel towards one of the tun / tap / KNI devices and will be received by the RTC thread. Figure 25b Conceptually illustrates an RTC thread that uses the Linux kernel to communicate with a service process to provide a service.

[0177] V. Computing Device & Virtualization Software

[0178] Virtualization software (also known as a Managed Forwarding Element (MFE) or hypervisor) allows a computing device to host a set of virtual machines (VMs) and perform packet forwarding operations (including L2 switching and L3 routing operations). These computing devices are therefore also referred to as host machines. The packet forwarding operations of the virtualization software are managed and controlled by a set of central controllers, and thus in some embodiments, the virtualization software is also referred to as a Managed Software Forwarding Element (MSFE). In some embodiments, when the virtualization software of a host machine instantiates a local instance of a logical forwarding element as a physical forwarding element operation, the MSFE performs its packet forwarding operations for one or more logical forwarding elements. Some of these physical forwarding elements are Managed Physical Routing Elements (MPREs) that perform L3 routing operations for logical routing elements (LREs), and some of these physical forwarding elements are Managed Physical Switching Elements (MPSEs) that perform L2 switching operations for logical switching elements (LSEs). Figure 26 Illustrates a computing device 2600 that serves as a host machine (or host physical endpoint) running virtualization software of some embodiments of the present invention.

[0179] As shown, the computing device 2600 can access the physical network 2690 through a physical NIC (PNIC) 2695. The host 2600 also runs virtualization software 2605 and hosts VMs 2611 - 2614. The virtualization software 2605 serves as an interface between the hosted VMs and the physical NIC 2695 (and other physical resources such as processors and memories). Each of the VMs includes a virtual NIC (VNIC) for accessing the network through the virtualization software 2605. Each VNIC in the VM is responsible for exchanging packets between the VM and the virtualization software 2605. In some embodiments, the VNIC is a software abstraction of a physical NIC implemented by a virtual NIC emulator.

[0180] The virtualization software 2605 manages the operations of VMs 2611-2614 and includes several components for managing the access of VMs to the physical network (in some embodiments, by implementing a logical network to which the VMs are connected). As shown, the virtualization software includes several components, including MPSE 2620, a group of MPRE 2630, a controller agent 2640, a VTEP 2650, and a group of uplink pipelines 2670.

[0181] The VTEP (VXLAN Tunnel Endpoint) 2650 allows the host machine 2600 to act as a tunnel endpoint for logical network traffic (e.g., VXLAN traffic). VXLAN is an overlay network encapsulation protocol. The overlay network created by VXLAN encapsulation is sometimes referred to as a VXLAN network, or simply VXLAN. When a VM on host 2600 sends a packet (e.g., an Ethernet frame) to another VM in the same VXLAN network but on a different host, the VTEP will encapsulate the data packet using the VNI of the VXLAN network and the network address of the VTEP before sending the packet to the physical network. The packet is tunneled through the physical network (i.e., the encapsulation makes the underlying packet transparent to intermediate network elements) to the destination host. At the destination host, the VTEP de-encapsulates the packet and forwards only the original internal data packet to the destination VM. In some embodiments, the VTEP module only serves as a controller interface for VXLAN encapsulation, and the encapsulation and de-encapsulation of VXLAN packets are completed at the uplink module 2670.

[0182] The controller agent 2640 receives control plane messages from a controller or a controller cluster. In some embodiments, these control plane messages include configuration data for configuring various components of the virtualization software (such as MPSE 2620 and MPRE 2630) and / or virtual machines. In Figure 26 the example shown, the controller agent 2640 receives control plane messages from the controller cluster 2660 from the physical network 2690 and then provides the received configuration data to the MPRE 2630 through the control channel without passing through the MPSE 2620. However, in some embodiments, the controller agent 2640 receives control plane messages from a direct data pipeline (not shown) independent of the physical network 2690. In some other embodiments, the controller agent receives control plane messages from the MPSE 2620 and forwards the configuration data to the router 2630 through the MPSE 2620.

[0183] The MPSE 2620 delivers network data to the physical NIC 2695 interfacing with the physical network 2690 and receives network data from the physical NIC 2695. The MPSE also includes a plurality of virtual ports (vPort) communicatively interconnecting the physical NIC with the VMs 2611 - 2614, the MPRE 2630, and the controller agent 2640. In some embodiments, each virtual port is associated with a unique L2 MAC address. The MPSE performs L2 link layer packet forwarding between any two network elements connected to its virtual ports. The MPSE also performs L2 link layer packet forwarding between any network element connected to any one of its virtual ports and an reachable L2 network element on the physical network 2690 (e.g., another VM running on another host). In some embodiments, the MPSE is a local instantiation of a logical switching element (LSE) that operates across different host machines and can perform L2 packet switching between VMs on the same host machine or different host machines. In some embodiments, the MPSE performs the switching functions of several LSEs according to the configurations of those logical switches.

[0184] The MPRE 2630 performs L3 routing on data packets received from the virtual ports on the MPSE 2620. In some embodiments, this routing operation requires resolving the L3 IP address to a next-hop L2 MAC address and a next-hop VNI (i.e., the VNI of the L2 segment of the next-hop). Each routed data packet is then sent back to the MPSE 2620 to be forwarded to its destination according to the resolved L2 MAC address. The destination can be another VM connected to a virtual port on the MPSE 2620 or an reachable L2 network element on the physical network 2690 (e.g., another VM running on another host, a physical non-virtualized machine, etc.).

[0185] As described above, in some embodiments, an MPRE is a local instantiation of a logical routing element (LRE) that operates across different host machines and can perform L3 packet forwarding between VMs on the same host machine or different host machines. In some embodiments, a host machine can have multiple MPREs connected to a single MPSE, where each MPRE in the host machine implements a different LRE. In some embodiments, even though the MPRE and MPSE are implemented in software, the MPRE and MPSE are referred to as "physical" routing / switching elements to distinguish them from "logical" routing / switching elements. In some embodiments, the MPRE is referred to as a "software router" and the MPSE is referred to as a "software switch". In some embodiments, the LRE and LSE are collectively referred to as logical forwarding elements (LFE), while the MPRE and MPSE are collectively referred to as managed physical forwarding elements (MPFE). Some of the logical resources (LR) mentioned throughout this document are LREs or LSEs that have corresponding local MPREs or local MPSEs running in each host machine.

[0186] In some embodiments, the MPRE 2630 includes one or more logical interfaces (LIF), each of which serves as an interface to a specific segment of the network (L2 segment or VXLAN). In some embodiments, each LIF can be addressed by its own IP address and serves as the default gateway or ARP proxy for network nodes (e.g., VMs) of its specific segment of the network. In some embodiments, all MPREs in different host machines can be addressed by the same "virtual" MAC address (or vMAC), while each MPRE is also assigned a "physical" MAC address (or pMAC) to indicate in which host machine the MPRE is operating.

[0187] The uplink module 2670 relays data between the MPSE 2620 and the physical NIC 2695. The uplink module 2670 includes an egress chain and an ingress chain, each of which performs multiple operations. Some of these operations are preprocessing and / or postprocessing operations for the MPRE 2630. The operations of the LIF, uplink module, MPSE, and MPRE are described in U.S. Patent Application 14 / 137,862, titled "Logical Router", filed on December 20, 2013 and published as U.S. Patent Application Publication 2015 / 0106804.

[0188] As Figure 26As shown, the virtualization software 2605 has multiple MPREs for multiple different LREs. In a multi-tenant environment, a host machine can operate virtual machines from multiple different users or tenants (i.e., connected to different logical networks). In some embodiments, each user or tenant has a corresponding MPRE instantiation of its LRE in the host for handling its L3 routing. In some embodiments, although different MPREs belong to different tenants, they share the same vPort on the MPSE 2620 and thus share the same L2 MAC address (vMAC or pMAC). In some other embodiments, each different MPRE belonging to different tenants has its own port to the MPSE.

[0189] The MPSE 2620 and MPRE 2630 make it possible to forward data packets between VMs 2611 - 2614 without sending them through the external physical network 2690 (as long as the VMs are connected to the same logical network, since VMs of different tenants will be isolated from each other). Specifically, the MPSE performs the function of a local logical switch by using the VNIs of various L2 segments (i.e., their corresponding L2 logical switches) of various logical networks. Similarly, the MPRE performs the function of a logical router by using the VNIs of those various L2 segments. Since each L2 segment / L2 switch has its own unique VNI, the host machine 2600 (and its virtualization software 2605) can direct packets of different logical networks to their correct destinations and effectively separate the traffic of different logical networks from each other.

[0190] VI. Electronic System

[0191] Many of the above features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more processing units (e.g., one or more processors, processor cores, or other processing units), they cause the (one or more) processing units to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard disk drives, EPROMs, etc. Computer-readable media do not include carrier waves and electronic signals transmitted wirelessly or through a wired connection.

[0192] In this specification, the term "software" refers to firmware residing in read-only memory or applications stored in magnetic storage that can be read into memory for processing by a processor. Additionally, in some embodiments, multiple software inventions may be implemented as sub-parts of a larger program while retaining distinct software inventions. In some embodiments, multiple software inventions may also be implemented as separate programs. Finally, any combination of separate programs that implement the software inventions described herein is within the scope of the present invention. In some embodiments, when a software program is installed to operate on one or more electronic systems, the software program defines one or more specific machine implementations of the operations that run and execute the software program.

[0193] Figure 27 The conceptual diagram shows an electronic system 2700 that implements some embodiments of the present invention. The electronic system 2700 can be used to execute any of the control, virtualization, or operating system applications described above. The electronic system 2700 can be a computer (e.g., desktop computer, personal computer, tablet computer, server computer, mainframe, blade computer, etc.), a telephone, a PDA, or any other type of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 2700 includes a bus 2705, one or more processing units 2710, a system memory 2725, a read-only memory 2730, a permanent storage device 2735, an input device 2740, and an output device 2745.

[0194] The bus 2705 collectively represents all system, peripheral, and chipset buses that communicatively connect many internal devices of the electronic system 2700. For example, the bus 2705 communicatively connects the one or more processing units 2710 to the read-only memory 2730, the system memory 2725, and the permanent storage device 2735.

[0195] The one or more processing units 2710 retrieve instructions to be executed and data to be processed from these various memory units in order to execute the processes of the present invention. The one or more processing units can be a single processor or a multi-core processor in different embodiments.

[0196] The read-only memory (ROM) 2730 stores static data and instructions required by the one or more processing units 2710 and other modules of the electronic system. On the other hand, the permanent storage device 2735 is a read-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 2700 is turned off. Some embodiments of the present invention use a mass storage device (such as a magnetic disk or an optical disk and its corresponding disk drive) as the permanent storage device 2735.

[0197] Other embodiments use removable storage devices (such as floppy disks, flash drives, etc.) as the permanent storage device. Like the permanent storage device 2735, the system memory 2725 is a read-write memory device. However, unlike the storage device 2735, the system memory is a volatile read-write memory, such as random access memory. The system memory stores some of the instructions and data required by the processor during operation. In some embodiments, the processes of the present invention are stored in the system memory 2725, the permanent storage device 2735, and / or the read-only memory 2730. The (one or more) processing units 2710 retrieve the instructions to be executed and the data to be processed from these various memory units in order to execute the processes of some embodiments.

[0198] The bus 2705 is also connected to input and output devices 2740 and 2745. The input device enables a user to convey information to and select commands for the electronic system. The input device 2740 includes an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device 2745 displays images generated by the electronic system. The output devices include printers and display devices, such as a cathode ray tube (CRT) or a liquid crystal display (LCD). Some embodiments include devices that serve as both input and output devices, such as a touch screen.

[0199] Finally, as Figure 27 shown, the bus 2705 also couples the electronic system 2700 to a network 2765 via a network adapter (not shown). In this way, the computer can be part of a network of computers (such as a local area network ("LAN"), a wide area network ("WAN"), or an intranet, or a network of networks, such as the Internet). Any or all of the components of the electronic system 2700 can be used in conjunction with the present invention.

[0200] Some embodiments include electronic components such as a microprocessor, a storage device that stores computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as a computer-readable storage medium, a machine-readable medium, or a machine-readable storage medium), and a memory. Some examples of such computer-readable media include RAM, ROM, a read-only compact disc (CD-ROM), a recordable compact disc (CD-R), a rewritable compact disc (CD-RW), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini SD card, micro SD card, etc.), magnetic and / or solid state disk drives, read-only and recordable discs, ultra-dense optical discs, any other optical or magnetic medium, and floppy discs. A computer-readable medium can store a computer program executable by at least one processing unit and including a set of instructions for performing various operations. Examples of computer programs or computer code include machine code such as that produced by a compiler, and files including higher-level code executed by a computer, an electronic component, or a microprocessor utilizing an interpreter.

[0201] Although the foregoing discussion has mainly referred to a microprocessor or multi-core processor that executes software, some embodiments are executed by one or more integrated circuits, such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). In some embodiments, such integrated circuits execute instructions stored on the circuit itself.

[0202] As used in this specification, the terms "computer", "server", "processor", and "memory" all refer to electronic or other technical devices. These terms do not include a person or group of people. For the purposes of this specification, the term display or displaying means displaying on an electronic device. As used in this specification, the terms "computer-readable medium", "plural computer-readable media", and "machine-readable medium" are strictly limited to tangible, physical objects that store information in a form readable by a computer. These terms do not include any wireless signals, wired download signals, and any other transient signals.

[0203] In this document, the term "packet" refers to a set of bits sent across a network in a specific format. Those of ordinary skill in the art will recognize that the term packet can be used herein to refer to various formatted sets of bits that can be sent across a network, such as Ethernet frames, TCP segments, UDP datagrams, IP packets, and the like.

[0204] Throughout this specification, computing and network environments including virtual machines (VMs) are mentioned. However, a virtual machine is just one example of a data computing node (DCN) or a data computing end node (also referred to as an addressable node). A DCN can include a non-virtualized physical host, a virtual machine, a container that runs on top of a host operating system without a hypervisor or a separate operating system, and a hypervisor kernel network interface module.

[0205] In some embodiments, a VM operates with its own guest operating system on a host whose resources are virtualized by virtualization software (e.g., a hypervisor, a virtual machine monitor, etc.). A tenant (i.e., the owner of the VM) can choose which applications to operate on top of the guest operating system. On the other hand, some containers run on top of the host operating system without the need for a hypervisor or a separate guest operating system. In some embodiments, the host operating system uses namespaces to isolate containers from each other and thus provide operating system-level separation of different groups of applications operating within different containers. This separation is similar to the VM separation provided in a hypervisor virtualization environment that virtualizes system hardware and can thus be considered a form of virtualization that isolates different groups of applications operating in different containers. Such containers are lighter than VMs.

[0206] In some embodiments, a hypervisor kernel network interface module is a non-VM DCN that includes a network stack having a hypervisor kernel network interface and receive / transmit threads. An example of a hypervisor kernel network interface module is the vmknic module that is part of the VMware ESXi hypervisor. TM part of the hypervisor.

[0207] Those of ordinary skill in the art will recognize that although this specification refers to VMs, the examples given can be any type of DCN, including physical hosts, VMs, non-VM containers, and hypervisor kernel network interface modules. In fact, in some embodiments, an example network can include a combination of different types of DCNs.

[0208] Although the present invention has been described with reference to many specific details, those of ordinary skill in the art will recognize that the present invention can be embodied in other specific forms without departing from the spirit of the invention. Additionally, the plurality of figures (including Figure 5 , Figure 14 , Figure 20 and Figure 23 ) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in a series of consecutive operations, and different specific operations may be performed in different embodiments. Additionally, a process may be implemented using several sub-processes or as part of a larger macro-process. Thus, those of ordinary skill in the art will understand that the present invention is not limited by the above illustrative details but is defined by the appended claims.

Claims

1. A method for a computing device to implement a gateway data path between a first network and a second network, the method comprising: Receiving a packet from the first network; Executing a plurality of pipeline stages of the gateway data path to determine a next destination of the packet within the second network, wherein the plurality of pipeline stages include a centralized logical router stage and a distributed logical router stage, both the centralized logical router stage and the distributed logical router stage being executed within a single context of the computing device, wherein the centralized logical router is implemented only by the gateway data path at a first set of computing devices between the first network and the second network, and the distributed logical router is implemented by the gateway data path of the first set of computing devices and by a second set of computing devices within the second network.

2. The method according to claim 1, wherein the plurality of pipeline stages are executed at a processor as a single run-to-completion thread of the gateway data path.

3. The method according to claim 1, wherein: The plurality of pipeline stages further include a transit logical switch stage between the centralized logical router stage and the distributed logical router stage; The transit logical switch has a first logical port for the centralized logical router and a second logical port for the distributed logical router; The transit logical switch stage exchanges between the centralized logical router and the distributed logical router; And The transit logical switch stage is executed within a single context of the computing device.

4. The method according to claim 1, wherein the distributed logical router performs routing in each of the first set and the second set of computing devices according to the same routing table.

5. The method according to claim 1, wherein the second network is a provider network, one or more logical networks are implemented on the provider network, and the first network is an external physical network.

6. The method according to claim 1, further comprising: Receiving a second packet at the computing device from the second network; Executing a second plurality of pipeline stages of the gateway data path to determine a next destination of the packet within the first network, wherein the second plurality of pipeline stages include the centralized logical router stage but do not include the distributed logical router stage.

7. The method according to claim 6, wherein the distributed logical router was previously applied to the second packet at another computing device.

8. A machine-readable medium storing a program, the program implementing the method according to any one of claims 1-7 when implemented by at least one processing unit.

9. A computing device, comprising: A set of processing units; And A machine-readable medium storing a program, the program implementing the method according to any one of claims 1-7 when implemented by at least one of the processing units.

10. A system, comprising means for implementing the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Logical Router

    US20150106804A1

  • Logical router

    US9785455B2

  • Logical router with multiple routing components

    US9787605B2

  • Multifunctional wideband gateway and communication method thereof

    CN1595918A

  • Multiple Active L3 Gateways for Logical Networks

    US20150063364A1