Performing services at the edge of a network using a service plane
By introducing service classification and transmission mechanisms into edge forwarding elements, the complexity of providing edge services for various types of businesses in data centers is solved, and efficient service classification and forwarding management is achieved.
Patent Information
- Application Number
- CN202180027018.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-17
- Filing Date
- 2021-02-01
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-02-01
AI Technical Summary
Existing data centers struggle to effectively simplify the service classification of various business types when providing edge services, resulting in complex mechanisms and low efficiency.
By introducing service classification operations into edge forwarding elements, specific types of service sets of data messages are identified, and data is forwarded to the corresponding service nodes using different transmission mechanisms. Combined with connection tracker storage devices and service insertion rules, stateful services and service paths can be managed.
It enables efficient and flexible provision of edge services for different types of businesses within the data center, simplifies the service classification process, and improves the efficiency and reliability of data forwarding.
Smart Images

Figure CN115398864B_ABST
Abstract
Description
BACKGROUND
[0001] Today's data centers provide edge services for a variety of different types of traffic. In the past, edge services for different types of traffic used different mechanisms to perform service classification. New use cases for edge services require more mechanisms to provide edge services. To simplify the provision of edge services for a variety of types of traffic, there is a need in the art for a new approach to providing edge services. SUMMARY
[0002] Some virtualized computing environments provide edge forwarding elements that are located between an external network and an internal network (e.g., a logical network). In some embodiments, the virtualized computing environment provides additional edge forwarding elements (and edge services) for sub-networks within the virtualized computing environment. For example, different logical networks (e.g., tenant networks or provider networks) each having their own edge devices can be implemented within a data center (e.g., by a provider network) that provides edge forwarding elements between an external network and networks within the data center. Each logical network (e.g., a tenant network or a provider network) includes at least one edge forwarding element that, in some embodiments, executes on an edge host or as an edge compute node to provide the logical network with access to the external network and vice versa. In some embodiments, the edge forwarding element provides a set of services (e.g., edge services) for traffic handled by the edge forwarding element.
[0003] Some embodiments provide a novel approach for providing different types of services for a logical network associated with an edge forwarding element that operates between the logical network and an external network. The edge forwarding element receives a data message for forwarding and performs a service classification operation to select a particular type of set of services for the data message. The particular type of service is one of a plurality of different types of services that use different transport mechanisms to forward data to a set of service nodes (e.g., service virtual machines or service devices, etc.) that provide the service. The edge forwarding element then receives the data message after the selected set of services has been performed and performs a forwarding operation to forward the data message. In some embodiments, the approach is also performed by an edge forwarding element that is located at an edge of a logical network segment within the logical network.
[0004] In some embodiments, the transport mechanism includes a logical service forwarding plane (implemented as a logical service forwarding element) that connects the edge forwarding element to a set of service nodes that each provide a service in the set of services. In selecting the set of services, the service classification operation of some embodiments identifies a chain of multiple service operations that must be performed on the data message. In some embodiments, the service classification operation includes selecting a service path for the identified chain of services to provide the multiple services. After selecting the service path, the data message is sent along the selected service path to provide the services. Once the services have been provided, the data message is returned by the last service node in the service path that performs the last service operation to the edge forwarding element, and the edge forwarding element performs a next hop forwarding or performs a forwarding operation to forward the data message.
[0005] Some embodiments provide stateful services in the service chains identified for some data messages. To support stateful services in the service chains, some embodiments generate connection tracking records in a connection tracker storage device used by the edge forwarding element to track service insertion decisions made for multiple data message flows that require multiple different sets of services (i.e., service chains). The edge forwarding element (e.g., a router) receives a data message at a particular interface of the edge forwarding element that is traversing the edge forwarding element in a forward direction between two machines. In some embodiments, the data message is the first data message in a forward data message flow (e.g., a set of data messages that share the same set of attributes) that, together with a reverse data message flow between the two machines, constitutes a bidirectional flow.
[0006] The edge forwarding element identifies (1) a set of stateful services for the received data message, and (2) a next hop associated with the identified set of stateful services in the forward direction and a next hop associated with the identified set of stateful services in the reverse direction. Based on the identified set of services and the next hops for the forward and reverse directions, the edge forwarding element generates and stores first and second connection tracking records for the forward and reverse data message flows, respectively. The first and second connection tracking records each include the next hop identified for the forward and reverse data message flows, respectively. The edge forwarding element forwards the received data message to the next hop identified for the forward direction, and for subsequent data messages of the forward and reverse data message flows received by the edge forwarding element, uses the stored connection tracking records to identify the next hop for forwarding.
[0007] Some embodiments configure the edge forwarding elements to perform service insertion operations to identify stateful services to perform on data messages received at a plurality of virtual interfaces of the edge forwarding elements for forwarding. In some embodiments, the service insertion operations include applying a set of service insertion rules. The service insertion rules (1) specify a set of criteria and a corresponding action (e.g., a redirect action and a redirect destination) to take on data messages matching the criteria, and (2) are associated with a set of interfaces for which the service insertion rules apply. In some embodiments, a universal unique identifier (UUID) is used to specify the action, which is then used as a matching criterion to identify a type of service insertion and a subsequent policy lookup for a next hop data set. The edge forwarding elements are configured to, for each virtual interface, apply the relevant set of service insertion rules to data messages received at the virtual interface (i.e., make a service insertion decision).
[0008] As described above, the edge forwarding elements are configured with a connection tracker storage device that stores connection tracking records for data message flows based on results of service insertion operations performed for first data messages in the data message flows. In some embodiments, the connection tracker storage device is a common storage device for all interfaces of the edge forwarding elements, and each connection tracking record includes identifiers of service insertion rules used to identify a set of stateful services and next hops for the data message flow corresponding to the connection tracking record.
[0009] In some embodiments, the service insertion operations include a first lookup in the connection tracker storage device to identify a connection tracking record for a data message received at an interface, if it exists. If the connection tracking record exists, all connection tracking records that include a set of data message attributes (e.g., data message flow identifiers) that match data message attributes of the received data message are identified as a set of possible connection records for the data message. Based on the service insertion rule identifiers and the interface on which the data message was received, a connection tracking record in the set of possible connection records that stores an identifier of a service insertion rule applied to the interface is identified as a connection tracking record to store an action for the received data message. If a connection tracking record for the received data message is identified, the edge forwarding element forwards the data message based on the action stored in the connection tracking record. If a connection tracking record is not identified (e.g., the data message is a first data message in a data message flow), the edge forwarding element uses the service insertion rules to identify an action for the data message and generates a connection tracking record, and stores the connection tracking record in the connection tracker storage device.
[0010] Some embodiments provide a method of performing a stateful service that tracks state changes of service nodes to update a connection tracker record when necessary. At least one global state value indicative of a state of a service node is maintained at an edge device. In some embodiments, different global state values are maintained for service chain service nodes (SCSNs) and Layer 2 bump-in-the-wire service nodes (L2 SNs). The method generates a record in a connection tracker storage device that includes a current global state value as a flow state value for a first data message in a data message flow. Each time a data message is received for the data message flow, the stored state value (i.e., the flow state value) is compared to the relevant global state value (e.g., SCSN state value or L2 SN state value) to determine whether the stored action has been updated.
[0011] After a global state value associated with a flow changes, the global state value and the flow state value do not match, and the method checks a flow programming table to determine whether the flow has been affected by a flow programming instruction(s) that caused the global state value to change (e.g., increment). In some embodiments, the instructions stored in the flow programming table include a data message flow identifier and an updated action (e.g., drop, allow, update selected service path, update next hop address). If the data message flow identifier stored in the flow programming table does not match the current data message flow identifier, the flow state value is updated to the current global state value, and the data message is processed using the action stored in the connection tracker record. However, if at least one data message flow identifier stored in the flow programming table matches the current data message flow identifier, the flow state value is updated to the current global state value, and the action stored in the connection tracker record is updated to reflect the execution of the instructions with the matching flow identifier stored in the flow programming table, and the updated action is used to process the data message.
[0012] In some embodiments, an edge forwarding element is configured to provide services using service logic forwarding elements as transport mechanisms. The edge forwarding element is configured to use different transport mechanisms to connect different sets of virtual interfaces of the edge forwarding element to different network elements of a logical network. For example, a first set of virtual interfaces is configured to be connected to a set of forwarding elements inside the logical network using logical forwarding elements that connect source machines and destination machines of traffic of the logical network. In some embodiments, traffic received on the first set of interfaces is forwarded by the edge forwarding element toward the destination to a next hop without being returned to the forwarding element that received it. A second set of virtual interfaces is configured to be connected to a set of service nodes to provide services for data messages received at the edge forwarding element.
[0013] Each connection established for the second set of virtual interfaces can use a different transport mechanism, such as a service-logic forwarding element, a tunneling mechanism, and a line-in-card mechanism, and in some embodiments, some or all of the transport mechanisms are used to provide data messages to the service nodes. Each virtual interface in the third set of virtual interfaces is configured to connect to a service-logic forwarding element that connects the edge forwarding element to at least one of the set of internal forwarding elements. The virtual interfaces are configured to (1) receive data messages from the at least one internal forwarding element that are to be serviced by at least one of the set of service nodes, and (2) return the serviced data messages to the internal forwarding element network.
[0014] Some embodiments facilitate providing services that are reachable at a virtual internet protocol (VIP) address. A client uses the VIP address to access a set of service nodes in a logical network. In some embodiments, a data message from the client machine to the VIP is directed to an edge forwarding element, where the data message is redirected to a load balancer that load balances among the set of service nodes to select a service node to provide a service requested by the client machine. In some embodiments, the load balancer does not change the source IP address of the data message received from the client machine, so that the service node receives the data message to service with the client machine IP address identified as the source IP address. The service node services the data message using the service node's IP address as the source IP address, the client node's IP address as the destination IP address, and sends the serviced data message to the client machine. Because the client sent the original address to the VIP address, the client will not recognize the source IP address of the serviced data message as a response to the request sent to the VIP address, and the serviced data message will not be properly handled (e.g., it will be dropped, or not associated with the original request).
[0015] In some embodiments, facilitating the provision of services includes using service logic forwarding elements to return serviced data messages to the load balancer to track the state of the connection. To use service logic forwarding elements, some embodiments configure the egress data path of the service node to intercept serviced data messages before forwarding the serviced data messages to the logical forwarding elements in the data path from the client to the service node, and determine whether the serviced data messages need to be routed by the edge forwarding element through the routing services provided as the service. If the data message needs to be routed by the routing services (e.g., for serviced data messages), the serviced data messages are forwarded through the service logic forwarding elements to the edge forwarding element. In some embodiments, the serviced data messages and the VIP associated with the service are provided to the edge forwarding element, in other embodiments, the edge forwarding element determines the VIP based on the port used to send the data message on the service logic forwarding element. The edge forwarding element uses the VIP to identify the load balancer associated with the serviced data message. The serviced data messages are then forwarded to the load balancer for the load balancer to maintain state information for the connection to which the data message belongs, and to modify the data message to identify the VIP as the source address for forwarding to the client.
[0016] In some embodiments, the transport mechanism includes a tunneling mechanism (e.g., virtual private network (VPN), Internet Protocol Security (IPSec), etc.) that connects the edge forwarding element to at least one service node through a corresponding set of virtual tunnel interfaces (VTIs). In addition to the VTIs used to connect the edge forwarding element to the service nodes, the edge forwarding element uses other VTIs to connect to other network elements for which it provides forwarding operations. At least one of the VTIs used to connect the edge forwarding element to other (i.e., non-service node) network elements is identified to perform service classification operations, and is configured to perform service classification operations on data messages received at the VTI for forwarding. In some embodiments, the VTIs used to connect the edge forwarding element to the service nodes are not configured to perform service classification operations, but are configured to mark data messages returned to the edge forwarding element as serviced. In other embodiments, the VTIs used to connect the edge forwarding element to the service nodes are configured to perform limited service classification operations using a single default rule applicable to the VTIs that marks data messages returned to the edge forwarding element as serviced.
[0017] For traffic that exits the logical network through a particular VTI, some embodiments perform a service classification operation on different data messages to identify different VTIs that connect the edge forwarding element to service nodes to provide services required by the data messages. In some embodiments, each data message is then forwarded to the identified VTI to receive the required services (e.g., from service nodes that are connected to the edge forwarding element through the VTI). The identified VTIs do not perform service classification operations and only allow the data messages to reach the service nodes. The service nodes then return the serviced data messages to the edge forwarding element. In some embodiments, the VTIs are not configured to perform service classification operations, but are configured to mark all traffic directed from the service nodes to the edge forwarding element as serviced. The marked serviced data messages are then received at the edge forwarding element and forwarded through the particular VTI to the destination of the data messages. In some embodiments, the particular VTI does not perform additional service insertion operations because the data messages are marked as serviced.
[0018] The foregoing summary is intended to serve as a brief introduction to some embodiments of the application. It is not intended to be a complete description of all inventive subject matter disclosed herein. The detailed description and the accompanying drawings are to further describe embodiments described in the summary and other embodiments. Thus, to understand the full scope of all embodiments described herein, the summary, detailed description, drawings, and claims must be consulted. Moreover, the claimed subject matter is not limited by the illustrative details in the summary, detailed description, drawings. BRIEF DESCRIPTION OF DRAWINGS
[0019] The novel features of the application are set forth with particularity in the appended claims. However, the application is desirably understood to extend to other and novel embodiments as well. To that end, some embodiments of the application are presented in the following detailed description and the accompanying drawings to provide a thorough understanding of the application.
[0020] Figure 1 A process performed by an edge device to perform a service classification operation to select a particular set of services for a data message and to identify forwarding information for the data message is conceptually illustrated.
[0021] Figure 2 A process to identify whether a connection tracker record is stored in a connection tracker storage device used in some embodiments is conceptually illustrated.
[0022] Figure 3 A process to forward a data message at an edge forwarding component that is forwarded by Figure 1 The process provides a service type and forwarding information.
[0023] Figure 4A logical network with two tiers of logical routers (i.e., availability zone (AZ) logical gateway routers) is conceptually illustrated.
[0024] Figure 5 A possible management plane view of a logical network is shown, in which both AZGs and VPCGs include centralized components.
[0025] Figure 6 A logical network with two tiers of logical routers (i.e., availability zone (AZ) logical gateway routers) is conceptually illustrated. Figure 5 A physical implementation of the management plane structure of the two-tier logical network is shown, in which both VPCGs and AZGs include SRs and DRs.
[0026] Figure 7 Logical processing operations for availability zone (T0) logical router components included in an edge data path executed by an edge device for a data message are shown.
[0027] Figure 8 A TX SR acting as a source of traffic on a logical service forwarding element is shown.
[0028] Figure 9 A service path including two service nodes accessed by a LSFE including a TX SR is shown.
[0029] Figure 10 A second embodiment including two edge devices and respectively executing availability zone gateway data paths and virtual private cloud gateway data paths is shown.
[0030] Figure 11 A set of operations performed by a service router at T0 or T1 for a first data message in a data message stream requiring a service from a set of service nodes of a custom service path are shown.
[0031] Figure 12 A set of operations performed by a service router at T0 or for a data message in a data message stream requiring a service from a set of service nodes of a custom service path are shown.
[0032] Figure 13 A process for verifying or updating an identified connection tracker record for a data message stream is conceptually illustrated.
[0033] Figure 14 A set of connection tracker records in a connection tracker storage device and a set of exemplary flow programming records in a flow programming table are shown.
[0034] Figure 15 An object data model of some embodiments is shown.
[0035] Figure 16 Several operations that the network manager and controller perform in some embodiments to define rules for service insertion, next service hop forwarding, and service processing are conceptually illustrated.
[0036] Figure 17 A process for configuring a logical forwarding element to connect to a logical service forwarding plane is conceptually illustrated.
[0037] Figure 18 A set of operations performed by a service router at T0 or T1 for a first data message in a data message flow that requires service from a service node reachable through a tunneling mechanism is illustrated.
[0038] Figure 19 A set of operations performed by a service router at T0 or T1 for a first data message in a data message flow that requires service from a service node reachable through a L2BIW mechanism is illustrated.
[0039] Figure 20A A data message sent from a compute node in a logical network (e.g., logical network A) implemented in a cloud environment to a compute node in an external data center is conceptually illustrated.
[0040] Figure 21A A data message sent from a compute node in an external data center to a compute node in a logical network implemented in a cloud environment is conceptually illustrated.
[0041] Figure 22 A first method for providing service for a data message at an uplink interface in a set of uplink interfaces is conceptually illustrated.
[0042] Figure 23 A second method for providing service for a data message at an uplink interface in a set of uplink interfaces is conceptually illustrated.
[0043] Figure 24 A logical network that provides service classification operations at multiple routers of the logical network is conceptually illustrated.
[0044] Figure 25 An edge forwarding element that connects to a service node using multiple transport mechanisms is conceptually illustrated.
[0045] Figure 26 A logical network that includes three VPC service routers 2630 belonging to two different tenants is illustrated.
[0046] Figure 27 A logical network is shown that includes three VPC service routers 2630 belonging to three different tenants.
[0047] Figure 28 A process for accessing services provided at availability zone edge forwarding elements from a VPC edge forwarding element is shown conceptually.
[0048] Figure 29 A process that an availability zone service router is to perform when it receives a data message from a VPC service router as part of the process is shown conceptually.
[0049] Figure 30 A VPC service router handling a data message sent from a first compute node to a second compute node in a second network segment served by a second VPC service router is shown conceptually.
[0050] Figure 31 A VPC service router handling a data message sent from an external network to a compute node is shown conceptually.
[0051] Figure 32A -B shows a set of data messages used to provide a service addressable at a VIP to clients served by the same virtual private cloud gateway (e.g., a virtual private cloud gateway service and a distributed router).
[0052] Figure 33 An electronic system with which some embodiments of the present application are implemented is shown conceptually. DETAILED DESCRIPTION
[0053] In the following detailed description of the application, numerous specific details, examples and embodiments of the application are set forth in order to provide a thorough understanding of the present application. However, it will be clear to those skilled in the art that the present application is not limited to the embodiments described and that the present application can be practiced with some of the specific details and examples discussed.
[0054] Some virtualized computing environments / logical networks provide edge forwarding elements that sit between external networks and internal networks (e.g., logical networks). In some embodiments, the virtualized computing environment provides additional edge forwarding elements (or edge services) for sub-networks within the virtualized computing environment. For example, different logical networks, each with their own edge devices, can be implemented within a data center (e.g., by a provider network) that provides edge forwarding elements between external networks and networks internal to the data center. Each logical network (e.g., a tenant network or a provider network) includes at least one edge forwarding element, which in some embodiments, executes in an edge host or as an edge compute node to provide the logical network with access to external networks, and vice versa. In some embodiments, the edge forwarding elements provide a set of services (e.g., middlebox services) for traffic that the edge forwarding elements process.
[0055] As used in this document, a data message refers to a set of bits in a particular format that is sent across a network. Those of ordinary skill in the art will recognize that the term "data message" is used in this document to refer to various formatted sets of bits that are sent across a network. The formatting of these bits can be specified by a standardized protocol or a non-standardized protocol. Examples of data messages that follow a standardized protocol include Ethernet frames, IP packets, TCP segments, UDP datagrams, and the like. Further, as used in this document, references to L2, L3, L4, and L7 layers (or Layer 2, Layer 3, Layer 4, and Layer 7) are references to the second data link layer, the third network layer, the fourth transport layer, and the seventh application layer, respectively, of the OSI (Open Systems Interconnection) layer model.
[0056] Further, in this example, each logical forwarding element is a distributed forwarding element that is implemented by configuring a plurality of software forwarding elements (SFEs) (i.e., managed forwarding elements) on a plurality of host computers. To this end, in some embodiments, each SFE or a module associated with the SFE is configured to encapsulate data messages of the LFE with an overlay network header that contains a virtual network identifier (VNI) associated with an overlay network. Thus, in the following discussion, the LFE is considered to be an overlay network fabric that spans a plurality of host computers.
[0057] In some embodiments, LFEs also span configured hardware forwarding elements (e.g., top-of-rack switches). In some embodiments, each LFE is a logical switch implemented by configuring multiple software switches (referred to as virtual switches or vswitches) or related modules on multiple host computers. In other embodiments, LFEs can be other types of forwarding elements (e.g., logical routers), or any combination of forwarding elements (e.g., logical switches and / or logical routers) that form a logical network or part thereof. Many examples of LFEs, logical switches, logical routers, and logical networks exist today, including those provided by VMware's NSX network and service virtualization platform.
[0058] Some embodiments provide novel methods for providing different types of services for a logical network associated with an edge forwarding element, which are performed by an edge device that acts between the logical network and an external network. The edge device receives a data message for forwarding, and performs a service classification operation to select a particular set of services for the data message. Figure 1 The process 100 performed by an edge device to perform a service classification operation to select a particular set of services for a data message and identify forwarding information for the data message is shown conceptually.
[0059] In some embodiments, the process is performed as part of an edge data path, for data messages entering the network, before a routing operation. In some embodiments, the process is performed by a network interface card (NIC) that is designed or programmed to perform the service classification operation. In some embodiments, the process 100 is additionally or alternatively performed by the edge device as part of logical processing at a plurality of virtual interfaces of a logical edge forwarding element, including a set of virtual tunnel interfaces (VTIs) used to connect the edge forwarding element to compute nodes outside the data center. In some embodiments, a particular interface is configured to perform the service classification operation (e.g., by having the service classification tag toggled to "1"), while other interfaces are not configured to perform the service classification operation (e.g., if the service classification tag is set to "0"). In some embodiments, as part of a processing pipeline, a centralized (e.g., service) router invokes a set of service insertion and service transport layer modules (such as the modules in element 735 of Figure 7
[0060] Process 100 continues by receiving (at 110) a data message at an interface (e.g., NIC, VTI) connected to an external network (e.g., a router external to a data center implementing a logical network). In some embodiments, the data message is received from the external network as part of a communication between a client in the external network and a compute node (e.g., a server or service node) in the logical network (or vice versa). In some embodiments, the data message is a data message between two compute nodes in the external network that are receiving service at the edge of the logical network.
[0061] Process 100 continues by determining (at 120) whether a connection tracker record is stored in a connection tracker for a data message flow to which the data message belongs. Figure 2 A process 200 for identifying whether a connection tracker record is stored in a connection tracker storage used in some embodiments is conceptually illustrated. The determination (at 120) includes determining (at 221) whether the connection tracker storage stores any record (i.e., entry in the connection tracker storage) with a flow identifier that matches a flow identifier of the received data message. In some embodiments, the flow identifier is a set of header values (e.g., a five-tuple), or a value generated based on a set of header values (e.g., a hash of the set of header values). If no matching entry is found, process 200 determines that no connection tracker record is stored for the data message, and process 200 produces a “no” at operation 120 of process 100. In some embodiments, the connection tracker storage stores multiple possible matching entries, distinguished by a label indicating the type of stateful operation that created the connection tracker record (e.g., a preliminary firewall operation or a service classification operation). In other embodiments, separate connection tracker storages are maintained for different types of stateful operations. In some embodiments, a connection tracker record created by a service classification operation includes a rule identifier associated with a service insertion rule that (1) was applied to a first data message in the data message flow, and (2) determined the contents of the connection tracker record.
[0062] If at least one matching connection tracker record is found in the connection tracker storage, process 200 determines (at 222) a tag (e.g., a flag bit) identifying whether the record was created as part of the service classification operation or as part of a different stateful process (e.g., a standalone firewall operation). In some embodiments, the tag is compared to a value stored in a buffer associated with the data message, which is used during logical processing to store data other than data typically included in the data message (e.g., context data, the interface over which the data message was received, etc.). If the tag of the record(s) with the matching flow identifier does not indicate that it is related to the service classification operation, process 200 produces a “no” at operation 120 of process 100.
[0063] However, if at least one record includes both the matching flow identifier (at 221) and the matching service classification operation tag (at 222), the process identifies (at 223) the interface to which the service insertion rule used to generate each of the potentially matching records was applied (i.e., the interface in the “applied to” field of the rule that was hit by the first data message of the potentially matching record). In some embodiments, the rule identifier is stored in the connection tracker record, and the rule identifier is associated with (e.g., points to) a data storage (e.g., a container) storing a list of interfaces to which it was applied. In such embodiments, identifying the interface to which the rule used to generate each of the potentially matching records was applied includes identifying the interface stored in the data storage associated with the rule.
[0064] The process then determines (at 224) whether any of the interfaces to which the rule was applied is the interface over which the current data message was received. In some embodiments, data messages of the same data message flow are received at different interfaces based on load balancing operations (e.g., equal cost multi-path (ECMP)) performed by forwarding elements (e.g., routers) in the external network. Moreover, some data messages are necessarily received at multiple interfaces as different service rules are applied as part of the processing pipeline. For example, a data message received at a first VTI that applies a particular service rule identifies a second VTI to which the data message is to be redirected to provide a service required by the data message. The second VTI is connected to a service node that provides the required service, and after the data message is serviced, the data message is returned to the second VTI. The flow identifier matches the connection tracker record of the original data message, but the service insertion rule identified in the connection tracker record is not applied to the data message received at the second VTI (e.g., the applied to field of the service insertion rule does not include the second VTI), such that the data message is not redirected to the second VTI to be serviced again.
[0065] In some embodiments, the interface is identified by a UUID (e.g., a 64-bit or 128-bit identifier) that is too large to be stored in the connection tracker record. (At 223) the UUID (or other identifier) of the identified interface is compared to the UUID of the interface on which the data message was received. As described above, in some embodiments, the UUID of the interface is stored in the buffer associated with the data message. If no interface that applied the (potentially matching connection tracker record's) rule matches the interface on which the data message was received, process 200 produces a "no" at operation 120 of process 100. However, if a connection tracker record is associated with the interface on which the data message was received (i.e., the rule used to generate the connection tracker record was applied to the interface on which the data message was received), process 200 produces a "yes" at operation 120 of process 100. In some embodiments, additional state values associated with the service node state are checked, as will be discussed in connection with Figure 13 .
[0066] If process 100 determines (at 120) that the data message belongs to a flow with a connection tracker record, process 100 retrieves (at 125) a service action based on information in the connection tracker record. In some embodiments, the service action includes a service type and a set of forwarding information stored in the connection tracker record. Additional details regarding retrieving the service action are described in connection with Figure 12 and Figure 13 In some embodiments, the service type identifies a transport mechanism (e.g., a logical service forwarding element, an L3 VPN, or an L2 in-line plugin). In some embodiments, the forwarding information includes different types of forwarding information for different types of service insertion types. For example, the forwarding information for a service provided by a service chain includes a service path identifier and a next hop MAC address. The forwarding information for a service node that is an in-line plugin service node or a service node connected through a virtual private network includes a next hop IP. The service type and the forwarding information are then provided (at 170) to an edge forwarding element (e.g., a virtual routing and forwarding (VRF) context of the edge forwarding element), and the process ends. In some embodiments, the service type and the forwarding information are provided (at 170) to a transport layer module that redirects the data message to the service node using the transport mechanism identified by the service type to reach the destination identified by the forwarding information, as described in connection with Figure 3 .
[0067] If process 100 determines (at 120) that there is no connection tracker storage entry for the received data message due to any reason identified in process 200, process 100 performs (at 130) a first service classification lookup against a set of service insertion rules to find a highest priority rule defined for data messages having a set of attributes shared by the received data message. In some embodiments, the set of data message attributes in a particular service insertion rule can include any of: header values at layer 2, layer 3, or layer 4, or hash values based on any header values, and can include wildcard values for certain attributes (e.g., fields) that allow any value for that attribute. In embodiments described with respect to process 100, the service insertion rule identifies a universally unique identifier (UUID) associated with a set of actions for data messages matching the service insertion rule. In other embodiments, the service insertion rule includes a set of actions to perform for the received data message (e.g., redirect to a particular address using a particular transport mechanism). In some embodiments, a lowest priority (e.g., default) rule that applies to all data messages (e.g., specifies all wildcard values) is included in the set of service insertion rules and is identified if no other service insertion rule with a higher priority is identified. In some embodiments, the default rule will specify a no-op that causes the data message to be provided to a routing function of the edge forwarding element to be routed without any service performed on the data message. In other embodiments, the default rule will cause the data message to be provided to the routing function with an indication that the data message does not require additional service classification operations.
[0068] After identifying the UUID associated with the service insertion rule (and data message) (at 130), the process 100 performs a policy lookup based on the UUID identified based on the service insertion rule (at 130) (at 140). In some embodiments, the separation of the service insertion rule lookup and the UUID (policy) lookup is used to simplify policy updates for multiple service insertion rules by changing the policy associated with a single UUID rather than having to update each service insertion rule. The UUID lookup is used to identify a forwarding information set and to identify a particular service type (e.g., a service using a particular transport mechanism). For example, for different data messages, the UUID lookup can identify any of the following: a next hop IP address (for a tunneling mechanism), a virtual next hop IP address (for an in-line appliance mechanism), or a forwarding data set including at least a service path ID, a service index, and a next hop layer 2 (e.g., MAC) address (for a mechanism using a service logic forwarding element). In some embodiments, the type of transport mechanism is inferred from the type of forwarding information identified for the data message. Some embodiments using a service logic forwarding element use a layer 3 (e.g., IP) address to identify a next hop. In such embodiments, it can be necessary to include a service type identifier.
[0069] In using the UUID to identify a forwarding information set and to identify a particular service type, some embodiments perform a load balancing operation to select among multiple next hops to provide the identified service. In some embodiments, the identified next hop is a service node that provides a different service. In some embodiments, the service node includes at least one of a service virtual machine and a service appliance. In some embodiments, the load balancing operation is based on any of: a round robin mechanism, a load based selection operation (e.g., selecting a service node with the lowest current load), or a distance based selection operation (e.g., selecting a nearest service node as measured by a selected metric).
[0070] After determining the service action and forwarding information (at 140), the data message flow identifier and forwarding information for the reverse flow are identified (at 150). In many cases, the data message flow identifier for the reverse flow is based on the same set of header values for the forward data message flow, with the source and destination addresses being swapped. The forwarding information for the reverse data message flow for certain types of service insertion (i.e., particular types of transport mechanisms) is different for the forward and reverse flows. For certain types of service insertion (i.e., transport mechanisms), the forwarding information for the reverse data message flow identifies the next hop for the reverse flow that was the previous hop for the forward flow. For other types of service insertion (e.g., tunneling mechanisms), the reverse forwarding information identifies the same next hop (e.g., the next hop IP address of the tunnel endpoint). In some embodiments, operation 150 is skipped since reverse connection tracker records are not necessary. For example, some rules specify that they apply only to data messages in a particular direction.
[0071] Based on the data message flow identifiers and forwarding information identified for the forward and reverse flows, a set of connection tracker records are generated (at 160) for the forward and reverse data message flows, with state information (e.g., data message identifiers and forwarding information) for the forward and reverse data message flows, respectively. In some embodiments, generating the connection tracker records includes querying a state value stored in the flow programming table that reflects a current state version of the set of service node types associated with the type of service identified for the data message in the flow programming table. In some embodiments, the flow IDs for the forward and reverse data message flows are the same, except for a directionality bit that indicates whether it is a forward or reverse data message.
[0072] In some embodiments, the reverse flow identifier is different than the reverse flow identifier that would be generated based on a data message received in the forward direction. For example, a simple reverse identifier generation operation would swap the source and destination IP (L3) and MAC (L2) addresses and generate an identifier based on the swapped header values, but if the service node performs a NAT operation, then a data message received in the reverse direction would generate a reverse flow identifier based on the translated address rather than based on the original (forward) data message header address. In some embodiments, a return data message with a different set of flow identifiers (e.g., header values, etc.) would be treated as a new flow and new connection tracker record for the forward and reverse of the data message flow associated with the reverse data message flow of the original data message.
[0073] Additional details regarding connection tracker records and flow programming tables are described in Figure 14discussed. In some embodiments, after creating the connection tracker record for the forward and reverse data message flows, the data message is provided (at 170) to a component of the edge forwarding element along with the forwarding information and the service type for the component of the edge forwarding element to process, which is responsible for providing the data message to the service node, as discussed below with respect to Figure 3
[0074] Figure 3 Conceptually illustrated is a process 300 for forwarding a data message at an edge forwarding component, which is provided a service type and forwarding information by process 100. In some embodiments, process 300 is performed by an edge forwarding element executing on an edge device. In some embodiments, the edge forwarding element executes as a virtual machine, while in other embodiments, the edge forwarding element is a managed forwarding element (e.g., a virtual routing and forwarding context) executing on the edge device. In some embodiments, some operations of the process are performed by service insertion and service transport layer modules (e.g., 720-729 of element 720) invoked by a service (e.g., centralized) router (e.g., 730). Process 300 begins by receiving (at 310) a data message along with a service type and forwarding information for the data message determined using service classification operations of process 100 in some embodiments. Figure 7
[0075] Process 300 determines (at 320) a service insertion type associated with the received data message. In some embodiments, the determination is made based on the service type information received from the service classification operations. In other embodiments, the determination is made implicitly based on the type of forwarding information received from the service classification operations. For example, an IP address provided as the forwarding information for a particular data message for a virtual tunnel interface (VTI) indicates that the transport mechanism is a tunnel mechanism. Alternatively, a pseudo IP address provided as the forwarding information indicates that the transport mechanism is a layer 2 inline mechanism. If the forwarding information includes a service path identifier and a next hop MAC address, the transport mechanism is understood to be a logical service forwarding plane of a service chain.
[0076] If the process 300 determines (at 320) that the service type uses a tunneling mechanism, the process 300 identifies (at 332) an egress interface based on the IP address provided by the service classification operation. In some embodiments, the egress interface is identified by a routing function associated with the service transport layer module. Based on the identified egress interface, the data message is provided (at 342) to a VTI that encapsulates the data message for delivery over a virtual private network (VPN) tunnel to a service node to provide the service. In some embodiments, the tunnel uses an Internet Protocol Security (IPsec) protocol to tunnel the data message to the service node. In some embodiments that use a secure VPN (e.g., IPsec), the data message is encrypted prior to being encapsulated for forwarding using the tunnel mechanism. In some embodiments, the encryption and encapsulation are performed as part of the data path for the virtual tunnel interface used to connect to the service node (e.g., hereinafter referred to as an L3 service node).
[0077] The encapsulated (and encrypted) data message is then sent over the VPN to the L3 service node for the L3 service node to provide the service and return the serviced data message back to the edge forwarding element. After the service node provides the service, the serviced data message is received (at 352) at the edge forwarding element (e.g., service transport layer module) and the data message is provided (at 380) to a routing function (e.g., a routing function implemented by the edge forwarding element) for forwarding to the destination. In some embodiments, the routing is based on the original destination IP address associated with the data message that is maintained in a memory buffer of the edge device associated with the data message, which in some embodiments stores additional metadata such as the interface on which the data message was received and data for features associated with the edge forwarding element such as IP fragmentation, IPsec, access control lists (ACLs), etc.
[0078] If the process 300 determines (at 320) that the service type uses a Layer 2 in-line plugin transport mechanism, the process 300 identifies (at 334) the source and destination interfaces based on a set of next hop pseudo IP addresses provided by the service classification operation. The next hop pseudo IP addresses are used to identify the source and destination Layer 2 (e.g., MAC) addresses associated with the in-line plugin service node (i.e., a service node that does not change the source and destination Layer 2 (e.g., MAC) addresses of the data message). In some embodiments, the set of next hop pseudo IP addresses includes a set of source and destination pseudo IP addresses that are resolved to source and destination Layer 2 (e.g., MAC) addresses associated with different interfaces of the edge forwarding element. In some embodiments, the different interfaces are identified by a routing function associated with the service transport layer module. In some embodiments, the different interfaces are used to distinguish between data messages traversing the edge device (e.g., edge forwarding element) in different directions (e.g., north to south traffic versus south to north traffic) such that data messages traveling in one direction (e.g., from within a logical network to an external network) use a first interface as the source and a second interface as the destination, and data messages traveling in the opposite direction (e.g., from an external network to a logical network) use the second interface as the source and the first interface as the destination.
[0079] The data message is then sent (at 344) from the source interface to the destination interface using the identified source and destination Layer 2 addresses. After the data message is sent (at 344) using the identified interfaces to the service node, the edge forwarding element receives (at 354) the serviced data message from the service node at the destination interface. The serviced data message is then provided (at 380) to a routing function (e.g., a routing function implemented by the edge forwarding element) for forwarding to the destination. In some embodiments, the routing is based on an original destination IP address associated with the data message that is maintained throughout the processing of the data message. In other embodiments, the original destination IP address is maintained in a memory buffer of the edge device associated with the data message, which in some embodiments stores additional metadata such as the interface on which the data message was received, as well as data associated with features of the edge forwarding element such as IP fragmentation, IPsec, access control lists (ACLs), etc.
[0080] If the process 300 determines (at 320) that the service type uses a service-logic forwarding element transport mechanism, the process 300 identifies (at 336) the interface associated with the service-logic forwarding element based on a table that stores the association between the logical forwarding elements of the service-logic forwarding plane and the interfaces connected to the logical forwarding elements. In some embodiments, the table is a global table provided by a network management or control compute node and includes information for all logical forwarding elements in the logical network that are connected to any one of the set of service-logic forwarding elements. In some embodiments, the interface associated with the logical service forwarding element is identified based on the forwarding information (e.g., based on a service path identifier or a service virtual network identifier provided as part of the forwarding information).
[0081] The data message is then sent (at 346) to the identified interface (or logical service plane data message processor) along with service path information and service metadata (SMD) to be encapsulated with a logical network identifier (LNI) for delivery to the first service node in the service path identified in the service path information. In some embodiments, the service path information provided as part of the forwarding information includes (1) a service path identifier (SPI) used by the logical forwarding element and each service node to identify the next-hop service node, (2) a service index (SI) indicating the position of the hop in the service path, and in some embodiments, (3) a time-to-live. In some embodiments, the LNI is a service virtual network identifier (SVNI). Further details of using a service forwarding plane can be found in U.S. Patent Application 16 / 444,826, filed June 18, 2019, which is incorporated by reference herein.
[0082] After being serviced by the service nodes in the service path, the data message is received (at 356) at the edge forwarding element. In some embodiments, the edge forwarding element receives the data message as a routing service node identified as the last hop in the service path identified for the data message. In such embodiments, the service router implements a service proxy to receive the data message according to a standardized protocol for service chaining using a service path. In some embodiments, the edge forwarding element receives the service data message along with service metadata that identifies the original source and destination addresses to be used for forwarding the data message to its destination. In some embodiments, the service metadata also includes any flow programming instructions sent by a service node or service insertion proxy on the service path. In some embodiments, the flow programming instructions include instructions for modifying how service classification operations select service chains, service paths, and / or forwarding data message flows along a service path. In other embodiments, the flow programming involves other modifications to how the service plane processes data message flows. Flow programming will be further described below.
[0083] Then, process 300 determines (at 366) whether the received service data message includes flow programming instructions. If process 300 determines that flow programming instructions are included with the serviced data message, the flow programming table is updated (at 375) by adding the flow programming instructions to the table for processing of subsequent data messages in the data message flow. In some embodiments, the flow programming instructions identify the flow to which the flow programming instructions pertain and a new service action (e.g., pf_value) for the identified flow. In some embodiments, the new service action is an instruction to skip a particular service node (e.g., a firewall service node) for the next data message or for all subsequent data messages in the data message flow (e.g., if the firewall service node determines to allow the data message flow), or to discard all subsequent data messages of the data message flow (e.g., if the firewall service node determines to not allow the data message flow).
[0084] In some embodiments, the connection tracker record for the flow identified in the flow programming instructions is updated during processing of the next data message in the data message flow. For example, in some embodiments, a flow program generation value (e.g., flow program gen) is updated (e.g., incremented) each time a flow programming instruction is added to the flow programming table to indicate that the flow programming instruction has been received and that state information generated using a previous flow program generation value can be out of date. Upon identifying a connection tracker record for a particular data message, the flow programming table is queried to see if the connection tracker record must be updated based on flow programming instructions contained in the flow programming table if the flow program generation value does not equal a current value. Reference is made to Figure 13 The use of the flow program generation value is discussed in more detail.
[0085] If process 300 determines (at 366) that there are no flow programming instructions, or after updating the flow programming table, the data message is provided (at 380) to a routing function (e.g., a routing function implemented by an edge forwarding element) for forwarding to the destination. In some embodiments, the original data message header set is carried through the service path in the service metadata. In other embodiments, the original header value set is stored in a buffer of the edge device and is restored after receiving the data message from the last hop in the service path. Those of ordinary skill in the art will appreciate that operations 366 and 375 are performed in parallel with operation 380 in some embodiments, as they are not dependent on each other.
[0086] In some embodiments, service classification operations are provided in a virtual network environment. In some embodiments, the virtual network environment is equivalent to the virtual network environment described in U.S. Patent 9,787,605, which is incorporated by reference herein. A basic introduction to the virtual network environment is given here and more details are provided in the above-referenced patent.
[0087] Figure 4 A logical network 400 with two tiers of logical routers is conceptually illustrated. As shown, the logical network 400 includes at the tier-3 level an availability zone logical gateway router (AZG) 405, several virtual private cloud logical gateway routers (VPCs) 410-420 for logical networks implemented in the availability zone. The AZG 405 and VPCGs 410-420 are sometimes referred to as tier-0 (TO) and tier-1 (Tl) routers, respectively, to reflect the hierarchical relationship between the AZG and VPCGs. The first virtual private cloud gateway 410 has attached two logical switches 425 and 430, with one or more data compute nodes coupled to each logical switch. For simplicity, only the logical switches attached to the first VPCG 410 are shown, although the other VPCGs 415-420 will typically have logical switches (to which data compute nodes are coupled) attached. In some embodiments, the availability zone is a data center.
[0088] In some embodiments, any number of VPCGs can be attached to an AZG, such as the AZG 405. Some data centers can have only a single AZG, with all VPCs implemented in the data center attached to the AZG, while other data centers can have multiple AZGs. For example, a large data center can wish to use different AZG policies for different VPCs, or can have too many different VPCs to attach all of the VPCs to a single AZG. The partial routing table of an AZG includes routes for all logical switch domains of its VPCGs, so attaching multiple VPCGs to an AZG creates several routes for each VPCG based on the subnets attached to the VPCG. As shown, the AZG 405 provides connectivity to an external physical network 435; some embodiments allow only the AZG to provide such connectivity, so that the data center (e.g., availability zone) provider can manage the connectivity. While part of the logical network 400, each of the separate VPCGs 410-420 is independently configured (although a single tenant can have multiple VPCGs, if so chosen).
[0089] Figure 5 One possible management plane view of the logical network 400 is shown, in which both the AZG 405 and the VPCGs 410 include centralized components. In this example, DR is used to distribute the routing aspects of the AZG 405 and VPCGs 410. However, because the configuration of the AZG 405 and VPCGs 410 includes providing stateful services, the management plane view (and thus, the physical implementation) of the AZG and VPCGs includes active and standby service routers (SRs) 510-520 and 545-550 for these stateful services.
[0090] Figure 5 A management plane view 500 of logical topology 400 is shown when VPCG 410 has centralized components (e.g., because stateful services are defined for the VPCG that cannot be distributed). In some embodiments, stateful services such as firewalls, NAT, load balancing, etc. are provided only in a centralized manner. However, other embodiments allow some or all such services to be distributed. For simplicity, only the details of the first VPCG 410 are shown; other VPCGs can have the same defined components (DR, transit LS, and two SRs), or just the DR if no stateful services requiring SRs are provided. AZG 405 includes a DR 505 and three SRs 510-520 connected together by a transit logical switch 525. In addition to the transit logical switch 525 within the AZG 405 implementation, the management plane defines a separate transit logical switch 530-540 between each VPCG’s and AZG’s DR 505. In the case of a VPCG that is fully distributed, the transit logical switch 530 connects to the DR that implements the VPCG’s configuration. Thus, as described in U.S. Patent 9,787,605, a packet sent by a data compute node attached to logical switch 425 to a destination in an external network will be processed through the pipeline of logical switch 425, the VPCG’s DR, transit logical switch 530, AZG 405’s DR 505, transit logical switch 525, and one of SRs 510-520. In some embodiments, the presence and definition of transit logical switches 525 and 530-540 are hidden from the user (e.g., administrator) configuring the network through an API, with the possible exception of troubleshooting purposes.
[0091] Figure 5 The shown partially centralized implementation of VPCG 410 includes a DR 560 and two SRs 545 and 550 to which logical switches 425 and 430 are attached. As in the AZG implementation, the DR and two SRs each have an interface to a transit logical switch 555. In some embodiments, this transit logical switch serves the same purpose as switch 525. For VPCGs, some embodiments implement the SRs in an active-standby manner, with one SR designated as active and the other as standby. Thus, as long as the active SR is running, a packet sent by a data compute node attached to one of logical switches 425 and 430 will be sent to the active SR rather than the standby SR.
[0092] The figure above shows a management plane view of a logical router in some embodiments. In some embodiments, an administrator or other user provides a logical topology (as well as other configuration information) through an API. This data is provided to the management plane, which defines the implementation of the logical network topology (e.g., by defining DRS, SRS, transit logical switches, etc.). In addition, in some embodiments, the user associates each logical router (e.g., each AZG or VPCG) with a set of physical machines (e.g., a predefined group of machines in a data center) to facilitate deployment. For a purely distributed router, the set of physical machines is not important, as the DR is implemented across managed forwarding elements that reside on hosts as well as data compute nodes that are connected to the logical network. However, if the logical router implementation includes SRs, these SRs will each be deployed on a particular physical machine. In some embodiments, the physical machine group is a set of machines designated for hosting SRs (as opposed to user VMs or other data compute nodes attached to logical switches). In other embodiments, the SRs are deployed on machines with user data compute nodes.
[0093] In some embodiments, the user definition of a logical router includes a particular number of uplinks. As described herein, an uplink is a northbound interface of a logical router in a logical topology. For a VPCG, its uplink connects to an AZG (in general, all uplinks connect to the same AZG). For an AZG, its uplink connects to an external router. Some embodiments require that all uplinks of an AZG have the same external router connectivity, while other embodiments allow uplinks to connect to different sets of external routers. Once the user has selected a machine group for a logical router, then if the logical router requires SRs, the management plane assigns each uplink of the logical router to a physical machine in the selected machine group. The management plane then creates an SR on each machine to which an uplink is assigned. Some embodiments allow multiple uplinks to be assigned to the same machine, in which case the SR on the machine has multiple northbound interfaces.
[0094] As described above, in some embodiments, an SR can be implemented as a virtual machine or other container, or as a VRF context (e.g., in the case of a DPDK-based SR implementation). In some embodiments, the choice of SR implementation can be based on the services selected for the logical router and which type of SR best provides these services.
[0095] Furthermore, in some embodiments, the management plane creates transit logical switches. For each transit logical switch, the management plane assigns a unique VNI to that logical switch, creates ports on each SR and DR connected to the transit logical switch, and assigns IP addresses to any SR and DR connected to the logical switch. Some embodiments require that the subnet assigned to each transit logical switch be unique within a logical L3 network topology (e.g., network topology 400) with multiple VPCGs, and each VPCG can have its own transit logical switch. That is, in Figure 5 In this implementation, each of the following logic switches—525 within the AZG, 530-540 between the AZG and VPCG, and 520 (and any other VPCG implementation's transition logic switch)—requires a unique subnet. Furthermore, in some embodiments, the SR may need to initiate connections to VMs in the logical space, such as HA agents. To ensure proper return operations, some embodiments avoid using link-local IP addresses.
[0096] Some embodiments impose various restrictions on the connectivity of logical routers in a multi-tier configuration. For example, while some embodiments allow any number of logical router tiers (e.g., an AZG tier connected to an external network, along with multiple VPCG tiers), other embodiments allow only a two-tier topology (one VPCG tier connected to an AZG). Furthermore, some embodiments allow each VPCG to connect to only one AZG, and each logical switch created by the user (i.e., a non-transfer logical switch) is only allowed to connect to one AZG or one VPCG. Some embodiments also add the restriction that the southbound ports of logical routers must each be in a different subnet. Therefore, two logical switches connected to the same logical router may not have the same subnet. Finally, some embodiments require that different uplinks of the AZG must exist on different gateway machines. It should be understood that some embodiments do not include any of these requirements, or may include various different combinations of these requirements.
[0097] Figure 6 Conceptually demonstrated Figure 5 The diagram illustrates the physical implementation of the management plane structure of a two-layer logical network, where both VPCG 410 and AZG 405 include both SR and DR. It should be understood that this diagram only shows the implementation of VPCG 410 and not many other VPCGs that can be implemented on many other hosts and whose SRs can be implemented on other gateway machines.
[0098] The figure assumes that there are two VMs attached to each of two logical switches 425 and 430 that reside on four physical hosts 605-620. Each of these hosts includes a management forwarding element (MFE) 625. In various different embodiments, these MFEs can be either flow-based forwarding elements (e.g., Open vSwitch) or code-based forwarding elements (e.g., ESX) or a combination of the two. These different types of forwarding elements implement the various logical forwarding elements in different ways, but in each case, they perform a pipeline for each logical forwarding element that can need to process a packet.
[0099] Thus, as Figure 6 indicated, the MFEs 625 on the physical hosts include configurations for implementing both logical switches 425 and 430 (LSA and LSB), the DR 560 and transit logical switch 555 for the VPCG 410, and the DR 505 and transit logical switch 525 for the AZG 405. However, when the VPCG of the data compute nodes that reside on the hosts does not have a centralized component (i.e., SR), some embodiments implement only the distributed components of the AZG on the host MFEs 625 (those that are coupled to the data compute nodes). As described below, northbound packets sent from the VMs to the external network will be handled by their local (first hop) MFE until the transit logical switch pipeline specifies that the packet is to be sent to the SR. If the first SR is part of the VPCG, the first hop MFE will not perform any AZG processing, so there is no need for the centralized controller(s) to push the AZG pipeline configuration to those MFEs. However, since one of the VPCGs 415-420 can not have a centralized component, some embodiments always push the distributed aspects of the AZG (the DR and the transit LS) to all MFEs. Other embodiments only push the configuration for the AZG pipeline to the MFEs that are also receiving the configuration for the fully distributed VPCGs (those without any SRs).
[0100] In addition, Figure 6 The physical implementation shown includes four physical gateway machines 630-645 (also referred to as edge nodes in some embodiments), to which the SRs of the AZG 405 and the VPCG 410 are assigned. In this case, the administrator that configures the AZG 405 and the VPCG 410 selects the same set of physical gateway machines for the SRs, and the management plane assigns one of the SRs for both logical routers to the third gateway machine 640. As shown, the three SRs 510-520 for the AZG 405 are each assigned to a different gateway machine 630-640, while the two SRs 545 and 550 for the VPCG 410 are each assigned to a different gateway machine 640 and 645.
[0101] This figure illustrates the SR as separate from the MFE 650 running on the gateway machine. As noted above, different embodiments can implement the SR differently. Some embodiments implement the SR as a VM (e.g., when the MFE is a virtual switch integrated into the virtualization software of the gateway machine), in which case the SR processing is performed outside of the MFE. On the other hand, some embodiments implement the SR as a VRF within the data path of the MFE (when the MFE uses DPDK for data path processing). In either case, the MFE treats the SR as part of the data path, but in the case where the SR is a VM (or other data compute node), sends the packet to the separate SR to be processed by the SR pipeline (which can include the performance of various services). As with the MFE 625 on the host, the MFE 650 of some embodiments is configured to perform all of the distributed processing components of the logical network.
[0102] Figure 7 and Figure 10 A set of logical processing operations related to the availability zone (TO) and VPC (Tl) logical routers are shown. Figure 7 Logical processing operations of the availability zone (TO) logical router components included in the edge data path 710 performed by the edge appliance 700 for a data message are shown. In some embodiments, the edge data path 710 is performed by an edge forwarding element of the edge appliance 700. The edge data path 710 includes logical processing stages for a number of operations, including an availability zone (TO) service (e.g., centralized) router 730 and an availability zone (TO) distributed router 740. As shown, the TO SR 730 invokes a set of service insertion and service transmission layer modules 735 to perform service classification operations (or service insertion (SI) classification operations). In some embodiments, the edge data path 710 includes logical processing operations for VPC (Tl) service and distributed routers. For the availability zone (TO) SR, in some embodiments, the VPC (Tl) SR invokes a set of service insertion and service transmission layer modules to perform service classification operations.
[0103] The service insertion layer and service transport layer module 735 includes a service insertion pre-processor 720, a connection tracker 721, a service transport layer module 722, a logical switch service plane processor 723, a service plane layer 2 interface 724, a service routing function 725, a line-in-the- wire (BIW) interface 726, a virtual tunnel interface 727, a service insertion post-processor 728, and a flow programming table 729. In some embodiments, the service insertion pre-processor 720 performs process 100 to determine the service type and forwarding information for a received data message. In some embodiments, the service transport layer module 722 performs process 300 to direct the data message to the appropriate service node to perform the required service and return the data message to the T0 SR 730 for routing to the destination of the data message.
[0104] The functions of the modules of the service insertion layer and service transport layer 735 are described in more detail below with respect to Figure 11 , 12 , 18, and 19. In some embodiments, the service insertion pre-processor 720 is invoked for data messages received on each of a set of interfaces of an edge forwarding element that is not connected to a service node. The service insertion (SI) pre-processor 720 applies service classification rules (e.g., service insertion rules) that are defined (e.g., by a provider or by a tenant with multiple VPCGs after a single AZG) for application at the T0 SR 730. In some embodiments, each service classification rule is defined in terms of a flow identifier that identifies a data message flow that requires a service insertion operation (e.g., service by a set of service nodes). In some embodiments, the flow identifier includes a set of data message attributes (e.g., any one or a combination of a set of header values (e.g., a 5-tuple) that define the data message flow), a set of context data associated with the data message, or a value derived from the set of header values or the set of context data (e.g., a hash of the set of header values or an application identifier of an application associated with the data message).
[0105] In some embodiments, the interfaces connected to the service nodes are configured to mark data messages returned to the edge forwarding element as serviced so that they are not provided to the SI pre-processor 720 again. After the SI pre-processor 720 performs the service classification operation, the results of the classification operation are passed to the service transport layer module 722 for forwarding the data message to the set of service nodes that provide the required set of services.
[0106] After the service node(s) process the data message, the serviced data message is returned to the service transport layer module 722 for post-processing at the SI post-processor 728 and then returned to the T0 SR 730 for routing. The T0 SR 730 routes the data message and provides the data message to the T0 DR 740. In some embodiments, the T0 SR 730 is connected to the T0 DR 740 through a transit logical switch (not shown), as described above with reference to Figure 5 and 6 The T0 SR 730 and the T0 DR 740 perform logical routing operations to forward the incoming data message to the correct virtual private cloud gateway and ultimately to the destination compute node. In some embodiments, the logical routing operations include identifying an egress logical port of a logical router to forward the data message to a next hop based on a destination IP address of the data message.
[0107] In some embodiments, the edge data path 710 also includes logical processing stages for T1 SR and T1 DR operations and the T0 SR 730 and the T0 DR 740. Some embodiments insert a second service classification operation performed by a service insertion layer and a set of service transport layer modules invoked by the T1 SR. The SI pre-processor invoked by the VPCG applies service classification rules defined for the VPCG (e.g., service insertion rules) (e.g., service insertion rules for a particular VPC logical network behind the VPCG). In some embodiments, the service classification rules specific to the VPCG are included in the same set of rules as the service classification rules specific to the AZG and distinguished by a logical forwarding element identifier. In other embodiments, the service classification rules specific to the VPCG are stored in a separate service classification rules storage device or database used by the SI pre-processor invoked by the VPCG.
[0108] The SI pre-processor invoked by the VPCG performs the same operations as the SI pre-processor 720 to identify data messages that require a set of services and forwarding information and service types for the identified data messages. For the SI pre-processor 720, the SI pre-processor performs service classification operations and, after providing services, returns the data message to a logical processing stage of the T1 SR. The T1 SR routes the data message and provides the data message to the T1 DR. In some embodiments, the T1 SR is connected to the T1 DR through a transit logical switch (not shown), as described above with reference to Figure 5 and Figure 6 The T1 SR and the T1 DR perform logical routing operations to forward the incoming data message to the destination compute node through a set of logical switches, as described above with reference to Figure 5 and Figure 6described. In some embodiments, the logical routing operation includes identifying an egress logical port for a logical router used to forward the data message to a next hop based on a destination IP address of the data message. Multiple T1 SRs and DRs can be identified by the T0 DR 740, and in some embodiments, the above discussion applies to each T1 SR / DR in the logical network. Thus, one of ordinary skill in the art will appreciate that, in some embodiments, the edge device 700 performs edge processing for multiple tenants, each tenant sharing the same set of AZG processing stages but having its own VPCG processing stage.
[0109] For outgoing messages, the edge data path is similar, but in some embodiments, the edge data path will include T1 and T0 DR components, respectively, only when the source compute node is executing on the edge device 700 or the T1 SR is executing on the edge device 700. Otherwise, the host of the source node (or the edge device executing the T1 SR) will perform the logical routing associated with the T1 / T0 DR. In addition, for outgoing data messages, the data message is logically routed by the SR (e.g., T0 SR 730) prior to invoking the service insertion layer and service transport layer modules. The functions of the service insertion layer and service transport layer modules are similar to the forward (e.g., incoming data messages discussed above), and will be discussed in greater detail below. For data messages requiring services, the serviced data messages are returned to the SR (e.g., T0 SR 730) to be sent over the interface identified by the logical routing processing.
[0110] Figure 8 A TX SR 1130 is shown acting as a source of traffic on a logical service forwarding element 801 (e.g., a logical service switch). The logical service forwarding element (LSFE) is implemented by a set of N software switches 802 executing on N devices. The N devices include a set of devices on which service nodes (e.g., service virtual machines 806) execute. The TX SR 1130 sends data messages requiring service by the SVM 806 through SIL and STL modules 1120 and 1122, respectively. The SI layer module 1120 identifies the forwarding information required to send the data message to the SVM 806 on the LSFE, as discussed above with reference to Figure 1 and will be discussed below with reference to Figure 11 and Figure 12 The forwarding information and data message are then provided to the STL module 1122 for processing, facilitating delivery on the LSFE to the SVM 806 using port 810. Because the SVM 806 executes on a separate device, the data message sent out from the software switch port 815 is encapsulated by an encapsulation processor 841 for transmission on an intermediate network.
[0111] The encapsulation processor 842 then decapsulates the encapsulated data message and provides it to the port 816 for delivery to the SVM 806 through its STL module 826 and SI agent 814. Return data messages pass through the modules in the reverse order. The operation of the STL module 826 and SI agent 814 are discussed in greater detail in U.S. Patent Application 16 / 444,826.
[0112] Figure 9 A service path is shown that includes two service nodes 906 and 908 accessed by the TX SR 1130 to the LSFE 801. As shown, the TX SR 1130 sends a reference Figure 8 The first data message is described. The data message is received by the SVM 1 906, which provides the first service in the service path and forwards the data message to the next hop in the service path, in this case the SVM 2 908. The SVM 2 908 receives the data message, provides the second service, and forwards the data message to the TX SR 1130, which is identified in some embodiments as the next hop in the service path. In other embodiments, the TX SR 1130 is identified as the source of the serviced data message that returns to it after the last hop (e.g., the SVM 2 908) has provided its service. As to the Figure 8 Additional details of the processing at each module are explained in greater detail in U.S. Patent Application 16 / 444,826.
[0113] Figure 10 A second embodiment is shown that includes two edge devices 1000 and 1005 that respectively perform the AZ gateway data path 1010 and the VPC gateway data path 1015. Figure 7 And Figure 10 The functions of like numbered elements appearing in Figure 7 And Figure 10 The difference between the embodiments is that in Figure 10 the VPC edge data path (the T1 SR 1060 as well as the service insertion layer and service transport layer modules 1065) is performed in the edge device 1005 rather than the edge device 1000. As noted above, in some embodiments, the distributed router is performed on any device that performs the previous processing step.
[0114] In some embodiments, the edge forwarding element is configured to provide services using the service logic forwarding element as a transport mechanism, as referenced in Figure 11The edge forwarding element is configured to connect different sets of virtual interfaces of the edge forwarding element to different network elements of the logical network using different transport mechanisms. For example, a first set of virtual interfaces is configured to connect to a set of forwarding elements inside the logical network using a set of logical forwarding elements that connect source and destination machines of traffic of the logical network. In some embodiments, traffic received on the first set of interfaces is forwarded by the edge forwarding element to a next hop toward the destination without being returned to the forwarding element from which it was received. A second set of virtual interfaces is configured to connect to a set of service nodes to provide services for data messages received at the edge forwarding element.
[0115] Each connection established for the second set of virtual interfaces can use a different transport mechanism, such as a service logical forwarding element, a tunneling mechanism, and an in-line appliance mechanism, and in some embodiments, some or all of the transport mechanisms are used to provide data messages to service nodes as discussed below with reference to Figure 11 、 12 , 18, and 19. Each virtual interface in a third set of virtual interfaces (e.g., a subset of the second set) is configured to connect to a logical service forwarding element that connects the edge forwarding element to at least one of the set of internal forwarding elements as discussed below with reference to Figures 30 to 32A -B. The virtual interfaces are configured to be used to (1) receive data messages from the at least one internal forwarding element that are to be serviced by at least one of the set of service nodes, and (2) return the serviced data messages to the set of internal forwarding elements.
[0116] In some embodiments, the transport mechanisms include a logical service forwarding element that connects the edge forwarding element to a set of service nodes that each provide a service of a set of services. In selecting the set of services, a service classification operation of some embodiments identifies a chain of multiple service operations that must be performed on the data message. In some embodiments, the service classification operation includes selecting a service path that provides the multiple services for the identified chain of services. After selecting the service path, the data message is sent along the selected service path to provide the services. Once the services are provided, the data message is returned by the last service node in the service path that performed the last service operation to the edge forwarding element, and the edge forwarding element performs a forwarding operation to forward the data message as further discussed below with reference to Figure 11 and 12 .
[0117] Figure 11A set of operations performed by a service router at T0 or T1 (e.g., TX SR 1130) for a set of service insertion layers and a set of service transport layer modules 1135 invoked for a first data message 1110 in a data message flow that requires services from a set of service nodes that define a service path are shown. Figure 11 TX SR 1130 and service insertion layer (SIL) and service transport layer (STL) modules 1135 are shown. In some embodiments, TX SR 1130 and SIL and STL modules 1135 represent the functionality of a centralized service router and SIL and STL modules at either T0 and T1. In some embodiments, the T0 and T1 data paths share the same set of SIL and STL modules, while in other embodiments, the T0 and T1 data paths use separate SIL and STL modules. SIL and STL modules 1135 include service insertion pre-processor 1120, connection tracker 1121, service layer transport modules 1122, logical switch service plane processor 1123, service plane layer 2 interface 1124, service insertion post-processor 1128, and flow programming table 1129.
[0118] Data message 1110 is received at an edge device and provided to edge TX SR 1130. In some embodiments, TX SR 1130 receives data messages at an uplink interface of TX SR 1130 or at a virtual tunnel interface. In some embodiments, SI pre-processor 1120 is invoked at different processing operations for the uplink interface and virtual tunnel interface (VTI). In some embodiments, invocation for data messages received at the uplink interface and virtual tunnel interface are implemented by different components of the edge device. For example, in some embodiments, SI pre-processor 1120 is invoked for data messages received at the uplink interface by the NIC driver as part of the standard data message pipeline, while SI pre-processor 1120 is invoked for data messages received at the VTI after (before) the decapsulate and decrypt (encrypt and encapsulate) operations as part of a separate VTI processing pipeline. In some embodiments where SI pre-processor 1120 is implemented differently for the uplink and VTI, the same connection tracker is used to maintain consistent state for each data message even as it traverses the VTI and uplink.
[0119] The SI preprocessor 1120 performs a set of operations analogous to those of the process 100. The SI preprocessor 1120 performs a lookup in the connection tracker storage 1121 to determine whether there is a connection tracker record for the data message stream to which the data message belongs. As described above, this determination is based on a flow identifier that includes or is derived from flow attributes (e.g., header values, context data, or values derived from header values and alternatively or in combination from context data). In the illustrated example, the data message 1110 is the first data message in a data message stream, and there is no connection tracker record identified for the data message stream to which the data message 1110 belongs. The connection tracker storage lookup is equivalent to the operation 120 of the process 100, and if a most recent connection tracker record has been found, the SI preprocessor 1120 will forward information in the identified connection tracker record to the LR-SR as in operations 120, 125, and 170 of the process 100.
[0120] Since no connection tracker record is found in this example, the SI preprocessor 1120 performs a lookup in the service insertion rule storage 1136 to determine whether there are any service insertion (service classification) rules that apply to the data message. In some embodiments, SI rules for different interfaces are stored in the SI rule storage 1136 as separate sets of rules that are queried based on an incoming interface identifier (e.g., as metadata stored in the edge device’s buffer for the incoming interface UUID). In other embodiments, SI rules for different interfaces are stored as a single set of rules with potential matching rules that are checked to see if they apply to the interface on which the data message was received. As will be discussed below, the set(s) of SI rules are received from a controller that generates the rule sets based on policies defined at the network manager (by an administrator or by the system). In some embodiments, the SI rules 1145 in the SI rule storage 1136 are specified according to flow attributes that identify the data message stream to which the rule applies and a service action. In the illustrated example, the service action is a redirect to a UUID that identifies a service type and forwarding information.
[0121] Assuming the lookup in the SI rules store 1136 results in identification of service insertion rules applicable to the data message 1110, the process queries the policy table 1137 using the UUID identified from the service applicable insertion rules. In some embodiments, the UUID is used to simplify management of service insertion such that if a particular service node in a set of service nodes fails, then there is no need to update each individual rule specifying the same set of service nodes, but rather the set of service nodes associated with the UUID can be updated, or the selection (e.g., load balancing) operation can be updated for the UUID. The current example shows a UUID that identifies a service chain identifier associated with multiple service paths identified by a set of multiple service path identifiers (SPIs) and a set of selection metrics. The selection metrics can be selection metrics for a load balancing operation that is any one of: a round robin mechanism, a load based selection operation (e.g., select the service node with the lowest current load), or a distance based selection operation (e.g., select the closest service node as measured by the selected metric). In some embodiments, the set of service paths is a subset of all possible service paths for the service chain. In some embodiments, the subset is selected by a controller that allocates different service paths to different edge devices. In some embodiments, allocating service paths to different edge devices provides a first level of load balancing across service nodes.
[0122] Once a service path is selected, the SI pre-processor 1120 identifies forwarding information associated with the selected service path by performing a lookup in a forwarding table 1138. The forwarding table 1138 stores forwarding information for service paths (e.g., a MAC address of a first hop in the service path). In some embodiments, the forwarding information includes a service index that indicates a length of the service path (i.e., a number of service nodes included in the service path). In some embodiments, the forwarding information also includes a time to live (TTL) value that indicates a number of service nodes in the service path. In other embodiments, the next hop MAC address, service index, and TTL value are stored in the policy table 1137 with the SPI, and the forwarding table 1138 is unnecessary.
[0123] In some embodiments, selecting a service path for a forward data message flow includes selecting a corresponding service path for a reverse data message flow. In such embodiments, the forwarding information is determined at this time for each direction. In some embodiments, the service path for the reverse data message flow includes the same service nodes as the service path for the forward data message flow, but in reverse order through the service nodes. In some embodiments, the service path for the reverse data message flow includes the same service nodes as the service path for the forward data message flow, but in reverse order through the service nodes when at least one service node modifies the data message. For some data message flows, the service path for the reverse data message flow is the same as the service path for the forward flow. In some embodiments, an SR can be used as a service node that provides L3 routing services, and is identified as the last hop for each service path. In some embodiments, the SR L3 routing service node is also the first hop for each service path to ensure that traversing the service path in reverse order will end at the SR, and the SR performs the first hop processing for the service path as a service node.
[0124] Once the service path is selected and the forwarding information is identified, a connection tracker record is created for the forward and reverse flows and provided to the connection tracker storage 1121. In some embodiments, the service post-processor 1128 is queried for a state value (e.g., a flow program generation value “flow_prog_gen”) that indicates the current state of the set of service nodes (e.g., the set of service nodes associated with the identified service type). As described below, the connection tracker record includes the forwarding information (e.g., SPI, service index, next hop MAC address, and service insertion rule identifier for the service insertion rule identified as matching the attributes of the data message 1110) used to process subsequent data messages in the forward and reverse data message flows. In some embodiments, the connection tracker record also includes the flow program generation value to indicate the current flow program generation value at the time the connection tracker record is created to facilitate comparison with the then-current value at the time of subsequent data messages in the data message flows for which the record is created.
[0125] The data message 1152 is then provided to the STL module 1122 along with the forwarding information 1151. In this example, the forwarding information for the data message that requires services provided by the service chain includes service metadata (SMD) that, in some embodiments, includes any or all of a service chain identifier (SCI), SPI, service index, TTL value, and direction value. In some embodiments, the forwarding information also includes the MAC address of the next hop and a service insertion type identifier for identifying the data message as being transmitted using the logical service forwarding element transport mechanism.
[0126] As shown, the STL module 1122 provides the data message 1153 along with an encapsulation header 1154, which in some embodiments includes SMD and liveness attributes that indicate that the L3 routing service node is still operational for the logical switch service plane processor 1123, which prepares the data message for sending to the service plane L2 interface 1124 based on the information included in the encapsulation header 1154. In some embodiments, instead of an encapsulation header, forwarding information is sent or stored as separate metadata, which in some embodiments includes SMD and liveness attributes that indicate that the L3 routing service node is still operational. The function of the logical switch service plane processor 1123 is similar to the port proxy described in U.S. Patent Application 16 / 444,826, filed June 18, 2019. As shown, the logical switch service plane processor 1123 removes the header 1154 and records the SMD and next hop information. The data message is then provided to the service plane L2 interface 1124 (e.g., a software switch port associated with the logical service forwarding element).
[0127] The data message is then encapsulated for delivery by an interface (e.g., a port or virtual tunnel endpoint (VTEP)) of the software forwarding element to the first service node in the service path to produce 1157. In some embodiments, the SMD is a modified set of SMDs that enables the original data 1110 message to be reconstructed when the serviced data message is returned to the logical switch service plane processor 1123. In some embodiments, encapsulation is only needed when the next hop service node is executing on another device, so that the encapsulated data message 1157 can traverse an intermediate network fabric.
[0128] In some embodiments, the encapsulation encapsulates the data message with a cover header to produce the data message 1157. In some embodiments, the cover header is a Genve header that stores the SMD and STL attributes in one or more TLVs thereof. As described above, the SMD attributes in some embodiments include an SCI value, an SPI value, an SI value, and a service direction. Other encapsulation headers are described in U.S. Patent Application 16 / 444,826, filed June 18, 2019. The illustrated data path for the data message 1110 assumes that the first service node in the service path is on an external host (not a host of the edge device). Conversely, if the edge device hosts the next service node in the service path, the data message will not need to be encapsulated, but instead is sent on the logical service forwarding plane using the SVNI associated with the logical service plane and the MAC address of the next hop service node to the next service node.
[0129] If flow programming instructions are included in the encapsulation header 1158, the flow programming instructions 1159 are provided to the flow programming table 1129, and the flow programming version value is updated (e.g., incremented). In some embodiments, the flow programming instructions in the flow programming table 1129 include a new action (e.g., pf_value) that indicates that the subsequent data messages should be dropped, allowed, or a new service path identified to skip a particular service node (e.g., a firewall that has determined to allow the connection) while traversing other service nodes in the original service path. Reference will be made to Figure 13 The use of the flow programming version value is further discussed.
[0130] Figure 12 A set of operations performed by the service router at T0 or T1 (e.g., TX SR 1130) for a set of service insertion layer and service transport layer modules 1135 invoked for a data message 1210 in a data message flow that requires services from a set of service nodes that define a service path is shown. Figure 12 TX SR 1130 and service insertion layer (SIL) and service transport layer (STL) modules 1135 are shown. In some embodiments, TX SR 1130 and SIL and STL modules 1135 represent the functionality of a centralized service router and SIL and STL modules at either T0 and T1. In some embodiments, the T0 and T1 data paths share the same set of SIL and STL modules, while in other embodiments, the T0 and T1 data paths use separate SIL and STL modules. SIL and STL modules 1135 include service insertion pre-processor 1120, connection tracker 1121, service layer transport modules 1122, logical switch service plane processor 1123, service plane layer 2 interface 1124, service insertion post-processor 1128, and flow programming table 1129.
[0131] Data message 1210 is received at the edge device and provided to edge TX SR 1130. In some embodiments, TX SR 1130 receives data messages at an uplink interface of TX SR 1130 or at a virtual tunnel interface. In some embodiments, SI preprocessor 1120 is invoked at different processing operations for the uplink interface and the virtual tunnel interface (VTI). In some embodiments, invocation for data messages received at the uplink interface and the virtual tunnel interface is implemented by different components of the edge device. For example, in some embodiments, SI preprocessor 1120 is invoked for data messages received at the uplink interface by a NIC driver that is part of a standard data message pipeline, while SI preprocessor 1120 is invoked for data messages received at the VTI after (before) decapsulation and decryption (encryption and encapsulation) operations that are part of a separate VTI processing pipeline. In some embodiments where SI preprocessor 1120 is implemented differently for the uplink and the VTI, the same connection tracker is used to maintain consistent state for each data message, even as it passes through the VTI and the uplink.
[0132] SI preprocessor 1120 performs a similar set of operations as operations of process 100. SI preprocessor 1120 performs a lookup in connection tracker storage 1121 to determine whether there is a connection tracker record for the data message stream to which the data message belongs. As described above, the determination is based on a flow identifier that includes or is derived from flow attributes (e.g., header values, context data, or values derived from header values and alternatively or in combination from context data). In the illustrated example, data message 1210 is a data message in a data message stream that has a connection tracker record in connection tracker storage 1121. The connection tracker storage lookup begins with operation 120 of process 100. Figure 12 An additional set of operations is illustrated that will serve as an example of the operations discussed in Figure 13
[0133] Some embodiments provide a method of performing a stateful service that tracks state changes of service nodes to update a connection tracker record as necessary. At least one global state value indicative of a state of a service node is maintained at an edge device. In some embodiments, different global state values are maintained for service chain service nodes (SCSNs) and Layer 2 line card insert service nodes (L2 SNs). The method generates a record in a connection tracker storage device that includes a current global state value as a flow state value for a first data message in a data message flow. Each time a data message is received for the data message flow, the stored state value (i.e., the flow state value) is compared to the relevant global state value (e.g., the SCSN state value or the L2 SN state value) to determine whether the stored action has been updated.
[0134] After a global state value associated with a flow changes, the global state value and the flow state value do not match, and the method checks a flow programming table to determine whether the flow has been affected by flow programming instruction(s) that caused the global state value to change (e.g., increment). In some embodiments, the instructions stored in the flow programming table include a data message flow identifier and an updated action (e.g., drop, allow, update selected service path, update next hop address). If the data message flow identifier stored in the flow programming table does not match the current data message flow identifier, the flow state value is updated to the current global state value, and the action stored in the connection tracker record is used to process the data message. However, if at least one of the data message flow identifiers stored in the flow programming table matches the current data message flow identifier, the flow state value is updated to the current global state value, and the action stored in the connection tracker record is updated to reflect the execution of the instructions with the matching flow identifier stored in the flow programming table, and the updated action is used to process the data message.
[0135] Figure 13 A process 1300 for verifying or updating a connection tracker record identified for a data message flow is conceptually illustrated. In some embodiments, the process 1300 is performed by an edge forwarding element executing on an edge device. In Figure 12 In the example of FIG. 11, the process is performed by the SI preprocessor 1120. The process 1300 begins by identifying (at 1310) a connection tracker record for a data message received at the edge forwarding element. A flow programming version value is stored in the connection tracker record that reflects the flow programming version value at the time the connection tracker record was generated. Alternatively, in some embodiments, the flow programming version value reflects the flow programming version value at the last update of the connection tracker record. The connection tracker record stores forwarding information for the data message flow. However, the stored information can be outdated if there are flow programming instructions for the data message flow.
[0136] For some data message flows, a previous data message in the data message flow will have been received that includes a set of flow programming instructions. In some embodiments, the data message is a served data message that has been served by a set of service nodes in a service path, and the set of flow programming instructions is based on flow programming instructions from the set of service nodes in the service path. In some embodiments, the set of flow programming instructions includes flow programming instructions for both forward and reverse data message flows affected by the flow programming instructions. In some embodiments, the forward and reverse flow IDs are the same, and a direction bit distinguishes between forward and reverse data message flows in the connection tracker record.
[0137] The set of flow programming instructions is recorded in a flow programming table, and a flow programming version value is updated (incremented) at the flow programming table to reflect that flow programming instructions have been received that can require updating information in at least one connection tracker record. In some embodiments, the flow programming instructions are based on any one of the following events: a failure of a service node, an identification of a service node that is no longer needed as part of a service path, a decision to drop a particular data message in a data message flow, or a decision to drop a particular data message flow. Based on the event, the flow programming instructions include forwarding information that specifies a different service path (based on a service node failure or an identification of a service node that is no longer needed) or a new action (e.g., a pf_value) than previously selected (based on the event) for a next data message (based on a decision to drop a particular data message in a data message flow) or for a data message flow (based on a decision to drop a particular data message flow). In some embodiments, the flow programming table stores records related to individual flows as well as records of failed service nodes (or service paths to which they belong) for use in determining available service paths during service path selection operations. The records stored for individual flows persist until they are executed, which in some embodiments occurs upon receipt of a next data message of a data message flow, as described below.
[0138] The process 1300 then determines (at 1320) whether the flow programming generation value is current (i.e., not equal to the flow programming version value stored by the flow programming table). In the described embodiments, determining (at 1320) whether the flow programming version value (e.g., flow_prog_gen or BFD_gen) is current includes a query of the flow programming table 1129 that includes only the flow programming version value to perform a simple query operation to determine whether a further, more complex query must be performed. If the flow programming version value is current, then the action stored in the connection tracker record can be used to forward the data message to a service node to provide the required service, and the process ends.
[0139] If the process 1300 determines (at 1320) that the stream programming version value is not current, the process 1300 determines (at 1330) whether there are stream programming instructions applicable to the received data message (i.e., the data message stream to which the data message belongs). In some embodiments, this second determination is made using a query (e.g., 1271) that includes the stream ID, which is used as a key to identify a stream programming record stored in the stream programming table. In some embodiments, the query also includes a service path identifier (SPI) that can be used to determine whether a service path has failed. In some embodiments, the stream programming generation value is not current because a stream programming instruction has been received that caused the stream programming version value to update (e.g., increment). In some embodiments, the stream programming instruction is related to the data message stream, while in other embodiments, the stream programming instruction is related to a different data message stream or a failed service node.
[0140] If the process 1300 determines (at 1330) that there are no relevant stream programming instructions for the received data message, the stream programming version value stored in the connection tracker record is updated (at 1340) to reflect the stream programming version value returned from the stream programming table. The data message is then processed (at 1345) based on the action stored in the connection tracker record, and the process ends. However, if the process 1300 determines (at 1330) that there are relevant stream programming instructions for the received data message, the data message is processed (at 1350) using the action in the stream programming table. In some embodiments, the determination that there are relevant stream programming instructions for the received data message is based on receiving a non-empty response to the query (e.g., 1272). The connection tracker record is then updated (at 1360) based on the query response to update the service action and the stream programming version value, and the process ends. Those of ordinary skill in the art will appreciate that operations 1350 and 1360 are performed together or in a different order without affecting functionality.
[0141] For some data messages, processing the data message according to the forwarding information based on the stream programming record and the connection tracker record includes forwarding the data message to a service transport layer module 1122, which forwards the data message along the service path identified in the forwarding information. For other data messages, processing the data message according to the forwarding information includes discarding (or allowing) the data message based on the stream programming instructions. Similar processing is performed for L2 BIW service nodes based on a bidirectional forwarding detection (BFD) version value (e.g., BFD_gen), which is a state value associated with the failure of a service node connected through an L2 BIW transport mechanism and stored in the connection tracker record at creation.
[0142] After processing the data message by the SI post-processor 1128, the served data message 1162 is provided to the TX SR 1130, marked (e.g., using a tag or metadata associated with the data message) as served, so that the SI pre-processor is not invoked a second time to classify the data message. The TX SR 1130 then forwards the data message 1163 to the next hop. In some embodiments, the served mark is maintained while forwarding the data message, while in some other embodiments, the mark is removed as part of the logical routing operations of the TX SR 1130. In some embodiments, metadata is stored for the data message indicating that the data message is served. Figure 31 An example of an embodiment in which the identification of a data message is maintained as served to avoid creating a loop from the Tl SR serving classification operation will be discussed.
[0143] Figure 14 A set of connection tracker records 1410-1430 in the connection tracker storage 1121 and a set of example flow programming records 1440-1490 in the flow programming table 1129 are shown. The connection tracker storage 1121 is shown as storing connection tracker records 1410-1430. The set of connection tracker records 1410-1430 includes a set of connection tracker records for different transport mechanisms, each including separate records for forward and reverse data message flows.
[0144] The set of connection tracker records 1410 is a set of connection tracker records for forward and reverse data message flows using a logical service plane (e.g., a logical service forwarding element) transport mechanism. In some embodiments, the set of connection tracker records 1410 includes a connection tracker record 1411 for a forward data message flow and a connection tracker record 1412 for a reverse data message flow. Each connection tracker record 1411 and 1412 includes a flow identifier (e.g., a flow ID or a flow ID’), a set of service metadata, a flow programming generation value (e.g., flow program gen), an action identifier (e.g., a pf value), and a rule ID identifying the service rule used to create the connection tracker record. In some embodiments, the flow IDs for the forward and reverse data message flows are different flow IDs based on a swap of source and destination addresses (e.g., IP and MAC addresses). In other embodiments, the different flow IDs for the forward and reverse data message flows are based on different values for the source and destination addresses as a result of network address translation provided by a service node in a set of service nodes. In some embodiments, the forward flow ID and the reverse flow ID are the same except for a bit indicating directionality. In some embodiments, the directionality bit is stored in a separate field and the forward flow ID and the reverse flow ID are the same.
[0145] In some embodiments, service metadata (SMD) includes a service path ID (e.g., SPI 1 and SPI 1'), a service index (e.g., SI, which should be the same for forward and reverse data message flows), a time-to-live (TTL), and a next hop MAC address (e.g., hoplmac and hopMmac). In the foregoing reference to Figure 3 and Figure 11 The use of SMD in processing data messages has been described. In some embodiments, SMD includes the Network Service Header (NSH) attributes of RFC (Request for Comments) 8300 of IETF (Internet Engineering Task Force). In some embodiments, SMD includes a service chain identifier (SCI) and a direction (e.g., forward or reverse) as well as SPI and SI values for processing service operations of a service chain.
[0146] In some embodiments, the rule ID is used to identify (as described in operation 223) a set of interfaces to which the rule applies by using the rule ID as a key in an applied_to storage device 1401 that stores records including a rule ID field 1402 that identifies the rule ID and an applied_to field 1403 that contains a list of interfaces to which the identified rule applies. In some embodiments, the interfaces are logical interfaces of service routers identified by UUIDs. In some embodiments, the applied_to storage device 1401 is configured by a controller that is aware of service insertion rules, service policies, and interface identifiers.
[0147] The connection tracker record set 1420 is a set of connection tracker records for forward and reverse data message flows that use a Layer 2 in-line with wire (BIW) transport mechanism. In some embodiments, the connection tracker record set 1420 includes a connection tracker record 1421 for a forward data message flow and a connection tracker record 1422 for a reverse data message flow. Each connection tracker record 1421 and 1422 includes a flow identifier (e.g., flow ID or flow ID'), an IP address (e.g., a pseudo IP address associated with an interface connected to an L2 service node), a bidirectional forwarding detection (BFD) generation value (e.g., BFD_gen) that is a state value associated with a failure of a service node connected to an LR-SR (e.g., an AZG-SR or a VPCG-SR) through an L2 BIW transport mechanism, an action identifier (e.g., pf_value), and a rule ID that identifies a service rule used to create the connection tracker record. The flow ID for the forward and reverse data message flows is the same as the flow ID described for the connection tracker record 1410. In some embodiments, the pf_value is a value that identifies whether the flow should be allowed or dropped, bypassing the service node.
[0148] The connection tracker record set 1430 is a connection tracker record set for forward and reverse data message flows using a layer 3 tunneling mechanism. In some embodiments, the connection tracker record set 1430 includes a connection tracker record 1431 for a forward data message flow and a connection tracker record 1432 for a reverse data message flow. Each connection tracker record 1431 and 1432 includes a flow identifier (e.g., flow ID or flow ID'), an IP address (e.g., an IP address of a virtual tunnel interface that connects the LR-SR to a service node), an action identifier (e.g., pf_value), and a rule ID that identifies the service rule used to create the connection tracker record. The flow ID for the forward and reverse data message flows are the same as the flow ID described for the connection tracker record 1410. In some embodiments, the pf_value is a value that identifies whether the flow should be allowed or dropped, bypassing the service node.
[0149] In some embodiments, the flow programming table 1129 stores state values 1440-1470. The state value "flow_program_gen" 1440 is a flow program state value used to identify changes to the state of the flow programming table. As described above, the flow_program_gen value is used to determine whether the flow programming table should be consulted (e.g., if the connection tracker record stores an outdated flow_program_gen value) to determine forwarding information for a data message, or whether the forwarding information stored in the connection tracker record is current (e.g., the connection tracker record stores a current flow_program_gen value).
[0150] The state value "BFD_gen" 1450 is a liveness state value used to identify changes to the liveness value of a service node connected using a L2 BIW tunneling mechanism. Similar to the flow_program_gen value, the BFD_gen value is used to determine whether the BFD_gen value stored in the connection tracker record is a current BFD_gen value and the forwarding information is still valid, or whether the BFD_gen value stored in the connection tracker is outdated and the SI preprocessor needs to determine whether the forwarding information is still valid (e.g., to determine whether the service node corresponding to the stored IP address is still running). In some embodiments, a separate storage structure stores a list of failed service nodes using BFD (e.g., L2 BIW service nodes) to detect failures, which is consulted when the BFD_gen value stored in the connection tracker record does not match the global BFD_gen value.
[0151] The state value "SPI fail gen" 1460 is an up status value that identifies a change in the up status of a service path (i.e., an ordered set of service nodes) connected to the LR-SR using a logical service plane (e.g., logical service forwarding element) transport mechanism. In some embodiments, the SPI fail gen value is provided from a controller that implements a central control plane that is aware of service node failures and updates the SPI fail gen value when a service node failure is detected. Similar to the BFD gen value, SPI fail gen is used to determine whether the SPI fail gen value associated with a service path identifier associated with a UUID in the policy store is up to date. If the SPI fail gen value is not up to date, then it must be determined whether the service path currently listed as a possible service path is still operational. In some embodiments, a separate storage structure stores a list of failed service paths that is consulted when the SPI fail gen value is not up to date (i.e., does not match the stored SPI fail gen state value 1460).
[0152] The state value "SN gen" 1470 is an up status value that identifies a change in the up status of a service node connected using an L3 tunnel transport mechanism. Similar to the flow program gen value, the SN gen value is used to determine whether the SN gen value stored in a connection tracker record is the current SN gen value and the forwarding information is still valid, or whether the SN gen value stored in the connection tracker is stale and the SI preprocessor needs to determine whether the forwarding information is still valid (e.g., to determine whether the service node corresponding to the stored IP address is still operational). In some embodiments, a separate storage structure stores a list of failed L3 service nodes that is consulted when the SN gen value stored in a connection tracker record does not match the global SN gen value.
[0153] The flow program table 1129 also stores a set of flow program instructions. In some embodiments, a single flow program instruction received from a service node (through its service agent) generates a flow program record for each of a forward and reverse data message flow. The set of flow program records 1480 illustrates flow program records that update the pf_value for a forward data message flow identified by flow ID 1 (1481) and a reverse data message flow identified by flow ID 1' (1482). In some embodiments, flow ID 1 and flow ID 1' are identical except for a bit that identifies the flow ID as a forward or reverse data message flow. In some embodiments, the pf_value' included in the flow program table record 1480 is an action value that specifies that data messages of the data message flow should be dropped or allowed.
[0154] In some embodiments, the flow programming instructions are indicated by a flow programming tag, which can specify the following actions (1) a NONE action when no action is needed (this results in no flow programming action being performed), (2) a DROP action when other data messages of the flow should not be forwarded along the service chain, but should be discarded at the LR-SI classifier, and (3) an ACCEPT action when other data messages of the flow should not be forwarded along the service chain, but the flow should be forwarded by the LR-SR to the destination. In some embodiments, the flow programming tag can also specify a DROP_MESSAGE. The DROP_MESSAGE is used when the service node needs to communicate with the proxy (e.g., to respond to a ping request) and wishes to discard the user data message (if any) even though flow programming is not needed at the source.
[0155] In some embodiments, an additional action is available to the service proxy for internally communicating a failure of its SVM. In some embodiments, this action will direct the SI preprocessor to select another service path (e.g., another SPI) for the flow of data messages. In some embodiments, this action is carried in-band with the user data message by setting appropriate metadata fields in some embodiments. For example, as described further below, the service proxy communicates with the SI postprocessor (or controller computer responsible for generating and maintaining a list of available service paths) through OAM (operations, administration, and maintenance) metadata of the NSH attribute through in-band data message traffic on the data plane. Assuming that the flow programming action is subject to signaling delays and is susceptible to loss by design, the SVM or service proxy can still see data messages at the source belonging to a flow that is expected to be discarded, accepted, or redirected for some period of time after the flow programming action is communicated to the proxy. In this case, the service plane should continue to set the action to discard, allow, or redirect at the LR-SI classifier (or connection tracker record).
[0156] The set of flow programming records 1490 show flow programming records that update the set of service metadata for the forward data message flow identified by flow ID 2 (1491) and the reverse data message flow identified by flow ID 2' (1492). In some embodiments, the updated SPI (e.g., SPI 2 or SPI 2') represents a different set of service nodes. As described above, the updated service path can be based on a service node failure or based on a determination that a particular service node is no longer needed (e.g., a service node that provides a firewall decision that is suitable for all subsequent data messages that allow data messages).
[0157] Reference is made to Figure 15 and 16 Additional details related to service chain and service path creation and management are discussed. Figure 15An object data model 1500 for some embodiments is shown. In this model, objects shown in solid lines are provided by the user, while objects shown in dashed lines are generated by the service plane manager and controller. As shown, these objects include service manager 1502, service 1504, service profile 1506, vendor template 1507, service attachment 1508, service instance 1510, service deployment 1513, service instance runtime (SIR) 1512, instance endpoint 1514, instance runtime interface 1516, service chain 1518, service insertion rule 1520, service path 1522, and service path hop 1524.
[0158] In some embodiments, the service manager object 1502 can be created before or after the service object 1504 is created. An administrator or service management system can invoke the service manager API to create a service manager. The service manager 1502 can be associated with a service at any point in time. In some embodiments, the service manager 1502 includes service manager information such as vendor name, vendor identifier, restUrl (for callbacks), and authentication / certificate information.
[0159] As noted above, the service plane does not require the presence or use of a service manager, as service nodes can operate in a zero-aware mode (i.e., with zero awareness of the service plane). In some embodiments, the zero-aware mode allows only basic operations (e.g., redirecting traffic to the SVM of a service). In some such embodiments, there is no provision for integration to distribute object information (such as service chain information, service profiles, etc.) to service manager servers. Instead, these servers can poll the network manager for objects of interest.
[0160] The service object 1504 represents a type of service provided by a service node. The service object has a transport type attribute that specifies the mechanism it uses to receive service metadata (e.g., NSH, GRE, QinQ, etc.). Each service object also has a status attribute (which can be enabled or disabled) returned by the service manager, and a reference to the service manager that can be used to expose REST API endpoints to communicate events and perform API calls. It also includes a reference to the OVA / OVF attribute used to deploy the service instance.
[0161] The provider template object 1507 includes one or more service profile objects 1506. In some embodiments, the service manager can register provider templates, and service profiles can be defined on a per-service basis and based on provider templates with potentially specialized parameters. In some embodiments, a provider template object 1507 is created for L3 routing services, which can be used to represent LR-SR components with attributes that can be used to distinguish different LR-SR components of an edge forwarding element. Service chains can be defined by referencing one or more service profiles. In some embodiments, service profiles are not assigned labels and are not explicitly identified on a wire. To determine which functions to apply to traffic, service nodes perform a lookup (e.g., based on service chain identifier, service index, and service direction, as described above) to identify the applicable service profile. Whenever a service chain is created or modified, the mapping for this lookup is provided to the service manager by the management plane.
[0162] In some embodiments, the service profile object 1506 includes (1) a provider template attribute to identify its associated provider template, (2) one or more custom attributes when the template exposes configurable values through the service profile, and (3) action attributes such as forwarding action or replicate and redirect, which indicate that service agents forward received data messages to their service nodes, or forward a copy of received data messages to their service nodes, while forwarding the received data messages to the next service hop or back to the original source GVM when their service node is the last hop, respectively.
[0163] The service attachment object 1508 represents the service plane (i.e., is a representation of the service plane from the perspective of the user, such as a network administrator of a tenant in a multi-tenant data center or a network administrator in a private data center). The service attachment object is an abstraction that supports any number of different implementations of the service plane (e.g., logical L2 overlays, logical L3 overlays, logical network overlays, etc.). In some embodiments, each endpoint (on a service instance runtime (SIR) or GVM) that communicates through the service plane specifies a service attachment. The service attachment is the communication domain. Thus, services or GVMs outside of a service attachment can not be able to communicate with each other.
[0164] In some embodiments, service attachments can be used to create multiple service planes with hard separation between them. A service attachment has the following attributes (1) a logical identifier that identifies the logical network or logical forwarding element that carries the service attachment's traffic (e.g., SVNI of a logical switch), (2) the type of service attachment (e.g., L2 attachment, L3 attachment, etc.), and (3) an applied_To identifier that specifies the scope of the service attachment (e.g., transport node 0 and transport node 1 for north-south operations and a host cluster or set for east-west operations). In some embodiments, a control plane (e.g., a central control plane) translates the service attachment representation it receives from the management plane into a specific LFE or logical network deployment based on parameters specified by a network administrator (e.g., a data center administrator of a private or public cloud, or a network virtualization provider in a public cloud).
[0165] A service instance object 1510 represents an actual deployed instance of a service. Thus, each such object is associated with one service object 1504 by a service deployment object 1513 that specifies the relationship between the service object 1504 and the service instance object 1510. The deployed service instance can be a standalone service node (e.g., a standalone SVM), or it can be a high-availability (HA) service node cluster. In some embodiments, the service deployment object 1513 describes the service instance type, e.g., standalone or HA. As described below, an API of the service deployment object can be used in some embodiments to deploy several service instances for a service.
[0166] A service instance runtime (SIR) object 1512 represents an actual runtime service node operating in standalone mode, or an actual runtime service node of an HA cluster. In some embodiments, the service instance object includes the following attributes (1) a deployment mode attribute that specifies whether the service instance is operating in standalone mode, active / standby mode, or active / active mode, (2) a state attribute that specifies whether the instance is enabled or disabled, and (3) a deployed_to attribute that includes a reference to a service attachment identifier in the case of north-south operations.
[0167] In some embodiments, SVM provisioning is initiated manually. To this end, in some embodiments, the management plane provides APIs for (1) creating service instances of existing services, (2) deleting service instances, (3) growing service instances that have been configured as high-availability clusters by adding additional SIRs, and (4) shrinking service instances by removing one of their SIRs. When creating service instances of existing services, in some embodiments, the service instance can be created based on the template contained in the service. The caller can choose between a standalone instance or an HA cluster, in which case all VMs in the HA cluster will be provisioned. Also, in some embodiments, the API for service instance deployment allows deploying multiple service instances (e.g., for an HA cluster) with only one API call.
[0168] In some embodiments, the API that creates one or more SVMs specifies one or more logical locations (e.g., clusters, hosts, resource pools) where the SVMs should be placed. In some embodiments, the management plane attempts to place SVMs belonging to the same service instance on different hosts whenever possible. Anti-affinity rules can also be configured as appropriate to maintain the distribution of SVMs across migration events, such as VMotion events supported by VMware's Dynamic Resource Scheduler. Similarly, the management plane can configure affinity rules with specific hosts (or groups of hosts) when they are available, or the user provisioning the service instance can explicitly select the hosts or clusters.
[0169] As noted above, service instance runtime objects 1512 represent the actual SVMs running on hosts to implement the service. In embodiments where the LR-SR provides L3 routing services, service instance runtime objects 1512 also represent the edge forwarding elements. A SIR is a part of a service instance. Each SIR can have one or more traffic interfaces that are completely dedicated to service plane traffic. In some embodiments, each SIR runs at least one service agent instance to handle data plane signaling and data message format translation for the SIR as needed. When a service instance is deployed, in some embodiments, a SIR is created for each SVM associated with the service instance. The network manager also creates an instance endpoint for each service instance in the east-west service insertion. In some embodiments, each SIR object 1512 has the following attributes (1) a state attribute that is active for the SVMs that can handle traffic, and non-active for all other SVMs, regardless of the reason, and (2) a runtime state that specifies whether data plane liveness detection is on or off for the SIR.
[0170] An instance runtime interface 1516 is per-endpoint version of a service instance endpoint 1514. In some embodiments, the instance runtime interface 1516 is used to identify the interface of a SIR or GVM that can be a source or sink of service plane traffic. In east-west service insertion, in some embodiments, the lifecycle of the instance runtime interface is linked to the lifecycle of the service instance runtime. In some embodiments, no user action is required to configure the instance runtime interface.
[0171] In some embodiments, the instance runtime interface 1516 has the following attributes: endpoint identifier, type, reference to service attachment, and location. The endpoint identifier is a data plane identifier of the SIR VNIC. The endpoint identifier is generated when the SIR or GVM registers to the service transport layer and can be a MAC address or a portion of a MAC address. The type attribute can be shared or dedicated. SIR VNICs are dedicated, which means only service plane traffic can reach them, while GVM VNICs are shared, which means they will receive and send both service plane and regular traffic. The service attachment reference is a reference to the service attachment that implements the service plane for sending and receiving service plane traffic. In some embodiments, the reference is an SVNI for the service plane. In some embodiments, the location attribute specifies the location of the instance runtime interface, which is the UUID of the host where the instance runtime interface is currently located.
[0172] In some embodiments, the user defines service chain objects 1518 from an ordered list of service profiles 1506. In some embodiments, each service chain conceptually provides separate paths for forward and reverse traffic directions, but if only one direction is provided at creation, the other direction is automatically generated by reversing the service profile order. Either direction (or even both directions) of a service chain can be empty, which means no service will handle traffic in that direction. In some embodiments, even for empty service chains, the data plane will perform lookups.
[0173] Service chains are abstract concepts. They do not point to a specific set of service nodes. Instead, a network controller that is part of the service plane platform automatically generates a service path that points to a sequence of service nodes for a service chain and directs messages / flows along the generated service path. In some embodiments, a service chain is identified in the management plane or control plane by its UUID, which is a unique identifier for the service chain. Service nodes are provided with the meaning of the service chain ID through the management plane API that they receive through their service manager. More details are described in U.S. Patent Application 16 / 444,826, filed on June 18, 2019.
[0174] In some embodiments, a service chain tag can be used to identify a service chain in the data plane, as UUIDs are too long to be carried in encapsulation headers. In some embodiments, the service chain ID is a similar unsigned integer to the rule ID. Each data message that is redirected to a service carries the service chain tag of the service chain it traverses. When a service chain is created or modified, the management plane advertises the UUID to the service chain tag mapping. The service chain tag has a 1-to-1 mapping with the service chain UUID, while a single service chain can have 0 to many service path indexes.
[0175] In addition to the service chain ID, in some embodiments, a service chain has the following attributes: (1) a reference to all computed service paths, (2) a failure policy, and (3) a reference to a service profile. The reference to the computed service paths is described above. The failure policy is applied when it is not possible to traverse the service path selected for the service chain. In some embodiments, the failure policy can be PASS (forward traffic) and FAIL (drop traffic). The reference to the service profile for the service chain can include an egress list of service profiles that egress traffic (e.g., data messages traveling from GVMs to switches) must traverse, and an ingress list of service profiles that ingress traffic (e.g., data messages traveling from switches to GVMs) must traverse. In some embodiments, by default, the ingress list is initialized as the inverse of the egress list.
[0176] Different techniques can be used in some embodiments to define the service paths of a service chain. For example, in some embodiments, a service chain can have an associated load balancing policy, which can be one of the following policies. The load balancing policy is responsible for load balancing traffic across different service paths of the service chain. According to the ANY policy, the service framework is free to redirect traffic to any service path regardless of any load balancing considerations or flow pinning. Another policy is the LOCAL policy, which specifies that local service instances (e.g., SVMs executing on the same host computer as the source GVM) are to be preferred over remote service instances (e.g., SVMs executing on other host computers or external service devices).
[0177] Some embodiments generate a score for a service path based on how many SIRs are local and select the highest scoring, regardless of load. Another policy is the CLUSTER policy, which specifies that service instances implemented by VMs co-located on the same host are preferred, whether that host is the local host or a different host. The ROUND_ROBIN policy indicates that all active service paths are hit with equal probability or based on probabilities specified by a set of weight values.
[0178] The SI rule object 1520 associates a set of data message attributes with a service chain represented by the service chain object 1518. The service chain is implemented by one or more service paths, each defined by a service path object 1522. Each service path has one or more service hops represented by one or more service path hop objects 1524, each associated with an instance runtime interface 1516. In some embodiments, each service hop also references an associated service profile, an associated service path, and a next hop SIR endpoint identifier.
[0179] In some embodiments, the service path object has several attributes, some of which can be updated by the management or control plane when underlying conditions change. These attributes include a service path index, a state (e.g., enabled or disabled), a management mode (e.g., enable or disable) to use when the service path must be manually disabled (e.g., for debugging reasons), a host cross count (indicating how many times a data message traversing the service path crosses a host), a location count (indicating how many SIRs along the path are located on the local host), a backup service path list, a length of the service path, a reverse path (listing the same set of SIRs in reverse order), and a maintenance mode indicator (a bit that is true if any of the hops in the service path are in maintenance mode).
[0180] The host cross count is an integer and indicates how many times a data message traversing the service path must be issued from a PNIC. In some embodiments, this metric is used by the local or central control plane to determine a preferred path when multiple alternatives are available. The value is populated by the management plane or control plane and is the same for every host using the service path. In some embodiments, the locality count is not initialized by the management plane or control plane, but is computed by the local control plane when the service path is created or updated. Each LCP can compute a different number. This value is used by the local control plane to identify a preferred path when multiple alternatives are available. The service path length is a parameter used by the service plane to set an initial traffic index.
[0181] In some embodiments, the backup service path list is a pointer to an ordered list of all service paths for the same service chain. It lists all possible alternatives to try when a particular SIR along the path fails. The list can contain service paths for all possible permutations of SVMs in each HA cluster that the service path traverses. In some embodiments, the list will not contain SIRs belonging to different HA clusters.
[0182] In some embodiments, a service path is disabled when at least one service hop is not active. This condition is temporary and triggered by a service activity detection failure. A service path can be disabled in this way at any time. In some embodiments, a service path is also disabled when at least one service hop has no matching SIR. A service hop enters this condition when the SIR it refers to disappears, but the service path still exists in the object model.
[0183] The service plane must be able to uniquely identify each SPI. In some embodiments, a UUID generated by the control plane is sent for each service path. Due to data message header limitations in the service plane, in some embodiments, a large ID is not sent with each data message. In some embodiments, when the control plane generates a UUID for each service path, in these embodiments, it also generates a small unique ID for it, and this ID is sent with each data message.
[0184] To support using LR-SRs as service plane traffic sinks, in some embodiments, the network manager or controller generates an internal service representing the edge forwarding element and creates a vendor template representing L3 routing with configurable settings representing the LR-SR. For each LR-SR, in some embodiments, the network manager or controller creates (1) a service profile dedicated to the L3 routing vendor template, (2) a service instance, and (3) a service instance endpoint. The network manager or controller then allows the service profile in a service chain and configures a failure policy for the service path that includes the LR-SR. The LR-SR is then provided with a service link to the logical service plane, and the data plane is configured to inject service plane traffic into the LR-SR’s regular routing pipeline.
[0185] Figure 16 A number of operations that the network manager and controller perform in some embodiments to define rules for service insertion, next service hop forwarding, and service processing are conceptually illustrated. As shown, these operations are performed by a service registrar 1604, a service chain creator 1606, a service rule creator 1608, a service path generator 1612, a service plane rule generator 1610, and a rule distributor 1614. In some embodiments, each of these operations can be implemented by one or more modules of the network manager or controller, and / or can be implemented by one or more standalone servers.
[0186] Through a service partner interface 1602 (e.g., a set of APIs or a partner user interface (UI) portal), the service registry 1604 receives vendor templates 1605 that specify services performed by different service partners. These templates define partner services in terms of one or more service descriptors, including service profiles. The registry 1604 stores the service profiles in a profile store 1607 for use by a service chain creator 1606 in defining service chains.
[0187] In particular, through a user interface 1618 (e.g., a set of APIs or a UI portal), the service chain creator 1606 receives one or more service chain definitions from a network administrator (e.g., a data center administrator, a tenant administrator, etc.). In some embodiments, each service chain definition associates a service chain identifier for a service chain with an ordered sequence of one or more service profiles. Each service profile in a defined service chain is associated with a service operation that needs to be performed by a service node. The service chain creator 1606 stores the definition of each service chain in a service chain store 1620.
[0188] Through the user interface 1618 (e.g., a set of APIs or a UI portal), the service rule creator 1608 receives one or more service insertion rules from a network administrator (e.g., a data center administrator, a tenant administrator, etc.). In some embodiments, each service insertion rule associates a set of data message flow attributes with a service chain identifier. In some embodiments, the flow attributes are flow header attributes, such as L2 attributes or L3 / L4 attributes (e.g., five-tuple attributes). In these or other embodiments, the flow attributes are context attributes (e.g., AppID, process ID, active directory ID, etc.). Various techniques for capturing and using context attributes to perform forwarding and service operations are described in U.S. Patent Application 15 / 650,251, now published as U.S. Patent Publication 2018 / 0181423, the contents of which are incorporated herein. Any of these techniques can be used in conjunction with the embodiments described herein.
[0189] The service rule creator 1608 generates one or more service insertion rules and stores these rules in a SI rule store 1622. In some embodiments, each service insertion rule has a rule identifier and a service chain identifier. In some embodiments, the rule identifier can be defined in terms of flow identifiers (e.g., header attributes, context attributes, etc.) that identify the data message flow(s) to which the SI rule applies. On the other hand, the service chain identifier for each SI rule identifies the service chain that must be performed by the service plane for any data message flow that matches the rule identifier for the SI rule.
[0190] For each service chain that is part of a service rule, the service path generator 1612 generates one or more service paths, where each path identifies one or more service instance endpoints of one or more service nodes to perform the service operations specified by the sequence of service profiles of the chain. In some embodiments, the process of generating service paths for a service chain takes into account one or more criteria, such as (1) data message processing load on the service nodes (e.g., SVMs) that are candidate service nodes for the service path, (2) the number of host computers that data messages of the flow traverse as they pass through each candidate service path, and so on.
[0191] The generation of these service paths is further described in U.S. Patent Application 16 / 282,802, which is incorporated by reference herein. As described in this patent application, some embodiments identify service paths for a particular GVM on a particular host based on one or more metrics, such as a host traversal count (indicating how many times data messages traversing the service path traverse the host), a locality count (indicating how many SIRs along the path are located on a local host), and so on. Other embodiments identify service paths (i.e., select service nodes for the service path) based on other metrics, such as financial and licensing metrics.
[0192] The service path generator 1612 stores the identities of the generated service paths in a service path storage 1624. In some embodiments, this storage associates each service chain identifier with one or more service path identifiers, and for each service path (i.e., each SPI), it provides a list of service instance endpoints that define the service path. Some embodiments store the service path definitions in one data storage, while storing the associations between service chains and their service paths in another data storage.
[0193] The service rule generator 1610 then generates rules for service insertion, next service hop forwarding, and service processing from the rules stored in storage 1620, 1622, and 1624, and stores these rules in rule storage 1626, 1628, and 1630, from which the rule distributor 1614 can retrieve and distribute them to the SI preprocessor, service agents, and service nodes. In some embodiments, the distributor 1614 also distributes the path definitions from the service path storage 1624. The path definitions in some embodiments include the first-hop network address (e.g., MAC address) of the first hop along each path. In some embodiments, the service rule generator 1610 and / or the rule distributor 1614 specifies different sets of service paths for the same service chain and distributes them to different host computers, as different sets of service paths are optimal or preferred for different host computers.
[0194] In some embodiments, the SI classification rules stored in the rules storage 1626 associate flow identifiers with service chain identifiers. Thus, in some embodiments, the rules generator 1610 retrieves these rules from the storage 1622 and then stores them in the classification rules storage 1626. In some embodiments, the rules distributor 1614 retrieves the classification rules directly from the SI rules storage 1622. For these embodiments, the description of the SI classification rules storage 1626 is more of a conceptual illustration to highlight the three types of distribution rules and the next-hop forwarding rules and service node rules.
[0195] In some embodiments, the service rules generator 1610 generates next-hop forwarding rules for each hop service agent of each service path of each service chain. As noted above, in some embodiments, the forwarding table of each service agent has a forwarding rule that identifies the next-hop network address of each service path on which the agent’s associated service node resides. Each such forwarding rule maps a current SPI / SI value to a next-hop network address. The service rules generator 1610 generates these rules. For embodiments in which the SI preprocessor must look up the first-hop network address, the service rules generator also generates first-hop lookup rules for the SI preprocessor.
[0196] In addition, in some embodiments, the service rules generator 1610 generates service rules for service nodes that map service chain identifiers, service index values, and service directions to the service profiles of the service nodes. To do so, the service rules generator uses the service chain and service path definitions from the storages 1620 and 1624, and the service profile definitions from the service profile storage 1607. In some embodiments, when there is a service manager for the service nodes, the rules distributor forwards the service node rules to the service nodes through such service manager. In some embodiments, the service profile definitions are also distributed by the distributor 1614 to the host computers (e.g., their LCPs), so that these host computers (e.g., LCPs) can use these service profiles to configure their service agents, e.g., to configure the service agents to forward received data messages to their service nodes, or to duplicate received data messages and forward the copies to their service nodes while forwarding the original received data messages to their next service node hop or back to their source GVM when they are the last hop.
[0197] In some embodiments, the management and control plane dynamically modifies the service paths of a service chain based on the status of the service nodes of the service paths and the data message processing load on these service nodes, as described in U.S. Patent Application 16 / 444,826, filed June 18, 2019. In some embodiments, Figure 16The components of the system are also used to configure logical forwarding elements to use service chains.
[0198] Figure 17 A process 1700 is conceptually illustrated for configuring a logical forwarding element (e.g., a virtual routing and forwarding (VRF) context) to connect to a logical service forwarding plane. In some embodiments, the process 1700 is performed by a network controller computer to provide configuration information to an edge device to configure an edge forwarding element to connect to a logical service forwarding plane. The process begins by identifying (at 1710) a logical forwarding element to connect to a logical service forwarding plane. In some embodiments, the logical forwarding element is a logical router component (e.g., an AZG-SR, an AZG-DR, a VPCG-SR, or a VPCG-DR). In some embodiments, the logical router component is implemented as a virtual routing and forwarding (VRF) context.
[0199] For the identified logical forwarding element, a set of services available at the identified logical forwarding element is identified (at 1720). In some embodiments, the set of services available at the logical forwarding element is defined by an administrator or controller computer based on service insertion rules applicable at the logical forwarding element. In some embodiments, the set of services defines a set of service nodes (e.g., service instances) that connect to the logical service forwarding plane to provide the set of services.
[0200] Once the set of services is identified (at 1720), the process 1700 identifies (at 1730) a logical service forwarding plane that connects the logical forwarding element and the service nodes to provide the identified set of services. In some embodiments, the logical service forwarding element is identified by a service virtual network identifier (SVNI) selected from a plurality of SVNIs used from the logical network. In some embodiments, the set of service nodes that provide the identified services connect to a plurality of logical service forwarding planes identified by the plurality of SVNIs. In some embodiments, different SVNIs are used to distinguish between traffic for different tenants.
[0201] The process 1700 then generates (at 1740) configuration data to configure the logical forwarding element to connect to the identified logical service forwarding plane. In some embodiments, the configuration data includes an interface mapping table that maps the logical forwarding element (e.g., a VRF context) to interfaces of the logical service forwarding plane. In some embodiments, the logical forwarding element uses the interface mapping table to identify interfaces for forwarding data messages to service nodes that connect to the logical service forwarding plane.
[0202] Then, the process 1700 determines (at 1750) whether additional logical forwarding elements need to be configured to connect to the logical service forwarding plane. If additional logical forwarding elements are needed, the process 1700 returns to operation 1710 to identify the next logical forwarding element that needs to be connected to the logical service forwarding elements. If no additional logical forwarding elements need to be configured, the process 1700 provides (at 1760) configuration data to the set of edge devices on which the identified set of logical forwarding elements (e.g., edge forwarding elements) are implemented. In some embodiments, the configuration data includes service insertion data for configuring the logical forwarding elements as described above, and also includes service forwarding data for configuring the logical software forwarding elements that implement the logical service forwarding plane associated with the logical forwarding elements implemented by the set of edge devices.
[0203] Figure 18 A set of operations performed by the service router at T0 or T1 (e.g., TX SR 1130) for the service insertion layer and the set of service transport layer modules 1135 invoked for the first data message 1810 in a data message flow that requires service from a service node reachable through a tunneling mechanism (e.g., a virtual private network) are shown. The SI preprocessor performs the basic operations for service classification as described above Figure 11 Figure 18 The UUID identifies a virtual tunnel interface (VTI) or other identifier for the service node accessed through the VPN. In some embodiments, the UUID is associated with a set of multiple service nodes and a selection metric. The selection metric can be a selection metric used for a load balancing operation that is any one of: a round robin mechanism, a load based selection operation (e.g., selecting the service node with the lowest current load), or a distance based selection operation (e.g., selecting the closest service node as measured by the selected metric).
[0204] Once the service node is selected, the process identifies the forwarding information associated with the selected service node by performing a lookup in the forwarding table 1138. The forwarding table 1138 stores the forwarding information (e.g., IP address of the VTI) for the service nodes. In other embodiments, the IP address associated with the selected service node is stored in the policy table 1137 along with the VTI or service node identifier, and the forwarding table 1138 is unnecessary.
[0205] In some embodiments, selecting a service node for the forward data message flow includes selecting the same service node for the reverse data message flow. In such embodiments, the forwarding information for each direction is determined at this time (e.g., the IP address of the selected service node). Once the service node is selected and the forwarding information is identified, a connection tracker record is created for both the forward and reverse flows and provided to the connection tracker storage 1121. As described below, the connection tracker record includes the forwarding information (e.g., the IP address of the interface), the service action if one is defined for the data message flow, and the service insertion rule identifier for the service insertion rule identified as matching the attributes of the data message 1810. In some embodiments, the connection tracker record includes the service insertion type identifier. In some embodiments, the service node state value (e.g., SN_gen) is included in the connection tracker record as described above with reference to Figure 11 and Figure 14 The information stored in the connection tracker record is used to process subsequent data messages in the forward and reverse data message flows.
[0206] The data message 1822 is then provided to the STL module 1122 along with the forwarding information 1821. In this example, the forwarding information for data messages requiring service provided by the service node accessed through the VPN includes, in some embodiments, the next hop IP address for the virtual tunnel interface and the service insertion type identifier identifying the data message as using the tunnel transport mechanism.
[0207] As shown, the service routing processor 1125 routes the data message to the VTI based on the IP address identified by the SI preprocessor 1120. In some embodiments, the data message 1831 is provided to the VTI along with the original source and destination IP addresses and the original data message source and destination ports. In other embodiments, the destination IP address is changed to the IP address of the VTI with the original destination IP address stored in the metadata storage of the edge forwarding element for use by the edge forwarding element to restore the destination IP address of the serviced data message after receiving the serviced data message from the service node. The VTI receives the data message 1831 and, in some embodiments, processes the pipeline encrypts and encapsulates the data message to be delivered over the VPN as data message 1851. The return data message is then received at the VTI and processed as described above for return data messages of Figure 11 .
[0208] Figure 19A set of operations performed by the service router at T0 or T1 (e.g., TX SR 1130) for the service insertion layer and the service transport layer module set 1135 invoked for a first data message 1910 in a data message flow that requires service from a service node reachable through the L2 BIW mechanism is shown. The basic operations of the SI preprocessor for service classification are as above Figure 11 are described.
[0209] Figure 19 A UUID identifies a service node accessed through the L2 BIW transport mechanism. In some embodiments, the UUID is associated with a set of multiple service nodes and a selection metric. The selection metric can be a selection metric for a load balancing operation that is any one of: a round robin mechanism, a load based selection operation (e.g., select the service node with the lowest current load), or a distance based selection operation (e.g., select the closest service node as measured by the selected metric).
[0210] Once a service node is selected, the process identifies forwarding information associated with the selected service node by performing a lookup in a forwarding table 1138. The forwarding table 1138 stores forwarding information for service nodes (e.g., a set of pseudo IP addresses for the interfaces of the TX SR 1130). In some embodiments, the pseudo IP addresses are a set of source and destination IP addresses associated with first and second virtual interfaces (VIs) of the BIW interface pair 1126 that are each connected to the same service node. In other embodiments, the pseudo IP addresses associated with the selected service node are stored in the policy table 1137 along with the service node identifier, and the forwarding table 1138 is unnecessary.
[0211] In some embodiments, selecting a service node for a forward data message flow includes selecting the same service node for a reverse data message flow. For L2 BIW, in some embodiments, the forwarding information for the forward data message flow identifies the same pseudo IP addresses for the forward data message flow, but identifies the source IP address of the forward data message as the destination IP address of the reverse data message, and identifies the destination IP address as the source IP address. Once a service node is selected and forwarding information is identified, a connection tracker record is created for the forward and reverse flows and provided to the connection tracker storage 1121. As described below, the connection tracker record includes the forwarding information (e.g., the pseudo IP addresses of the destination interfaces), the service action if one is defined for the data message flow, and the service insertion rule identifier for the service insertion rule identified as matching the attributes of the data message 1910. In some embodiments, the connection tracker record includes a service insertion type identifier. In some embodiments, a service node state value (e.g., BFD_gen) is included in the connection tracker record as described above with reference to Figure 11and 14 The information stored in the connection tracker record is used to process subsequent data messages in both the forward and reverse data message flows as described. The information stored in the connection tracker record is used to process subsequent data messages in both the forward and reverse data message flows.
[0212] The data message 1922 is then provided to the STL module 1122 along with the forwarding information 1921. In this example, for data messages requiring a service provided by a service node accessed through an L2 BIW connection, in some embodiments, the forwarding information includes a next hop pseudo IP address set for the virtual interface and a service insertion type identifier for identifying the data message as using the L2 BIW transport mechanism.
[0213] As shown, the STL module 1122 provides the data message to the interface in the BIW interface pair 1126 identified as the source interface based on the pseudo IP address identified by the SI preprocessor 1120. In some embodiments, the data message 1932 is provided to the source interface in the BIW interface pair 1126 (associated with MAC address MAC 1) with the original source and destination IP addresses, but with the source and destination MAC addresses of the BIW interface pair 1126 associated with the L2 BIW service node. The data message is then processed by the L2 service node, which returns the serviced data message to the interface in the BIW interface pair identified as the destination interface (associated with MAC address MAC 2). The returned data message is then processed as described above for the return data message of Figure 11
[0214] As described above, in some embodiments, the transport mechanism includes a tunneling mechanism (e.g., virtual private network (VPN), Internet Protocol Security (IPSec), etc.) that connects the edge forwarding element to at least one service node through a corresponding set of virtual tunnel interfaces (VTIs). In addition to the VTIs used to connect the edge forwarding element to the service nodes, the edge forwarding element uses other VTIs to connect to other network elements for which it provides forwarding operations. At least one of the VTIs used to connect the edge forwarding element to other (i.e., non-service node) network elements is identified to perform service classification operations and is configured to perform service classification operations on data messages received at the VTI for forwarding. In some embodiments, the VTIs used to connect the edge forwarding element to the service nodes are not configured to perform service classification operations, but are configured to mark data messages returned to the edge forwarding element as serviced. In other embodiments, the VTIs used to connect the edge forwarding element to the service nodes are configured to perform limited service classification operations using a single default rule applied to the VTIs that marks data messages returned to the edge forwarding element as serviced.
[0215] For traffic that exits the logical network through a particular VTI, some embodiments perform a service classification operation on different data messages to identify different VTIs that connect the edge forwarding element to service nodes to provide services required by the data messages. In some embodiments, each data message is then forwarded to the identified VTI to receive the required services (e.g., from service nodes connected to the edge forwarding element through the VTIs). The identified VTIs do not perform service classification operations and only allow the data messages to reach the service nodes. The service nodes then return the serviced data messages to the edge forwarding element. In some embodiments, the VTIs are not configured to perform service classification operations, but are configured to mark all traffic directed from the service nodes to the edge forwarding element as serviced. The marked serviced data messages are then received at the edge forwarding element and forwarded through the particular VTIs to the destinations of the data messages. In some embodiments, the particular VTIs do not perform additional service insertion operations because the data messages are marked as serviced.
[0216] In some embodiments, the service classification operations are implemented separately from service classification operations for non-tunneled traffic received at the uplink interface of the edge forwarding element. In some embodiments, the different implementations are due to the fact that tunneled data messages are received at the uplink interface in encapsulated and encrypted format, which would result in incorrect service classification (e.g., incorrectly identifying the necessary set of services and forwarding information for the underlying (encapsulated) data message flow) if processed by the uplink service classification operations. Accordingly, some embodiments implement the service classification operations as part of the VTI data path after the incoming data messages have been decapsulated (and decrypted, if needed), or for outgoing data messages before encryption and encapsulation.
[0217] Figure 20A - Figures 2A-B conceptually illustrate data message flows through the above-described system. Figure 20A - Figure 2A conceptually illustrates data messages sent from a compute node 2060 in a logical network 2003 (e.g., logical network A) implemented in a cloud environment 2002 to a compute node 2080 in an external data center 2001. The compute node 2080 in the data center 2001 connects to the logical network using a VPN (i.e., tunneling mechanism) 2005 through an external network 2004 to the logical network 2003. In some embodiments, the tunnel uses a physical interface that is identified as an uplink interface of an edge device that performs the edge forwarding element, but is logically identified as a separate interface of the edge forwarding element. To conceptually clarify, different logical interfaces and associated service classification operations are presented to represent the logical structure of the network. In addition, internal elements of the data center 2001 beyond the tunnel endpoint and the destination compute node 2080 are also omitted for clarity.
[0218] Communication from compute node 2060 to 2080 begins with standard logical processing by the elements of logical network 2003. Thus, compute node 2060 uses logical switch 2050 to provide the data message to tenant distributed router 2040. In some embodiments, both logical switch 2050 and tenant distributed router 2040 are implemented by local managed forwarding elements on the same host as compute node 2060. In some embodiments, VPC distributed router 2040 in turn routes the data message to a transit logical switch (not shown) using the VPC service router 2030 as described with reference to Figures 4-6 VPC service router 2030 routes the data message to availability zone distributed router 2020. As described above, in some embodiments, VPC service router 2030 executes on a first edge device that also implements availability zone distributed router 2020, and in other embodiments, executes on the same edge device as availability zone service router 2010. In some embodiments, availability zone distributed router 2020 in turn routes the data message to a transit logical switch (not shown) using the availability zone service router 2010 as described with reference to Figures 4-6
[0219] Availability zone service router 2010 then routes the data message to the VTI that is the next hop for the data message. As part of the VTI processing pipeline, SI classifier 2007 (e.g., VTI-SI classifier) performs a service classification operation prior to encryption and encapsulation that identifies that the data message requires a service provided by L3 service node 2070 that is located outside of logical network 2003 based on service insertion rules applied at the VTI. The SI classifier identifies the VTI associated with VPN 2006 as the next hop to L3 service node 2070 and sends the data message for processing. The SI classifier located between availability zone service router 2010 and VPN 2006 does not perform a service classification operation for service insertion traffic, and the data message reaches L3 service node 2070 that performs the service for the data message.
[0220] Figure 20B The data message that has been serviced is shown being returned to Availability Zone Service Router 2010 for routing to destination compute node 2080 via VPN 2005. Although shown as a post-service insertion service from L3 service node 2070, in some embodiments, the marking of the data message as serviced (i.e., post-SI) is completed at the SI classifier located between VPN 2006 and Availability Zone Service Router 2010, based on a default rule that is the only SI rule applied at the SI classifier. In other embodiments, the marking is part of a processing pipeline configured for each interface connected to an L3 service node that does not have a service classification operation. For this data message, the SI classifier located between Availability Zone Service Routers 2010 does not perform a second service classification operation based on the data message marked as serviced, and the data message is processed (encapsulated or encrypted and encapsulated) to be delivered to compute node 2080 via VPN 2005. In some embodiments where data messages are marked as serviced using tags, further pipeline processing removes the tags after bypassing the SI classification operation based on the tags. In other embodiments, the data message is marked as served by a tag stored in the local metadata associated with the data message, and the tag is deleted once the data message has been processed and delivered to the external network at the availability zone service router 2010.
[0221] Figure 21A -B conceptually illustrates a data message sent from a computing node 2080 in an external data center 2001 to a computing node 2060 in a logical network 2103 (e.g., logical network A) implemented in a cloud environment 2002. Figure 20A -B and Figure 21A The components of -B are the same, and if in Figure 20A Data messages sent in -B are considered to be forward data message streams, so in Figure 21A The data message sent in -B can be considered a reverse data message stream. Communication begins with compute node 2080 sending a data message to the tunnel endpoint in data center 2001 connected to VPN 2005 (again, ignoring the internal components of data center 2001). The data message is encapsulated (or encrypted and encapsulated) and sent over external network 2004 using VPN 2005. The data message is then logically processed to reach VTI and undergo VTI's processing pipeline. The data message is decapsulated and, if necessary, decrypted. At this point, SI classifier 2007 performs a service classification operation to determine whether the data message requires any service.
[0222] SI classifier 2007 determines the service that the data message requires to be provided by L3 service node 2070 based on service insertion rules applied at VTI. The SI classifier identifies the VTI associated with VPN 2006 as the next hop to L3 service node 2070 and sends the data message for processing. SI classifiers located between availability zone service router 2010 and VPN 2006 do not perform service classification operations on service insertion traffic and the data message reaches L3 service node 2070 that performs the service on the data message.
[0223] Figure 21B The serviced data message is shown returning to availability zone service router 2010 to be routed through elements of logical network 2103 to destination compute node 2060. Although shown as post-service insertion traffic from L3 service node 2070, in some embodiments, marking the data message as serviced (i.e., post-SI) is done at SI classifiers located between VPN 2006 and availability zone service router 2010 based on a default rule that is the only SI rule applied at the SI classifiers. In other embodiments, the marking is part of the processing pipeline configured for each interface connected to L3 service nodes that do not have service classification operations. In some embodiments that use a tag to mark the data message as serviced, availability zone service router 2010 processes removing the tag before forwarding the data message to availability zone distributed router 2020. As discussed below with respect to FIG. 22, in some embodiments, the tag is a label that is removed by availability zone distributed router 2020 before forwarding the data message to compute node 2060. Figure 31 As discussed, some embodiments require the serviced tag to traverse logical router boundaries to avoid redundant service classification operations at VPC service routers. In other embodiments, marking the data message as serviced is a tag stored in local metadata associated with the data message and the tag is removed once the data message has completed processing to be delivered to the next hop router component (e.g., availability zone distributed router 2020). The data is then delivered to compute node 2060 through logical network that includes VPC service router 2030, VPC distributed router 2040, and logical switch 2050.
[0224] Figure 22 A first method for providing service for a data message at an uplink interface of a set of uplink interfaces is shown conceptually. In some embodiments, the data message is received from a source in external network 2004 and destined for a destination in external network 2004, but requires a service provided at an edge forwarding element of logical network 2203. Figure 22The services in the illustrated embodiment are provided by the service chain service nodes 2270a-c using a logical service forwarding plane transport mechanism (e.g., logical service forwarding element 2209). Those of ordinary skill in the art will appreciate that alternative transport mechanisms are used in other embodiments. In the illustrated embodiment, a data message arrives at a first uplink interface with the external network 2004, and a service classification operation occurs at the SI classifier 2007 that determines a set of services that are required based on service insertion rules that apply to the data message received at the uplink interface, and identifies forwarding information (e.g., SPI, next hop MAC, etc. as described above) to access the required set of services.
[0225] In the illustrated embodiment, the service classification operation is provided prior to the routing operation by the availability zone service router 2010. Based on the identified forwarding information, the availability zone service router 2010 provides the data message to the service chain service node 2270a to provide a first service to the data message, and passes the data message to the next hop in the service path (i.e., service chain service node 2270b). In some embodiments, the availability zone service router 2010 identifies the service chain service node functionality provided by the availability zone service router 2010 as the first hop, which then routes the data message to the service chain service node 2270a. In either embodiment, upon receiving the data message from the service chain service node 2270a, the service chain service node 2270b provides the next service in the service chain, and provides the data message to the service chain service node 2270c to provide additional services, and identifies the service chain service node functionality provided by the availability zone service router 2010 as the last hop in the service path. Each data message sent between service chain service nodes (e.g., SVMs) uses the logical service forwarding element 2209, and in some embodiments involves a service broker and service transport layer module, which are not shown here for clarity. The use of service brokers and service transport layer modules is described in more detail above with reference to Figure 11 and related U.S. Patent Application 16 / 444,826.
[0226] The serviced data message is then routed by the availability zone service router 2010 to a destination in the external network 2004. The routing identifies a second uplink interface with the external network 2004, and provides the serviced data message with a tag or metadata that identifies the data message as a serviced data message. Based on the identification, the service classifier at the second uplink interface does not provide an additional service classification operation, and the data message is forwarded to the destination. As described above, in some embodiments that use a tag to identify the data message as a serviced data message, the tag is removed prior to sending the data message through the uplink interface.
[0227] Figure 23 A second method for providing services for a data message at an uplink interface in a set of uplink interfaces is conceptually illustrated. In some embodiments, the data message is received from a source in the external network 2004 and destined for a destination in the external network 2004, but requires services provided at the edge forwarding elements of the logical network 2203. Figure 23 The services in the illustrated embodiment are provided by the service chain service nodes 2270a-c using a logical service forwarding plane transport mechanism (e.g., logical service forwarding elements 2209). Those of ordinary skill in the art will appreciate that alternative transport mechanisms are used in other embodiments. In the illustrated embodiment, the data message arrives at a first uplink interface with the external network 2004, and the service classification operation at the SI classifier 2007 fails to identify any required set of services, as the service classification rules are defined only for data messages received (in ingress or egress direction) at the second uplink interface. Thus, in this embodiment, the SI classifier of the first uplink interface provides the data message to the availability zone service router 2010 without service insertion forwarding information. The availability zone service router 2010 routes the data message to the second uplink interface based on the destination IP address of the data message, and the SI classifier of the second uplink interface determines that a set of services is required based on the service insertion rules applicable to data messages received at the second uplink interface, and identifies forwarding information (e.g., SPI, next hop MAC, etc. as described above) to access the required set of services. The remainder of the data message processing is as described above. Figure 22
[0228] Figure 24 A logical network 2203 providing service classification operations at multiple routers of the logical network is conceptually illustrated. As Figure 22 illustrated, a first service classification operation is performed by the availability zone service router 2010 prior to the availability zone service router 2010 identifying the set of services required for the data message. In this example, the set of services includes services provided by service chain service nodes 2270a and 2270b. As described above, the availability zone service router 2010 router provides the data message to the service chain service node 2270a, which provides services, and provides the serviced data message to the service chain service node 2270b, which provides additional services and returns the data message to the availability zone service router 2010. In the illustrated embodiment, the availability zone service router 2010 removes the label identifying the data message as a serviced data message, and forwards the data message to the VPC service router 2030 (through the availability zone distributed router 2020).
[0229] Before being routed by VPC service router 2030, the SI classifier 2007 associated with VPC service router 2030 performs a service classification operation. This operation determines the required service set based on service insertion rules applicable to data messages received at the uplink interface of VPC service router 2030 and identifies forwarding information (e.g., SPI, next-hop MAC, etc. as described above) to access the required service set. The data message is provided to VPC service router 2030, which uses the forwarding information to provide the data message to service chain service node 2270c. Service chain service node 2270c returns the serviced data message to VPC service router 2030. VPC service router 2030 then routes the serviced data message to the destination compute node 2060.
[0230] Figure 25 A conceptual illustration shows an edge forwarding element (AZG Serving Router 2010) connected to the serving node 2570a-e using multiple transport mechanisms. Logical network 2503 includes... Figure 20A -B The same logical edge forwarding elements: Availability Zone Serving Router 2010, Availability Zone Distributed Router 2020, VPC Serving Router 2030, and VPC Distributed Router 2040. In some embodiments, the different router components are each defined as separate VRF contexts. Figure 25 The dashed lines and dashed boxes in the diagram represent edge devices that implement different edge forwarding component assemblies. In the illustrated embodiment, different edge devices implement availability zone and VPC service routers, but the availability zone distributed router is implemented by both availability zone edge devices and VPC edge devices, respectively, for ingress and egress data messages, as shown in the reference. Figure 10 As explained above. Similarly, the VPC distributed router 2040 is implemented by both VPC edge devices and hosts, respectively, for ingress and egress data messages. As shown, the Availability Zone Service Router 2010 (1) connects to the service chain service node set 2570a-c via service link 2508 on the logical service forwarding plane (or logical service plane (LSP)) 2509, (2) connects to the L3 service node set 2570d via VPN set 2505, and (3) connects to the L2 BIW service node set 2570e via interface set. The Availability Zone Service Router 2010 uses service nodes to provide services, as referenced above. Figure 11 , 12 As described in 18 and 19. Because different data messages require different services provided by different types of service nodes, some embodiments use multiple service transport mechanisms to access different types of service nodes to provide services, such as Figure 25 As shown.
[0231] Figure 26 and Figure 27 A logical network in which multiple logical service forwarding planes are configured for different service routers is conceptually illustrated. Figure 26 A logical network 2603 is shown that includes three VPC service routers 2630 belonging to two different tenants. The logical network 2603 also shows three logical service forwarding planes 2609a-c connected to the VPC service routers 2630, one of which (2609c) is also connected to an availability zone service router 2010. The different logical service forwarding planes 2609a-c are connected to different sets of service chain service nodes. In Figure 26 In embodiments of the logical network 2603, the service chain service nodes 2670a-c are used by the VPC service routers 2630 of tenant 1, while the service chain service nodes 2670d-g are shared by the availability zone service router 2010 and the VPC service routers 2630 of tenant 2.
[0232] Figure 27 A logical network 2703 is shown that includes three VPC service routers 2630 belonging to three different tenants. The logical network 2703 also shows four logical service forwarding planes 2709a-c connected to the VPC service routers 2630, one of which (2709c) is also connected to an availability zone service router 2010 and a logical service forwarding plane 2709d, which is a second logical service forwarding plane connected only to the availability zone service router 2010. The logical service forwarding planes 2709a and 2709b are connected to a common set of service chain service nodes (2770a-c), while the logical service forwarding planes 2709c and 2709d are connected to different sets of service chain service nodes. In Figure 27 In embodiments of the logical network 2703, the service chain service nodes 2770a-c are used by the VPC service routers 2630 of both tenant 1 and tenant 2. Although the shared service chain nodes 2770a-c are used by two different tenants, the data message traffic of each tenant is kept separate by using different logical service forwarding planes 2709a and 2709b. As Figure 26 The service chain service nodes 2770d-g are shared by the availability zone service router 2010 and the VPC service routers 2630 of tenant 3, as shown. However, in Figure 27 The availability zone service router 2010 has a second logical service forwarding plane 2709d connected to it that is not shared by the VPC service routers 2630, as in Figure 30 and 31As discussed, in some embodiments, if the availability zone service router 2010 is configured to provide L3 routing services as a service chain service node, the service chain service node 2770h-j is accessible to the VPC service router 2630 for tenant 3.
[0233] Figure 28 A process for accessing services provided at availability zone edge forwarding elements from a VPC edge forwarding element is conceptually illustrated. In some embodiments, the process 2800 is performed by a VPC edge forwarding element (e.g., a VPC service router). In some embodiments, the process is performed by a service classification operation of a VPC edge forwarding element. The process 2800 begins by receiving (at 2810) a data message at an uplink interface of a VPC service router. In some embodiments, the data message is received from the VPC service router after a routing operation of the VPC service router, while in other embodiments, the data message is received from an availability zone distributed router.
[0234] The service classification operation determines (at 2820) that the data message requires a service provided at an availability zone service router. In some embodiments, the service classification operation performs the operations of process 100 to determine that the data message requires the service and identifies (at 2830) forwarding data for the data message. In some embodiments, the forwarding information identified for the data message includes service metadata (SMD) for sending the data message to the availability zone service router over a logical service forwarding plane, and additional service metadata for directing the availability zone service router to redirect the data message to a particular service node or set of service nodes. In some embodiments, the additional service metadata takes the form of an argument to a function call to a function exposed at the availability zone service router.
[0235] The data message is then sent (at 2840) to the availability zone service router over the logical service forwarding plane with the service metadata identifying the additional service required. Figure 29 A process 2900 to be performed as part of process 2800 when an availability zone service router receives a data message from a VPC service router is conceptually illustrated. The process 2900 begins by receiving (at 2910) a data message to be serviced that was sent (at 2840) from a VPC service router. The data message is received at a service link of the availability zone service router over a logical service forwarding plane.
[0236] Once the availability zone service router receives the data message, the availability zone service router determines (at 2920) that the data message requires routing service to at least one additional service node. In some embodiments, the determination is made based on additional metadata provided by the VPC service router, while in other embodiments, the determination is made based on arguments to a function call to a function (e.g., an API) available at the availability zone service router.
[0237] Once (at 2920) it is determined that the data message requires routing service to at least one additional service node, the availability zone service router provides service (at 2930) based on the received metadata. In some embodiments, the service is provided by the service node functionality of the service router, and the service is provided without redirection. In other embodiments, the service is provided by a set of service nodes reachable through one of the above-mentioned transport mechanisms, and the data message is redirected to the service nodes using the appropriate transport mechanism. Once the data message is redirected to the service nodes, the process proceeds very similarly to the process of 18 or 19, depending on the transport mechanism used to redirect the data message. Figure 11 、 12
[0238] The serviced data message is received (at 2940) at the availability zone service router, and is marked as serviced. As noted above, in this embodiment, the identification of the data message as serviced must be carried through the availability zone service router and the distributed router, so that the SI classifier of the VPC service router does not apply the same service insertion rules and redirect the data message to the same destination in a loop. In embodiments where the availability zone service router and the VPC service router are implemented in the same edge device, the metadata identifying the data message as serviced is stored in a shared metadata storage device that is used by the VPC service router SI classifier to identify the data message as serviced.
[0239] The data message is then routed to the destination (at 2950). In some embodiments, the data message is routed to an external destination or a VPC service router different from the VPC service router that sent the data message to the availability zone service router. In other embodiments, for data messages that originally went to a compute node in a network segment reached through a VPC service router (e.g., southbound data messages), the data message is routed to the VPC service router from which the data message was received. After routing the data message, the process 2900 ends. In embodiments where the data message is routed to the VPC service router from which the data message was received, the process 2800 receives a served data message identified as a served data message (at 2850) and the SI classifier does not perform a service classification operation based on the tag. The data message is then received at the VPC service router and routed to the destination of the data message.
[0240] Figure 30 A VPC service router 3030 is conceptually shown processing a data message sent from a first compute node 3060a to a second compute node 3060b in a second network segment served by a second VPC service router 3030. Prior to encountering the SI classifier 2007, the data message is processed through a logical switch 3050 connected to the source compute node 3060a, a VPC distributed router 3040, and the VPC service router 3030, which determines as described above with reference to Figure 28 the data message should be sent to an availability zone service router 2010 for an L3 service node 3070 to provide service. The data message is then sent through a logical service forwarding plane to the availability zone service router 2010, which redirects the direction to the identified service node (e.g., L3 service node 3070). The data message is then returned to the availability zone service router 2010 and routed to the compute node 3060b as described above.
[0241] In some embodiments, sending the data message to the availability zone service router 2010 using the logical service forwarding plane includes sending the data message through a Layer 2 interface 3005 of a software switch executing on the same device as the service router. The software switch is used to implement a logical service forwarding element (e.g., LSFE 801) represented by the LSP 3009. In some embodiments, the connection to the AZG service router 2010 is mediated by a service proxy implemented by the AZG service router 2010 to conform to an industry standard service insertion protocol.
[0242] Figure 31The VPC service router 3030 is conceptually shown processing a data message sent from the external network 2004 to the compute node 3060. Before encountering the SI classifier 2007, the data message is processed by the availability zone service router 2010 and the availability zone distributed router 2020, which are described above with reference to Figure 28 The SI classifier 2007 determines, as described above, that the data message should be sent to the availability zone service router 2010 for the L3 service node 3070 to provide service. The data message is then returned to the availability zone service router 2010 and routed to the compute node 3060. In routing the data message to the compute node 3060, the data message passes through the SI classifier 2007, but no service classification operation is performed because the data message is identified as a serviced data message that does not require a service classification operation. The data message is then processed by the VPC service router 3030, the VPC distributed router 3040, and the logical switch 3050 and delivered to the destination compute node 3060.
[0243] Some embodiments facilitate provisioning of services reachable at a virtual internet protocol (VIP) address. A client uses the VIP address to access a set of service nodes in a logical network. In some embodiments, a data message from a client machine to the VIP is directed to an edge forwarding element at which the data message is redirected to a load balancer that load balances among the set of service nodes to select a service node to provide a service requested by the client machine. In some embodiments, the load balancer does not change the source IP address of the data message received from the client machine so that the service node receives the data message to service with the client machine IP address identified as the source IP address. The service node services the data message and sends the serviced data message to the client machine using the IP address of the service node as the source IP address and the IP address of the client node as the destination IP address. Because the client sent the original address to the VIP address, the client will not identify the source IP address of the serviced data message as a response to a request sent to the VIP address and the serviced data message will not be properly processed (e.g., it will be discarded or not associated with the original request).
[0244] In some embodiments, facilitating the provisioning service includes using a service logic forwarding element to return the serviced data message to the load balancer to track the state of the connection. To use the service logic forwarding element, some embodiments configure the egress data path of the service node to intercept the serviced data message before forwarding it to the logical forwarding element in the data path from the client to the service node, and determine whether the serviced data message needs to be routed by the routing service provided by the edge forwarding element as a service. If the data message needs to be routed by the routing service (e.g., for serviced data messages), the serviced data message is forwarded by the service logic forwarding element to the edge forwarding element. In some embodiments, the serviced data message is provided to the edge forwarding element with the VIP associated with the service, and in other embodiments, the edge forwarding element determines the VIP based on the port used to send the data message on the service logic forwarding element. The edge forwarding element uses the VIP to identify the load balancer associated with the serviced data message. The serviced data message is then forwarded to the load balancer for the load balancer to maintain state information for the connection to which the data message belongs, and to modify the data message to identify the VIP as the source address for forwarding to the client.
[0245] Figure 32A -B shows a set of data messages for providing a service addressable at a VIP to clients served by the same virtual private cloud gateway (e.g., VPCG service and distributed router). Figure 32A A logical network 3203 is shown that includes two logical switches 3250a and 3250b served by the same VPC service router 3230. The logical switch 3250a is connected to a set of guest virtual machines (GVMs) 3261-3263 that provide a service reachable at a virtual IP (VIP) address. In some embodiments, the GVMs provide content rather than providing a service. The logical switch 3250b is connected to a client 3290 that accesses the service available at the VIP. Figure 32A A first data message is shown being sent from the client 3290 to the VIP address. The data message is forwarded to the VPC service router 3230, which identifies the load balancer 3271 as the next hop for the VIP address. The load balancer 3271 then performs a load balancing operation to select the GVM 3261 from the set of GVMs 3261-3263. The load balancer 3271 changes the destination IP address from the VIP to the IP address of the selected GVM 3261.
[0246] Figure 32BThe GVM 3261 is shown returning a served data message to the client device 3290. The served data message is intercepted at a service insertion (SI) preprocessor as described in U.S. Patent Application 16 / 444,826, which redirects the data message through the logical service forwarding plane 3209 to the VPC service router. In some embodiments, the preprocessor is configured to redirect all data messages through the logical service forwarding plane to the VPC service router. In some embodiments, the served data message is sent to the VPC service router with metadata identifying the VIP to which the data message was originally sent. In other embodiments, the VPC service router identifies the destination of the data message based on attributes of the data message (e.g., port or source address). The VPC service router 3230 routes the data message to the load balancer 3271. In some embodiments, the load balancer 3271 stores state information for the data message flow, which it uses to update the source IP address to the VIP address, and sends the data message to the client 3290 with the source IP address identified by the client 3290.
[0247] Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium). When these instructions are executed by one or more processing unit(s) (e.g., one or more processors, cores of processors, or other processing units), they cause the processing unit(s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard drives, EPROMs, etc. Computer readable media do not include carrier waves and electronic signals over wire, fiber optic, or other communication media.
[0248] In this specification, the term "software" is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger software
[0249] Figure 33A computer system 3300 with which some embodiments of the application are implemented is conceptually illustrated. The computer system 3300 can be used to implement any of the host, controller, and manager described above. Thus, it can be used to perform any of the processes described above. The computer system includes various types of non-transitory machine-readable media and interfaces for various other types of machine-readable media. The computer system 3300 includes a bus 3305, processing unit(s) 3310, a system memory 3325, a read-only memory 3330, a permanent storage device 3335, input devices 3340, and output devices 3345.
[0250] The bus 3305 generally governs communication between the various internal devices of the computer system 3300. For instance, the bus 3305 communicatively connects the processing unit(s) 3310 with the read-only memory 3330, the system memory 3325, and the permanent storage device 3335.
[0251] The processing unit(s) 3310 retrieve instructions and data from these various memory units to execute the processes of the application. In different embodiments, the processing unit(s) can be single or multi-core processors. The read-only memory (ROM) 3330 stores static data and instructions that are needed by the processing unit(s) 3310 and other modules of the computer system. The permanent storage device 3335, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the computer system 3300 is off. Some embodiments of the application use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 3335.
[0252] Other embodiments use a removable storage device (such as a flash drive) as the permanent storage device. Like the permanent storage device 3335, the system memory 3325 is a read-and-write memory device. However, unlike the permanent storage device 3335, the system memory is a volatile read-and-write memory, such as a random access memory. The system memory stores some of the instructions and data that the processor needs at runtime. In some embodiments, the processes of the application are stored in the system memory 3325, the permanent storage device 3335, and / or the read-only memory 3330. From these various memory units, the processing unit(s) 3310 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.
[0253] Bus 3305 also connects to input and output devices 3340 and 3345. Input devices enable the user to communicate information and select commands to the computer system. Input devices 3340 include alphanumeric keyboards and pointing devices (also referred to as "cursor control devices"). Output devices 3345 display information generated by the computer system. Output devices include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD). Some embodiments include devices such as a touchscreen that functions as both input and output devices.
[0254] Finally, as shown in Figure 33 Bus 3305 also couples computer system 3300 to a network 3365 through a network adapter (not shown) that connects to bus 3305. In this manner, the computer can operate in a networked environment (using
[0255] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine -readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media). Some examples of a computer-readable medium include RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD-RW), read-only digital versatile discs (e.g., DVD- ROM, dual-layer DVD-ROM), a variety of recordable / rewritable DVD (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc.), magnetic and / or solid state hard drives, read-only and recordable Blu-Ray® discs, ultra-density optical discs, and any other optical or magnetic media. The computer-readable medium can store the computer program instructions by using various technologies and / or formats, including, for example, information formats such as eXtensible Markup Language (XML), and / or compressed formats, such as Hyper Text Markup Language (HTML), Hyper Text Preprocessor (HTP), and / or Hypersonic Markup Language (HTML5). The computer program instructions can be executed by at least one processing unit of the computer system to provide various aspects of the functionality described above. The computer program instructions can be stored in a computer-readable medium that is non-transitory, meaning that the instructions are present on a tangible and / or physical computer-readable medium with the actual instructions (as opposed to a signal based, propagated message, for example). With reference to the computer system of FIG. 33, a computer program product, such as the computer program instructions 3320, can be stored in the memory 3325 of the computer system 3300. The computer program instructions 3320 can be executed by the processing unit(s) 3310 to cause the computer system 3300 to perform various computer-implemented steps, including the steps disclosed herein.
[0256] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, some embodiments are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself.
[0257] As used in this specification, the terms "computer," "server," "processor," and "memory" all refer to electronic or other technological devices. These terms exclude people or groups of people. For purposes of this specification, the terms display or displaying mean displaying on an electronic device. As used in this specification, the terms "computer-readable medium," "computer readable medium," and "machine readable medium" are entirely restricted to tangible, physical objects that store information in a
[0258] While the application has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the application can be embodied in other specific forms without departing from the spirit of the application. For instance, several of the figures conceptually illustrate processes. The specific operations of these processes can not be performed in the exact order shown and described. The specific operations can not be performed in one continuous sequence, and different specific operations can be performed in different embodiments. Furthermore, various additional operations can be performed, or described operations can be eliminated, in different embodiments.
[0259] Even though the service insertion rules in several of the examples above provide service chain identifiers, some of the inventions described herein can be implemented by having the service insertion rules provide service identifiers (e.g., SPIs) for different services specified by the service insertion rules. Similarly, several of the embodiments above perform distributed service routing by performing exact matching based on SPI / SI values, which relies on identifying the next service hop at each service hop. However, some of the inventions described herein can be implemented by having the service insertion preprocessor embed all service hop identifiers (e.g., service hop MAC addresses) as a set of service attributes of the data message and / or in encapsulation service headers of the data message.
[0260] In addition, some embodiments decrement the SI value in a different manner (e.g., at a different time) than the above-described methods. In addition, some embodiments do not perform next hop lookups based on just the SPI and SI values, but rather perform the lookups based on the SPI, SI, and service direction values, as these embodiments use a common SPI value for both the forward and reverse directions of data messages flowing between two machines.
[0261] The above-described approach is used in some embodiments to express path information in a single-tenant environment. Thus, one of ordinary skill will recognize that some embodiments of the present application are equally applicable to single-tenant data centers. Conversely, in some embodiments, the above-described approach is used to carry path information across different data centers of different data center providers when one entity (e.g., one company) is a tenant in multiple different data centers of different providers. In these embodiments, the tenant identifier embedded in the tunnel header must be unique across data centers or must be translated when they traverse from one data center to the next. Thus, those of ordinary skill in the art will appreciate that the present application is not limited by the foregoing illustrative details, but rather by the following claims.
Claims
1. A method for providing multiple services at a router of a data center, the method comprising: at the router, performing, for a data message received for routing, a service classification operation to determine a particular service chain for which multiple service operations must be performed on the data message; for the particular service chain, selecting a service path that provides the multiple services from a plurality of service paths associated with the particular service chain by performing a load balancing operation; sending the data message along the selected service path so that the multiple services are performed; receiving a serviced data message from a service node that performed a last service operation and determining whether the serviced data message includes flow programming instructions that indicate conditions for skipping or modifying service nodes in the service path for one or more subsequent data messages in a data message flow; and performing a next hop forwarding on the serviced data message. the router of the data center is a logical router of a logical network implemented by physical forwarding elements of the data center.
2. The method of claim 1, wherein, performing the next hop forwarding identifies a next hop within the logical network as the next hop for the serviced data message.
3. The method of claim 2, wherein, performing the next hop forwarding identifies a next hop in an external network as the next hop for the serviced data message.
4. The method of claim 2, wherein, the logical router is an edge router at a boundary between the logical network and an external network.
5. The method of claim 2, wherein, the serviced data message is a data message that crosses a boundary between the logical network and the external network.
6. The method of claim 5, wherein, 7. A method for providing multiple services at a router of a data center, the method comprising: at the router, performing, for a data message received for routing, a service classification operation to determine a particular service chain for which multiple service operations must be performed on the data message; for the particular service chain, selecting a service path that provides the multiple services, wherein the service path includes a plurality of service nodes connected to a logical service forwarding plane; sending the data message along the logical service forwarding plane to forward the data message through the selected service path so that the multiple services are performed on the data message, the logical service forwarding plane being implemented by two or more forwarding elements to forward the data message to two or more service nodes of the selected service path; receiving a serviced data message from a service node that performed a last service operation and determining whether the serviced data message includes flow programming instructions that indicate conditions for skipping or modifying service nodes in the service path for one or more subsequent data messages in a data message flow; and performing a next hop forwarding on the serviced data message. the logical service forwarding plane is implemented as a service logical forwarding element using a service virtual network identifier. the plurality of service nodes includes at least one of a service virtual machine and a service appliance.
8. The method of claim 7, wherein, each service node is associated with a service agent that operates between the logical service forwarding plane and the service node to facilitate implementing the service path.
9. The method of claim 7, wherein, 10. The method of claim 7, wherein, 11. A machine-readable medium storing a program which, when implemented by at least one processing unit, carries out the method according to any one of claims 1-10.
12. An electronic device, comprising: a set of processing units; and a machine-readable medium storing a program which, when implemented by at least one of the processing units, carries out the method according to any one of claims 1-10.
13. A system comprising means for carrying out the method according to any one of claims 1-10.
14. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1-10.
Citation Information
Patent Citations
Service path computation for service insertion
US11012351B2
Providing services with guest VM mobility
US11042397B2
Collecting and processing contextual attributes on a host
US20180181423A1
Logical router with multiple routing components
US9787605B2
Distributing service function chain data and service function instance data in a network
EP3300319A1